LiveHermes
Talk out loud to Hermes — the agent that gets things done. Hands-free, in real time, from your phone.
LiveHermes puts a live voice conversation in front of Hermes, an autonomous agent with its own tools, skills and memory. Speak naturally and Hermes listens, works on the task and answers out loud as soon as its first sentence is ready; you can interrupt it mid-sentence, and point the phone camera at something so it can see it too. It remembers what you tell it from one turn to the next. Everything runs on our own hardware: Qwen3-ASR listens, Qwen3.8-27B inside Hermes thinks, and Kokoro-82M speaks. Because Hermes can act and not only answer, LiveHermes is private: it is available to the platform owner only, behind sign-in.
Without it, using an agent like Hermes means typing into a chat window and reading the reply — not a hands-free conversation while you are doing something else.
This isn’t a proof of concept.
What you’re about to see is a fully containerized, production/enterprise-grade solution — the same build that deploys to Kubernetes on any of the three major cloud providers: AWS, Azure, or Google Cloud. There is no gap between this demo and what ships to production.
Built in, not bolted on
Observability — SSE, Phoenix & Live Logging
Every step streams live to the screen as it happens and mirrors into a self-hosted Phoenix tracing dashboard, exactly like the platform's production solutions.
Bring Your Own Model (BYOM)
Switch the underlying model with one setting — a fully local, on-prem model, GPT, Gemini, Claude, Amazon Bedrock, or Microsoft Foundry — with no code change.
Built as a Comparison, Not a Guarded Production Agent
This one is intentionally a side-by-side teaching demo, so it skips the enterprise governance/PII/prompt-injection layer that wraps the platform's production solutions — every other solution on this page carries that full layer.
Role-Based Access Control
Every account carries a role — user, admin or owner — and the server checks it before a solution opens, so a correct password alone isn't enough. Your username and role sit beside Log out, and every run starts with “Applying RBAC based access control” in the live stream.
What it’s built with
Python backend, served with Uvicorn
A FastAPI service running on Uvicorn handles every request — the same production Asynchronous Server Gateway Interface (ASGI) stack used across the whole platform, not a notebook or a prototype script.
2-way voice
Live two-way voice and live video support — talk to the platform and interrupt it mid-answer, with your camera feed shared on request. Runs on Jiina, our open-source voice stack (Qwen3-ASR, Qwen3.8-27B and Kokoro-82M) inside our own infrastructure, or on external Gemini models, chosen per session.
FastMCP tool server — MCP v2 support
Database access, document tools and integrations are exposed through the Model Context Protocol via a dedicated FastMCP server running on MCP v2, so agents call real, typed tools instead of hand-rolled function stubs.
Large MCP Payload Support
We go beyond the default MCP payload of 5MB. We support MCP payloads of 5MB+.
Next.js frontend
A React/Next.js interface talks to the backend over Server-Sent Events for live streaming — no page reloads, no polling.
Choice of multi-agent framework
The same solution can run on LangGraph, Google ADK, or Microsoft Agent Framework — switchable per deployment, not hard-wired to one vendor's orchestration engine.
Agent2Agent (A2A) protocol support
Every agent pipeline also exposes a standards-based Agent-to-Agent (A2A) endpoint with a real, discoverable agent card — so another agent system can call it directly, not just this UI, including mid-conversation clarification round-trips.