Planner (Restaurant Reservations and Pizza Orders by Phone)
Tell it the restaurant or the pizza you want. Jiina picks up the phone and makes the call for you.
Planner handles the small errands that need a phone call. Say or type “book a table at Nobu in Dallas for two on the 12th at 7:30” or “order a large pepperoni pizza to my address”. The agents ask for whatever is missing, such as the day, the time, the city, the toppings or the delivery address, then look up the business’s phone number on the web and ask your go-ahead before dialing. Jiina, IdeaNirvana’s AI agent, then makes the call and says plainly that she is an AI calling for you. She asks for a confirmation number, or for how long the pizza will take, reads it back to be sure, and shows you the whole conversation as it happens. The number or delivery time is read out to you if you used your voice. During testing, calls are simulated or ring only the owner’s own phone.
Without it, a reservation or a delivery order still means finding the number, waiting on hold and repeating the same details to a stranger, for something that should take one sentence.
This isn’t a proof of concept.
What you’re about to see is a fully containerized, production/enterprise-grade solution — the same build that deploys to Kubernetes on any of the three major cloud providers: AWS, Azure, or Google Cloud. There is no gap between this demo and what ships to production.
Built in, not bolted on
Agent Governance
Every agent call runs inside a governor with its own time and token budget and a kill-switch — a runaway or misbehaving step is stopped automatically, not left to run up cost or produce a bad answer, and every call leaves an audit-trail row.
PII / PHI Safe
User input is screened for personally identifiable and health information before it ever reaches a model, so sensitive data doesn't leak into a prompt, a log, or a third-party LLM call by accident.
Prompt-Injection Safe
A dedicated classifier checks every input for attempts to hijack the agent's instructions before it's acted on — the kind of "ignore your previous instructions" attack that a plain chatbot has no defense against.
SQL-Injection Safe
Where a solution talks to a database, every generated query is checked against a strict allow-list before it runs — no destructive statement (drop, delete, update, alter) can reach the database, however it's phrased.
Observability — SSE, Phoenix & Live Logging
Every step an agent takes streams live to the screen as it happens (no blank-screen wait for a final answer), and mirrors into a self-hosted Phoenix tracing dashboard — full request timelines, per-agent spans, and real token/cost usage, visible in real time, not reconstructed after the fact from a log file.
QoS — LLM-as-Judge
Before an answer ever reaches the user, a second, independent model call checks it for safety and can block it outright; a separate quality pass then scores completeness, correctness and relevance in the background — real automated review, not a cosmetic "checking quality..." status line.
Self-Improving (RSI via SKILLS.md)
Agents record what they learn from real runs into version-controlled SKILL.md files, which future runs read back — the system gets measurably better at its job over time instead of staying frozen at its original prompt.
Bring Your Own Model (BYOM)
Switch the underlying model with one setting — a fully local, on-prem model, GPT, Gemini, Claude, Amazon Bedrock, or Microsoft Foundry — with no code change and no vendor lock-in. Every option is a genuinely working, tested path, not a stub.
Durable Memory Across Sessions
Give the agent a user ID and it remembers your last few questions and answers — a follow-up like "which of those spent the most?" resolves correctly days later, without re-explaining context every time you come back.
Tenant Info Isolation
That remembered history is scoped strictly per user ID at the database level — one person's session data is never visible to, or blendable with, another's, even on the same solution.
Role-Based Access Control
Every account carries a role — user, admin or owner — and the server checks it before a solution opens, so a correct password alone isn't enough. Your username and role sit beside Log out, and every run starts with “Applying RBAC based access control” in the live stream.
What it’s built with
Python backend, served with Uvicorn
A FastAPI service running on Uvicorn handles every request — the same production Asynchronous Server Gateway Interface (ASGI) stack used across the whole platform, not a notebook or a prototype script.
2-way voice
Live two-way voice and live video support — talk to the platform and interrupt it mid-answer, with your camera feed shared on request. Runs on Jiina, our open-source voice stack (Qwen3-ASR, Qwen3.8-27B and Kokoro-82M) inside our own infrastructure, or on external Gemini models, chosen per session.
FastMCP tool server — MCP v2 support
Database access, document tools and integrations are exposed through the Model Context Protocol via a dedicated FastMCP server running on MCP v2, so agents call real, typed tools instead of hand-rolled function stubs.
Large MCP Payload Support
We go beyond the default MCP payload of 5MB. We support MCP payloads of 5MB+.
Next.js frontend
A React/Next.js interface talks to the backend over Server-Sent Events for live streaming — no page reloads, no polling.
Choice of multi-agent framework
The same solution can run on LangGraph, Google ADK, or Microsoft Agent Framework — switchable per deployment, not hard-wired to one vendor's orchestration engine.
Agent2Agent (A2A) protocol support
Every agent pipeline also exposes a standards-based Agent-to-Agent (A2A) endpoint with a real, discoverable agent card — so another agent system can call it directly, not just this UI, including mid-conversation clarification round-trips.