IdeaNirvana
Super Agents

Jiina

Meet Jiina — Jovial Intelligence of Idea Nirvana AI. Talk to it like a person: real-time voice, in your language.

Jiina is a super agent developed by IdeaNirvana AI, the AI lab of the parent company IdeaNirvana LLC, registered in Ashburn, Virginia, USA, and in business since 2007. About IdeaNirvana AI

What it does

Jiina is where you talk to Idea Nirvana Lab’s super agent. Speak naturally and Jiina listens, thinks and answers out loud — it starts speaking as soon as its first sentence is ready, and you can interrupt it mid-sentence the way a real conversation works. It searches the web for anything current, does exact arithmetic, knows the time and the weather where you are, and answers in the language you speak. Jiina can also see: she looks through your camera, reasons about and understands what she sees, and answers your questions about it — just point the camera and ask. She can take a snapshot of what the camera shows. She can draw a picture from your description, restyle your camera picture into a cartoon or a watercolour, and make a short video of up to 15 seconds — realistic footage or an animated explainer — with her own narration. Just ask. She can also create Word documents, PowerPoint decks and Excel workbooks for you, read and summarise your company documents, answer questions from a database (read-only), and write and run code in a safe sandbox for multi-step jobs. Choose Jiina or Gemini’s live voice API as the engine underneath the same interface.

Without it

Without it, a voice interface here means speak, wait, get a reply, speak again — the stop-and-go rhythm of a walkie-talkie, not a conversation.

Meet Jiina

Jovial Intelligence of Idea Nirvana AI

Jiina is Idea Nirvana Lab’s super agent, built entirely on an open-source, Apache 2.0 licensed stack — five models that run on our own hardware, with no proprietary model needed to run Jiina.

Qwen3-ASR

Listens — speech recognition that also identifies the language you spoke

Apache 2.0

Qwen3.8-27B

Thinks — reasoning, tool use and the words of every answer

Apache 2.0

Kokoro-82M

Speaks — natural text-to-speech, sentence by sentence as the answer is written

Apache 2.0

FLUX.2 klein 4B

Draws — turns a description into a picture in about a second, and restyles what your camera sees

Apache 2.0

FastWan (Wan 2.1 1.3B)

Films — makes a realistic video of up to 15 seconds in 5-second scenes, narrated by Jiina

Apache 2.0

Languages Jiina speaks

Jiina understands and replies in all of these. Speak in any one of them and it answers in the same language; ask it to switch to another at any time and it keeps that language until you ask for a different one.

This isn’t a proof of concept.

What you’re about to see is a fully containerized, production/enterprise-grade solution — the same build that deploys to Kubernetes on any of the three major cloud providers: AWS, Azure, or Google Cloud. There is no gap between this demo and what ships to production.

Enterprise capabilities

Built in, not bolted on

Observability — SSE, Phoenix & Live Logging

Every step streams live to the screen as it happens and mirrors into a self-hosted Phoenix tracing dashboard, exactly like the platform's production solutions.

Bring Your Own Model (BYOM)

Switch the underlying model with one setting — a fully local, on-prem model, GPT, Gemini, Claude, Amazon Bedrock, or Microsoft Foundry — with no code change.

Built as a Comparison, Not a Guarded Production Agent

This one is intentionally a side-by-side teaching demo, so it skips the enterprise governance/PII/prompt-injection layer that wraps the platform's production solutions — every other solution on this page carries that full layer.

Role-Based Access Control

Every account carries a role — user, admin or owner — and the server checks it before a solution opens, so a correct password alone isn't enough. Your username and role sit beside Log out, and every run starts with “Applying RBAC based access control” in the live stream.

Tech stack

What it’s built with

Python backend, served with Uvicorn

A FastAPI service running on Uvicorn handles every request — the same production Asynchronous Server Gateway Interface (ASGI) stack used across the whole platform, not a notebook or a prototype script.

2-way voice

Live two-way voice and live video support — talk to the platform and interrupt it mid-answer, with your camera feed shared on request. Runs on Jiina, our open-source voice stack (Qwen3-ASR, Qwen3.8-27B and Kokoro-82M) inside our own infrastructure, or on external Gemini models, chosen per session.

FastMCP tool server — MCP v2 support

Database access, document tools and integrations are exposed through the Model Context Protocol via a dedicated FastMCP server running on MCP v2, so agents call real, typed tools instead of hand-rolled function stubs.

Large MCP Payload Support

We go beyond the default MCP payload of 5MB. We support MCP payloads of 5MB+.

Next.js frontend

A React/Next.js interface talks to the backend over Server-Sent Events for live streaming — no page reloads, no polling.

Choice of multi-agent framework

The same solution can run on LangGraph, Google ADK, or Microsoft Agent Framework — switchable per deployment, not hard-wired to one vendor's orchestration engine.

Agent2Agent (A2A) protocol support

Every agent pipeline also exposes a standards-based Agent-to-Agent (A2A) endpoint with a real, discoverable agent card — so another agent system can call it directly, not just this UI, including mid-conversation clarification round-trips.