IdeaNirvana
From the Lab

Voice and LLM on One GPU: 17 of 25 Multi-Agentic Solutions, Two-Way Voice

A 29x token-counting bug, an accent we couldn't prompt our way out of, a demo that died before anyone could speak, and a memory shortfall that got worse when we tried to fix it the obvious way. Four bugs, one GPU, and what each one taught us.

October 2, 2026 · 6 min read


IdeaNirvana Labs runs over 25 Multi-Agentic Solutions — reporting assistants, compliance checkers, a cloud architecture advisor, a wealth management dashboard — all backed by one on-premise GPU box, no cloud API bill that scales with usage. Text worked well. Voice didn't exist. This is the story of building it: four separate bugs that each looked like a different kind of problem until we found their root cause, and how 17 of the 25 solutions ended up with real two-way voice.

The solution was to replace one resident model, Qwen3.8-27B, with two: MiniCPM-o-4.5 (9B, voice) and Qwen3.5-9B (text reasoning, full 262K context length) — both running together on one DGX Spark, with real memory to spare.

Why one box, two models

The machine is an NVIDIA DGX Spark — 128GB of memory shared between CPU and GPU, one GPU. Before this project it ran Alibaba's Qwen3.8-27B. Adding voice meant either finding room for a second resident model, or finding one model that could do everything the first one did, plus listen and speak.

We first picked Qwen3-Omni-30B-A3B, with three internal stages — one for audio-to-text, one for reasoning, one for text-to-audio. We tested it properly before trusting it: real microphone input transcribed correctly, spoken answers generated naturally, a photo of a jacket and shirt correctly identified as two separate garments, a closed-book linked-list-reversal coding question answered correctly in under two seconds. Every isolated test passed. The real system still broke — just not on any of the things we'd tested.

Bug one: a 29x token-counting error that looked nothing like a counting bug

The first real question we ran through it — "how many advisors are there?" — failed with a budget error. Our governance layer gives each reasoning step a token allowance sized for its actual job, a few thousand tokens for a short question. This one claimed to need twelve times that.

We pulled the real system prompt, about 15.7KB of text, and fired it directly at the model with no framework in between. 3,522 tokens reported — completely normal. We tried it with thinking mode off: still normal. Through our own code's basic, non-streaming call path: still normal, 3,675 tokens. Through a raw streaming request with usage reporting turned on: 3,516 tokens. Four different angles, four unremarkable numbers, nowhere close to the ninety-nine thousand the real failure had reported.

The one thing none of those four tests exercised was our own production streaming code — the method every one of our 25 tools actually uses to get live, token-by-token output instead of one big blocking response. We added one line of logging to it and reran the identical question.

chunk 1:  input_tokens: 3819
chunk 2:  input_tokens: 3819
chunk 3:  input_tokens: 3819
  ... (29 chunks total, same number every time)
chunk 29: input_tokens: 3819, output_tokens: 147

There it was. The real prompt size — 3,819 tokens, matching every clean test — arrived on every single one of 29 streamed chunks, not once. And our code was adding up every chunk's number. Summing a number that's already a total is exactly as wrong as it sounds: 29 times 3,819 is 110,751 — almost precisely the garbage figure the real failure had reported.

The fix was two lines, branched by provider: keep summing for the provider that actually streams incremental deltas, and for everyone else — whether they report their total once or repeat it on every chunk — take the latest non-zero reading instead of adding anything up. One shared code path, every one of the 25 tools fixed simultaneously, because they all go through the same streaming method.

The lesson: a bug that only reproduces through the real system, never through an isolated test, usually means the isolated tests simply aren't exercising the real code path. Four clean results in a row didn't mean the bug wasn't real — it meant we were testing the wrong thing four times.

Bug two: the voice itself, and the point where prompting stopped being the answer

Once the counting bug was fixed, the voice pilot genuinely worked — real microphone input, real spoken answers, rolled out across every tool with a free-text question box. Living with it daily surfaced a different problem entirely: on English speech, the voice carried a heavy, persistent accent, and would occasionally drift into outright gibberish mid-sentence, especially around dollar figures.

So we went looking for a replacement voice model that would fit on the DGX Spark. MiniCPM-o-4.5 had full-duplex voice and was built on an entirely different text-to-speech foundation.

Three dependency errors before it would even start

The same inference engine advertised support for this new model out of the box. It did not run on the first attempt, the second, or the third — each one failed further in, inside the same vocoder-initialization step:

Once resolved, MiniCPM-o-4.5's weights — 9B parameters, a fraction of the 30B model it replaced — loaded quickly.

Audio in the wrong shape, twice

Even running, the first speech-generation request came back as plain text — no audio at all, as though the model had decided to just answer instead of speak. Two things turned out to be true that weren't true of the first voice model: triggering speech output needed a different request field entirely (a template flag, not the modality parameter we'd used before), and once audio did come back, it arrived as a second, separate entry in the response rather than attached to the text answer — code that only checked the first entry found nothing, silently.

Bug three: the bug that hid itself from its own error log

Testing the new voice model's own reasoning engine — built on the same architecture family as the text model we'd go on to pick for reasoning — surfaced the same category of problem for a third time: transcripts occasionally came back wrapped in visible reasoning text instead of a clean answer, something the model's "thinking mode" setting is supposed to prevent when turned off.

We had already turned thinking mode off, or thought we had. The actual bug was more specific: we'd been omitting the setting entirely rather than explicitly setting it to off, and it turns out those aren't the same thing. We confirmed this directly, reading the model's own chat template logic:

if enable_thinking is defined and enable_thinking is false:
    inject empty thinking block (forces no reasoning)
# if the key is simply absent: model decides for itself

Worse, one real transcript came back with an unclosed reasoning tag — no closing marker anywhere in the response, which meant our first attempt at a fix, stripping reasoning text found between opening and closing markers after the fact, couldn't have worked even in principle. The actual fix had to happen before generation, not after: explicitly pass the setting as false, every single call, on both voice models. Verified live — a response that previously took 34–44 seconds while reasoning silently ran in the background dropped to about 2 seconds once that reasoning pass was skipped entirely.

Bug four: a memory shortfall that got worse when we tried the obvious fix

With voice handled by MiniCPM-o-4.5, the platform's reasoning and coding model still needed room on the same 128GB of shared memory. The obvious first candidate was to squeeze in the existing Qwen3.8-27B, already compressed to a lean 4-bit format. Two attempts to fit it in, at two different context-length settings, both failed — and the second attempt, at a quarter the context length of the first, failed by a wider margin, not a narrower one. That single result ruled out our working theory in one line of evidence: if a long context window were the real memory cost, shrinking it four-fold should have helped, not hurt. Something close to a fixed cost, independent of context length, was the real constraint — almost certainly the serving engine's own graph-compilation buffers and the extra bookkeeping a speculative-decoding accelerator carries alongside the model.

The fix wasn't a smaller context window. It was a smaller model: Qwen3.5-9B, in the same architecture family, served through a simpler engine with no speculative-decoding accelerator adding its own fixed overhead. It came up clean on the first real attempt, and both models now run together, confirmed live, at the model's full 262K architectural context length — with real memory to spare.

What this all adds up to

Four bugs, each one looking at first like a different kind of problem — a counting error, a voice quality issue, a reasoning-mode toggle, a memory constraint — and each one actually traceable to the same underlying habit: trusting an assumption carried over from a different context without checking whether it still held. The streaming-sum logic was right for one provider and silently wrong for a new one. The reasoning-mode default worked when explicitly set and silently didn't when merely implied. The memory math assumed the obvious cost was the real cost, and it wasn't.

None of these were exotic failures. Rigorous engineering made each one findable by actually running the real system and reading what it reported, rather than trusting what seemed like it should be true.

Source: MiniCPM-o-Demo on GitHub.