about 18× real time
Transcription finishes before you do.
Talk naturally. Ginu delegates work to your agents and keeps track of everything they're doing,
and brings you in only when your judgment is needed.
Local speech. Memory in your own files. The Claude Code, Codex or Gemini account you have.

When several agents are working at once, the scarce resource is not compute, it is your attention. Every time you stop to read a terminal you pay a context switch, and the agents wait for you at exactly the point where you are slowest.
Saying “spawn someone to chase that flaky test, and show me the auth diff when it’s ready” takes three seconds and no context switch. Ginu runs both. The harder half is the other direction, and it is the part I actually built: it stops interrupting you while you are thinking. A routine completion waits for a gap in the conversation. A worker that has been blocked on you for three minutes earns the right to interrupt. Something irreversible does the opposite of shouting — it waits for a genuine break and is framed slowly, because a decision you cannot undo must never arrive as a yes-or-no that a reflexive “yeah” could approve.
You can believe that because none of it is a notification setting. Every piece of news is scored against the room — whether a turn is in flight, whether Ginu is already speaking, whether you are mid-sentence, how long it has been quiet — by a policy that is a file you can read. An agent cannot mark its own news urgent: it reports facts, and the policy turns facts into a number, which is the only reason everything here isn’t P1. And every decision it made is a line in a log with the evidence it had.
about 18× real time
Transcription finishes before you do.
about 4× real time
The first sentence is spoken while the rest arrives.
43 ms
Speech can be stopped mid-word.
Playback block size, not end-to-end interruption latency.
None.
Your voice stays off speech servers.
| What was measured, and how | Value |
|---|---|
| Transcription throughput against real-time audio: warm, timed in a real session on an M-series Mac. | about 18× real time |
| Speech synthesis throughput: warm, timed in a real session on an M-series Mac. | about 4× real time |
| Playback block size, 43 ms; a cut-in stops playback at the next block. This does not measure recognition or end-to-end interruption latency. | 43 ms |
| Time to the first agent response in per-turn timing logs, measured with Claude Code. Includes provider latency; it is not a local speech-processing figure. | about 3 s |
| Cloud round trips for speech, by architecture | none |
Measured on an M-series Mac. Yours will differ with hardware; the same timing line is written to your transcript on every turn.
File the next few things. The board shows what is running, what finished, and the one card that is blocked on an answer from you.
Queued tasks are records on disk, so they are still there after a restart. A high-priority one starts the moment you file it, and nothing already running is paused to make room.
One decision needs you. The rest keeps moving.
An illustrated example using Ginu’s task states.
Ginu keeps listening while it speaks, so cutting in is just talking. It stops, hears you, and answers the new thing.
A short “mm” or “yeah” is read as backchannel and does not stop it. Words do.
Try sayingExplore conversation“Wait—check the audio test first.”
A delegated task is a separate process on its own git branch and worktree. Nothing it does reaches your working tree until you accept it.
Open its live terminal to watch the real process, and type into it to redirect the work mid-run.
Try sayingExplore delegation“Look into that failing audio test while I work on the design.”
Queued tasks are records on disk, not lines in a chat. They survive a restart, each with its result and whether its branch merged.
A high-priority agent task starts the moment it is filed. Nothing already running is paused for it.
Try sayingExplore overnight“Work through these tasks and tell me what needs my attention.”
Permission asks and irreversible actions collect behind the bell instead of scrolling past. They wait: nothing here times out.
An agent cannot mark its own news urgent. It reports facts; a policy you can read turns them into a priority.
Try sayingExplore your say“Show me the changes before publishing anything.”
What Ginu learns about your work is a tree of notes on your own disk, each one keeping the belief it replaced.
A guess is labelled assumed until you confirm it, and correcting one is a sentence: “that’s wrong, it’s pnpm”.
Try sayingExplore your memory“What did we decide about this project last time?”
Open a pairing window on your computer, scan it with any phone, and the same conversation and tasks appear in its browser. Your computer still does the work.
Try sayingExplore from your phone“How are the tasks going? Is anything waiting for me?”
Speech is recognised and spoken on your computer, so no audio is sent anywhere. Memory is a tree of notes in files you own, ready to read and correct.
The text does. Your chosen CLI sends the transcribed turn and the context it needs to its own provider — Claude Code to Anthropic, Codex to OpenAI, Gemini to Google. Ginu has no server in that path.
Bring a real project. The first conversation is worth more than the first demo.
It's a closed beta, and it's free.