Front Cerebellum
Perceives streaming audio and visual input, handles real-time dialogue, and decides when to listen, speak, interrupt, or delegate.
Continuous omni perception, real-time interaction, asynchronous agentic action
01 / CAPABILITY DEMOS
Explore how Gander handles screen-grounded tasks, continuous perception, bidirectional interpretation, multi-party interaction, and everyday reasoning.
Original audio and on-screen captions are preserved.
02 / SYSTEM ARCHITECTURE
Gander unifies omni perception, real-time interaction, and agentic capabilities. Streaming speech, vision, and text remain part of the same evolving context, so interaction can continue while longer tasks run in the background.

Separate responsibilities, shared context. Three cooperating components connect immediate conversation with longer-horizon work.
Perceives streaming audio and visual input, handles real-time dialogue, and decides when to listen, speak, interrupt, or delegate.
Coordinates asynchronous tasks, tracks execution state, and routes progress and results back into the ongoing interaction.
Uses plug-and-play reasoning agents for complex reasoning, tool use, and longer-horizon task execution.

Inside the front cerebellum, a Thinker–Talker architecture separates semantic planning from speech generation. Incoming observations, interaction decisions, and model outputs are serialized along a shared timeline.
These are model settings, not a measurement of end-to-end response latency.

The front cerebellum starts from MiniCPM-o 4.5. An interaction corpus of approximately 2.7 million examples covers speech, audio-visual and agentic interaction, plus robustness and negative supervision.
Time-aligned examples include turn-taking, interruptions, backchannels, visual events, task lifecycles, and situations in which the model should remain silent.
Train the language model and audio projection layer while freezing perception encoders and speech-generation modules. Learn streaming responses, interaction decisions, and tool actions.
Freeze the Thinker and optimize only the TTS projector and decoder to produce aligned streaming speech.
03 / BENCHMARKS
The report evaluates interaction timing, spoken conversation, and omni understanding. The tables below retain the reported baselines, metric scales, and evaluation caveats.