Gander: Omni Interaction Agent

Continuous omni perception, real-time interaction, asynchronous agentic action

01 / CAPABILITY DEMOS

Gander’s continuous interaction capabilities.

Explore how Gander handles screen-grounded tasks, continuous perception, bidirectional interpretation, multi-party interaction, and everyday reasoning.

RECORDED DEMO

Original audio and on-screen captions are preserved.

02 / SYSTEM ARCHITECTURE

Perception, conversation, and action—connected.

Gander unifies omni perception, real-time interaction, and agentic capabilities. Streaming speech, vision, and text remain part of the same evolving context, so interaction can continue while longer tasks run in the background.

Gander overview
Figure 1 · Gander brings together continuous multimodal interaction and asynchronous agentic execution.

Cerebellum and Brain, working together.

Separate responsibilities, shared context. Three cooperating components connect immediate conversation with longer-horizon work.

Front Cerebellum

Perceives streaming audio and visual input, handles real-time dialogue, and decides when to listen, speak, interrupt, or delegate.

Agent Orchestration Runtime

Coordinates asynchronous tasks, tracks execution state, and routes progress and results back into the ongoing interaction.

Back Brain

Uses plug-and-play reasoning agents for complex reasoning, tool use, and longer-horizon task execution.

Cerebellum–Brain architecture
Figure 2 · The front cerebellum, orchestration runtime, and back brain share task progress and multimodal context.

One timeline for perception, decisions, and speech.

Inside the front cerebellum, a Thinker–Talker architecture separates semantic planning from speech generation. Incoming observations, interaction decisions, and model outputs are serialized along a shared timeline.

Streaming unit1 second
Sliding context128 chunks

These are model settings, not a measurement of end-to-end response latency.

Streaming Thinker–Talker
Figure 4 · The streaming Thinker–Talker architecture aligns perception, control, text, and speech generation.

Learning when to listen, respond, and delegate.

The front cerebellum starts from MiniCPM-o 4.5. An interaction corpus of approximately 2.7 million examples covers speech, audio-visual and agentic interaction, plus robustness and negative supervision.

  1. Interaction data

    Time-aligned examples include turn-taking, interruptions, backchannels, visual events, task lifecycles, and situations in which the model should remain silent.

  2. Stage 1 · Thinker training

    Train the language model and audio projection layer while freezing perception encoders and speech-generation modules. Learn streaming responses, interaction decisions, and tool actions.

  3. Stage 2 · Talker training

    Freeze the Thinker and optimize only the TTS projector and decoder to produce aligned streaming speech.

Training details: Section 4 of the technical report.

03 / BENCHMARKS

Measuring interaction, speech, and understanding.

The report evaluates interaction timing, spoken conversation, and omni understanding. The tables below retain the reported baselines, metric scales, and evaluation caveats.