rocky_mini: A Robot's Brain, Built Sim-First, Now Running on the Reachy Mini
An embodied AI character (Rocky, the Eridian engineer from Andy Weir's Project Hail Mary) on a Reachy Mini Wireless, running on a fully local, zero-cost stack. Every model, audio and camera service sits behind a Protocol with a Fake, so the brain was built with no robot, GPU, or local LLM, and its 241 tests still run that way. Since then: a Qwen2.5-7B LoRA trained and gated on a local RTX 4080, three voice fixes on the physical robot, and Rocky tested there and working.
rocky_mini is a personal homage project: Rocky, the Eridian engineer from Andy Weir's Project Hail Mary, role-played on a Reachy Mini Wireless robot from Pollen Robotics. He is a curious alien child on Earth, built to learn and remember across sessions (today that works in the sim; on the real brain the tool definitions are not yet sent to the model, see Still open), and he answers in chords and a cloned film voice with a light metallic edge. Rocky and the Eridians are Weir's creations, not mine, and the Reachy Mini is Pollen's hardware; I am naming both up front. The voice clones a short clip of James Ortiz's performance as Rocky in the film, taken from pedramamini's rocky_say gist.
The build ran in an unusual order. The first version of the software brain was written and tested entirely in simulation, with no robot, no GPU and no local language model in the loop. The real hardware came after: the RTX 4080 in my PC, which serves the model and trained Rocky's fine-tune, and then the robot, where bring-up found the places the simulation had been wrong. I have since tested Rocky on the physical Reachy Mini, and he works.
Why the brain came before the body
The Reachy Mini Wireless runs on a Raspberry Pi CM4 with 4 GB of memory, so a 7B language model does not run there in any useful sense. The PC becomes a LAN brain server for the language model, speech to text and speech synthesis; the robot becomes a thin client for voice activity detection, motion, audio mixing and, when enabled, the camera. There are no API keys anywhere. Even the robot's voice detector fits the board: a 2.3 MB Silero ONNX model on single-threaded onnxruntime, vendored because every published silero-vad wheel pulls in torch and gigabytes of CUDA wheels.
Sim-first began as a practical requirement: the whole software core had to be testable on a Windows PC with no robot, no model server and no GPU. What it bought was a ledger, with every piece of code on one side of a line, run or only written.
A Protocol with a Fake, not a mock
Every external dependency sits behind a Protocol: the language model, speech to text, speech synthesis, voice activity detection, the audio device, the camera, the face detector, the vision model and the image embedder. The Reachy SDK itself is import-guarded behind a pure tick() and one adapter, sdk_targets(). The sim and test paths use Fakes. A rule-based SimResponder plays Rocky well enough to drive a real conversation turn and the settings UI, and one config flag (ROCKY_LLM_BACKEND) swaps in OllamaLLM behind the same interface. Real clients are imported lazily, so the sim never touches Ollama, the speech models or the Reachy SDK.
A mock lives in a test file and knows what the test expects, so production code can keep an implicit dependency on a concrete client. With a Protocol, the interface is the production artifact, both implementations ship, and config picks one. The Fake gets exercised by the app, so when SimResponder breaks, it breaks in the settings UI while I am using it. Awkward interface questions also surface early. Streaming is the obvious one: if the LLM Protocol returned a string, the real path would be a rewrite, because the latency story depends on acting on text as it arrives.
class LLM(Protocol):
def stream(
self, messages: list[dict], tools: list[dict] | None = None
) -> AsyncIterator[StreamEvent]: # StreamEvent(delta: str, done: bool, tool_calls, prompt_eval_count, ...)
...
FakeLLM yields one word per event, the way a model streams tokens, so the sentence chunker, the tic linter and per-sentence speech synthesis were built against streaming from the first line of code.
Same Protocol seams, two worlds. Everything on the left runs in the 241-test sim suite. On the right, the solid boxes have run on the RTX 4080 PC or the physical Reachy Mini, and the dashed vision box has not yet run outside sim.
One flag swaps the language model. The rest of the swap lived somewhere else, and on the robot that split became one of the three reasons Rocky was deaf.
Four latency domains, one process
A late acknowledgement, a stuttering voice or two threads fighting over a joint each breaks the character, and each lives in a different time budget. On the robot, one process runs the asyncio conversation loop on its own thread beside motion, mixer and microphone threads, plus a 5 Hz vision thread whenever vision is enabled, which is the default on the robot path and runs against a Fake camera and detector until they are switched on. They communicate through setter calls on the motion manager, two audio queues in the Mixer, and a global generation counter.
MotionThread, 100 Hz
The single owner of set_target, one call per tick. A primary layer (breathing, keyframed emotes, and 13 recorded Pollen emotes sampled tick by tick instead of handed to the SDK's play_move) is smoothed by a 0.4 s blend, and the antennas perk to about 32 degrees while you talk. Additive offsets sit on top, including where the head points:
# MovementManager._secondary_deltas, trimmed; runs inside every 100 Hz tick
gaze_fresh = self._gaze is not None and (
self.speaking or (self.clock() - self._gaze_t) < GAZE_FRESH_S # 1.5 s
)
if gaze_fresh:
gaze_yaw, gaze_pitch = self._gaze # degrees, left by VisionThread via set_gaze()
deltas.append({"head_yaw": 0.6 * gaze_yaw, # head leads
"body_yaw": 0.2 * gaze_yaw, # body follows
"head_pitch": 0.5 * gaze_pitch})
elif self.doa_deg is not None: # degrees, left by AudioInThread via set_doa()
deltas.append({"head_yaw": 0.6 * self.doa_deg, "body_yaw": 0.2 * self.doa_deg})
tick() then returns clamp_pose(apply_deltas(self.current_primary, secondary)): pitch and roll within 40 degrees, head-versus-body yaw within 65, antennas between 3 and 45 and resting near 10, never zero. Nothing upstream can route around the clamp.
The vision and microphone threads never move the head; each leaves a number, and once per tick the manager decides which wins, so breathing, a face and a voice compose into one movement instead of a fight over the neck. One wrinkle: self.speaking is meant to hold the gaze while Rocky talks, but it depends on an audio-envelope feed that nothing in the running app calls yet. That branch never fires on the robot, and the speech head-bob that shares the feed stays at zero. Both are unit-tested. A unit test proves a function works; it says nothing about whether anything calls it.
Mixer, microphone and conversation loop
The Mixer is the single owner of push_audio_sample. Voice arrives on a generation-tagged bus, is resampled once to the robot's 16 kHz output, ring-modulated against a 140 Hz carrier and blended in at 20 percent, then summed with the chord bus, soft-clipped with tanh, and pushed in 320-sample frames. The microphone thread runs Silero voice detection with a 0.6 s hangover and polls the direction of arrival about ten times a second. The conversation loop is a state machine over idle, thinking, speaking and sleep-watch. A voice turn runs speech to text, fires the acknowledgement chord and thinking pose, then streams the model into a sentence chunker that pulls out emote and sound-effect tags and runs the tic linter. Each sentence goes to synthesis as soon as it is complete, so one plays while the next is generated.
The generation counter is ownership applied to time. Barge-in increments it, the Mixer drops anything queued under the old number, and on the robot clear_player() flushes audio already handed to the SDK. The generation bump keeps the Mixer's own queue single-writer. The SDK flush is not: clear_player() runs on whichever thread raises the interrupt, and a sentence the pump drained just before the bump can still be pushed after the flush. Half-duplex stays the default: the Wireless has hardware echo cancellation, but how well it removes Rocky's own voice is unmeasured, and motor noise is not echo, so the canceller has no reference to subtract it against.
What 241 tests cover, and what they cannot
The pytest suite is green at 241 tests and still needs no robot, GPU or model server; the latest run passed in a fresh Python 3.11 environment without the Reachy SDK installed. That is the best evidence the seam holds. The README's count is a smaller lesson: it said 126 while the suite ran 128, was corrected to 164, and still says 164. Trust the number pytest prints.
The suite follows the risk. DSP tests check under an FFT that ring modulation creates sidebands and suppresses the carrier, and that carrier phase stays continuous across chunk boundaries, because a phase jump is an audible click that ears rationalize and an FFT does not. The persona prompt gets a byte-stability test, because the prefix must be identical for Ollama to reuse its cache. The first live check of that cache said there was no reuse: it was counting prompt_eval_count, which on Ollama 0.32.x reports the whole prompt, cached or not. Timing the prefill instead showed turn two at 14 to 17 percent of turn one. The counter was the artifact.
Concurrency is tested at the seam, not by sleeping and hoping: the Mixer tests drive the generation counter and assert stale chunks drop, and the turn tests run the whole state machine against Fakes in milliseconds. The naivety suite, 20 red-team probes that tempt Rocky to state Earth facts he has not been taught, is a thresholded metric, because a zero-leak gate on a language model fails for reasons unrelated to the change and gets muted within a week. A separate Playwright script walks the settings UI in a browser. No test here proves a servo moves.
The brain: a fine-tune, and a gate that can say no
The brain is Ollama serving a Qwen2.5-7B-Instruct fine-tune, quantized to GGUF q5_K_M and held resident. The fine-tune is a QLoRA trained with Unsloth on the 4080 under WSL2. The training script's defaults are rank 16, alpha equal to rank, and two epochs over every attention and MLP projection of a 4-bit base; two because style is cheap to teach and overtraining a small set erodes general ability. The data stays local: authored replies, off-novel synthetic dialogue and neutral replay against forgetting, with the novel used only as short paraphrases for voice calibration and a held-out set of verbatim snippets never trained on.
Version 1 was trained, served and still wrong. It had learned the old, telegraphic persona and overfit Rocky's catchphrases, so when the persona moved toward the film's delivery, the stock model with a good prompt became the recommended robot setting. Live conversation also felt clipped, and the cause was the training mix: gold replies averaged 13 words, and the persona ordered "one main idea per turn." Version 2 changed the reply shape (react, add one thought, hand the turn back), added 18 livelier exemplars, grew neutral replay to 60 of 369 rows, and trained on the new persona bytes, so the prefix it learned matches the one the app sends.
Merging hit a memory wall: a 15 GB fp16 base and 7 GB of RAM in WSL. The way out is in the LoRA math. Each update is stored as two thin matrices, so the merged weight W + (B @ A) * alpha / r can be computed one tensor at a time, shard by shard:
def delta_for(base_key: str) -> torch.Tensor | None:
# base_key: "model.layers.0.self_attn.q_proj.weight"; W shape: (d_out, d_in)
if not base_key.endswith(".weight"):
return None
stem = base_key[: -len(".weight")]
a = lora.get(f"base_model.model.{stem}.lora_A.weight") # shape: (r, d_in)
b = lora.get(f"base_model.model.{stem}.lora_B.weight") # shape: (d_out, r)
if a is None or b is None:
return None # not an adapter target: copied unchanged
return (b.to(torch.float32) @ a.to(torch.float32)) * scale # (d_out, d_in); scale = alpha / r
Each shard is loaded, patched in fp32, cast back, written and freed. The script then checks the merged count against the adapter's targets and exits on a mismatch, so a key-naming slip cannot yield a silently half-merged model. Merging before quantizing was deliberate: Ollama's docs warn that an adapter applied over a differently quantized base gives erratic results.
The ship gate
A fine-tune ships only if it matches or beats the stock model on naivety leaks, question-particle rate and tool-call validity; clears floors on each (at most 2 leaks in 20, both rates at least 0.9); reproduces no run of five or more held-out book words; and keeps general capability within 0.15 of the baseline. Version 2 passed with 0 of 20 leaks, a particle rate of 1.0, tool validity of 1.0, a clean memorization probe, and capability 1.0 against the stock model's 0.875. It ships as rocky:latest, selected with one environment setting. One caveat on that tool score: validity defaults to 1.0 when the model makes no tool calls, and that run did not record the count.
Two details make the verdicts mean something. The capability probe runs under a neutral system message, because under Rocky's persona, refusing untaught Earth facts is correct, and the probe would otherwise score persona compliance. And determinism holds only within one server process: at temperature 0 with a fixed seed, the same stock baseline scored 1 of 20 leaks in one Ollama session and 16 of 20 in another. So the gate always scores candidate and baseline in the same invocation.
The voice
Piper sounded robotic, Piper with a chord layer more so, and an XTTS v2 clone dated. The voice that stuck is Chatterbox, Resemble AI's open cloning model, on a hand-checked 12-second cut of the film audio, picked by A/B across three reference cuts. On the way, an automatically picked reference with low adherence (cfg 0.35) drifted into the wrong accent, and loading Chatterbox after faster-whisper's ctranslate2 had bound cuDNN crashed the speech server, so Chatterbox now loads first. Piper stays the code default and the fallback.
The price is latency. The budget is a 2.5-second p50 from the user's last voiced frame, hangover included, and a sim test guards that plumbing with Fakes. The clone costs about 5.5 seconds per spoken sentence on the shared GPU, so first audio trails the acknowledgement chord by about 5 seconds. The decision log records the trade and its reversal: switch back to Piper for instant, robotic speech.
Bring-up: what the sim got wrong
Sim-first shaped bring-up. Everything a Fake could prove had been proven, so what remained were the places where a Fake encoded a wrong belief about the hardware, and those failed quietly.
Before the robot: auditing the Fakes
Before first boot, every SDK touch point was checked against the installed reachy-mini 1.9.0 source and, where it mattered, run live against the SDK's mockup daemon on the PC. The audit found seven broken surfaces and a gap, each a wrong assumption about the SDK that nothing in the sim could contradict. run() had the wrong signature. set_target was handed Rocky's own Pose in degrees when the SDK wants a 4x4 head matrix and radians (AttributeError: 'Pose' object has no attribute 'shape'). Direction of arrival lived on mini.media, in radians, and a getattr(..., None) lookup would have returned nothing forever, with no error. Piper's 22050 Hz audio played at the robot's 16000 Hz would have come out 27 percent slower and lower (16000 / 22050 is about 0.73). And the microphone thread had never been written. A Fake written from documentation is a hypothesis about the hardware, and it passes every test because the tests share its assumptions.
After the fixes, a 60-second mockup run held the motion loop at a mean 98.79 Hz, with a p99 tick of 10.33 ms and one thread calling set_target. That is a PC number; the CM4 measurement is still to do.
On the robot: three reasons Rocky was deaf
On the physical robot, Rocky could not hear, for three separate reasons, each fixed and verified live.
The vendored Silero model was called straight through onnxruntime on 512-sample frames with no carried context. It ran without error and returned near-zero probability for real speech. Silero documents 512 samples at 16 kHz, but its own wrapper also prepends 64 samples of the previous chunk, which this direct call skipped. Measured on the robot against clean 16 kHz synthesized speech, 63 percent of frames cleared the 0.6 threshold at 256 samples, against 1 percent at 512, so 256 shipped as the measured fix. Restoring the context at 512 is the untested alternative.
The ReSpeaker direction read is a USB control transfer that can throw a transient USBError (errno 5) under bus contention. That exception killed the microphone thread, and nothing logged it. A failed read now means "no reading" for that tick, and every microphone step is guarded, so a glitch costs one tick instead of Rocky's ears.
The third came from the seam. The swap from sim to real had ended up in two places: the app builder, which picks the language model, and the hardware entry point, which wires the robot's audio. The hardware path swapped in the real microphone but kept the sim's Fake synthesis and no speech-to-text client, so every voice turn raised "no STT configured". Nothing showed it: the turn runs as a coroutine submitted from the microphone thread, and an exception inside a future nobody reads vanishes. The fix wires the remote speech clients on the hardware path and attaches a callback:
fut = asyncio.run_coroutine_threadsafe(
self.state.loop.handle_voice_turn( # seg.pcm: float32 mono at 16 kHz, shape (n_samples,)
seg.pcm, seg.sample_rate, t_start=seg.t_last_voiced
),
aio, # the ConversationLoop's event loop, running on its own thread
)
def _log_turn_result(f) -> None:
try:
res = f.result() # re-raises whatever the turn raised
logger.info("turn done: reply=%r", (res.reply or "")[:120])
except Exception:
logger.exception("voice turn failed")
fut.add_done_callback(_log_turn_result)
Then the first robot turn after a speech-server restart timed out (httpx.ReadTimeout) while faster-whisper loaded cold. The server now pre-warms every model at boot, and the client timeout went from 15 to 30 seconds.
Every one of these was silent: a model returning zeros, a thread dying without a log line, an exception swallowed by a future. The sim had proven the logic. Bring-up was mostly the work of making failure loud.
After those fixes I tested Rocky on the physical Reachy Mini Wireless, and he works. The repo's own on-robot records cover the voice path: the microphone, voice detection, direction reads, speech to text and synthesis over the LAN, and the real language model in live conversation.
Giving an eyeless alien a camera
In the novel Rocky has no eyes and echolocates, so the first design used the direction of a voice as his only spatial input. He later gained a camera, a deliberate move to use the whole robot. The obvious route was wrong: the official Reachy Mini conversation app tracks faces daemon-side with mini.start_head_tracking(), which would put a second writer on the head. Instead a 5 Hz vision thread finds a face, maps its place in the frame to a yaw and pitch offset, and leaves it on the motion manager with set_gaze(). The reconciliation shown earlier is the whole integration. In character, the camera is an instrument Rocky built.
Two tools ride on it. look_around describes the room through a local vision model on the PC's Ollama (Qwen2.5-VL 7B by default, pending a timing probe). teach_object is a tool for naming an object in view: a frozen DINOv2 ViT-S/14 on the speech server embeds the frame, since Ollama has no image-embedding API, and later recognition accepts a cosine match of at least 0.60 that beats the runner-up by 0.08. All of it is built and tested in sim, and none of it has a robot run on record: the camera and face detector default to Fakes, the physical sign of the gaze is a bring-up check, and the thresholds are provisional. Today only unit tests call teach_object; the sim responder never does, and the real brain is not yet sent tool definitions.
Still open
- Motion-loop timing and CPU budget on the CM4 itself.
- The rest of the bring-up checklist: speaker pitch, the direction-of-arrival sign, and long-session failure drills (brain down, Wi-Fi pulled, cold Ollama, stop mid-sentence, rapid barge-in).
- The on-robot echo-cancellation test that would let open-mic barge-in replace half-duplex.
- Latency measured on the robot over the LAN, and a faster path for the cloned voice.
- A recorded robot session on LoRA v2 itself; the live conversations that shaped it came before it.
- Vision on hardware, and recorded emotes expected to push the antennas into the clamp.
- Gaps found by reading the code, not by watching the robot. The runtime turn never sends tool definitions to Ollama, so model-initiated
remember_fact,look_aroundandteach_objectcalls look unreachable on the real brain, although the gate scores them. Barge-in drops queued speech but not the rest of an in-flight reply, which arrives under the new generation and plays. On a clean checkout, the hardware path looks for the in-character fallback clips one directory above the tracked copies, and the emote voice clips load before that path is set. With default settings the robot path also starts the vision thread against the Fake camera, whose Fake detector always reports a centered face; that fresh zero gaze outranks the direction of arrival, so Rocky would hold his gaze forward instead of turning toward a voice until vision is switched off or a real camera is set. That is a code reading, not a robot observation.
The sim-first build, with the full concurrency walk-through, is written up as a companion post: Raising an Alien in Simulation. The code is at github.com/TirtheshJani/rocky_mini.
What it taught me
A Fake is a hypothesis about the hardware. The Protocol seam let the whole suite run with no hardware, and at the time of the audit it let seven wrong beliefs about the SDK pass all 128 tests. Auditing the Fakes against the real SDK before first boot was the step that finally tested those beliefs.
Make failure loud before making it rare. The costly bugs on the robot were the silent ones, and the fixes were small: a guarded loop, a done-callback, a warning logged once.
Keep a ledger of what has run, separate from what has been written. The diagram's solid and dashed boxes, the decision log and the open list above are one habit, and it is why "tested on the robot and working" can sit beside "vision has no robot run on record" without either weakening the other.
And compare a model to its baseline inside one server process, or the comparison measures the server. A green result is a claim about exactly what was measured, and the useful work is checking what it leaves out.