← ALL WORK
Writing2026

Raising an Alien in Simulation: Building a Robot's Brain Without the Robot

In July 2026 I built and tested Rocky's software brain in simulation, with no robot, no GPU and no local LLM in the loop, by putting every external dependency behind a Protocol with a Fake. Rocky has since been tested on the physical Reachy Mini and works; this corrected essay keeps the sim-first argument and records what changed.

Embodied AIConcurrencyLocal LLMRobotics

Update, 28 September 2026. Rocky has since been tested on the physical Reachy Mini and works. This is the July essay about building his brain in simulation before the robot was on my desk, with every detail that is wrong now, or was never true in the code, corrected. The current state of the project is on the project page.

I am building Rocky, the Eridian engineer from Andy Weir's Project Hail Mary, on a Reachy Mini Wireless robot. The framing is a curious alien child dropped on Earth: I teach him about the world, he remembers what he learns across sessions, and he answers in chords and a voice that is not quite human. This is a personal homage to Weir's novel and to Pollen Robotics' little robot, not an original character, and I want to name both sources before anything else.

The first build of that brain was written and tested entirely in simulation: no robot, no GPU, no local language model in the loop, and the test suite still runs that way. This essay is about how you build a robot's nervous system with none of the hardware present, and keep the codebase honest about which half has actually run. The GPU and the robot came later, and each surfaced problems the simulation could not.

The seam is the whole idea

The reason to work sim-first is boring and correct. A servo you can burn out, a 5.5 GB model you have to keep warm in VRAM, an audio backend that fights the operating system: none of that belongs in the inner loop where you are changing a persona rule for the fortieth time in an afternoon.

So every external service in rocky_mini sits behind a Protocol, and the sim and test paths use Fakes: the language model, speech to text, text to speech, voice activity detection, the robot's audio, and, added later, the camera, face detector, vision model and image embedder. A rule-based SimResponder plays Rocky well enough to drive the settings UI and a full conversation turn, and one config flag swaps in OllamaLLM behind the same interface. Real-client imports are lazy and guarded, so the sim never touches Ollama, the speech models or the Reachy SDK.

Diagram titled One Protocol seam, two worlds. The left column, built and tested in sim, holds the rule-based SimResponder, the asyncio turn loop, Fakes for the mic, speaker, servos and camera, JSONL memory with the settings UI, and a green suite of 241 pytest tests plus a separate Playwright browser pass. The right column is the real path, run on the PC and the robot. Four solid boxes have run: OllamaLLM with the Rocky LoRA v2, fine-tuned with Unsloth on the 4080 and served as rocky:latest after passing the ship gate; faster-whisper speech to text on the PC brain server; Chatterbox cloning the film audio for the live voice, with Piper as the code default; and the Reachy Mini's Raspberry Pi CM4 thin client handling the mic, VAD, direction of arrival, motion and audio mixing, with LLM, STT and TTS coming over the LAN. A dashed Vision box (face gaze, look_around on an Ollama VLM, DINOv2 object memory) is tested in sim with not yet run outside sim. A central panel lists the Protocol seams, and one config flag swaps the Fake for OllamaLLM.

The left column's first version, 164 tests, ran and passed before the robot arrived. It has kept growing since, and all 241 tests still run without the robot. The right column, written against the same Protocols, has since run: the language model, speech server and LoRA training on the RTX 4080 PC, and the voice path on the physical robot. The easy mistake in a sim-first build is letting a green suite imply that the robot works; for the whole build this essay first described, the robot had never moved.

Four latency domains, one process

A talking robot is a real-time problem wearing a toy costume. If Rocky takes a beat too long to acknowledge you, the illusion breaks; if two threads fight over a joint or the speaker, you get a jerk or a click. Each failure lives in a different time budget, so each gets one owner. On the robot, one process runs the asyncio conversation loop on its own thread plus worker threads for microphone input, the mixer and motion; a 5 Hz VisionThread, added later, starts with them unless vision is switched off; the real camera and face detector are opt-in, so out of the box the thread runs on Fakes.

Diagram titled One voice turn across four latency domains, one process. Four lanes run concurrently. AudioIn takes the mic through half-duplex voice activity detection to a speech event. The ConversationLoop, an asyncio loop, runs speech to text first, then fires the ack chord and thinking pose, then streams the LLM (SimResponder in sim, OllamaLLM on the real path) into a sentence chunker that pulls out tags and lints tics, then per-chunk TTS. Tools, the naivety auditor, curiosity and JSONL memory share the same loop. The Mixer, the single push_audio_sample owner, takes chords on a ChordBus and generation-tagged speech on a VoiceBus, ring-modulates the voice at a 0.2 wet blend, then sums, soft-clips and plays the result. The 100 Hz MotionThread, the single set_target owner, blends breathing and emotes, adds face gaze or the direction of arrival, and position-clamps every frame. A 5 Hz VisionThread can also feed face gaze to motion.

Follow one turn. AudioIn runs the microphone through voice activity detection with a 0.6-second hangover and emits a speech event, and the antennas perk up while you are still talking. The ConversationLoop sends the audio to speech to text, and only when the transcript returns does it fire the acknowledgement chord and thinking pose. Then the model streams tokens, a sentence chunker releases each finished sentence, pulls out emote and sound-effect tags and lints Rocky's verbal tics, and each sentence goes to text to speech at once, so one can play while the next is generated.

The Mixer is the only owner of push_audio_sample. Chords arrive on a ChordBus and speech on a generation-tagged VoiceBus; the Mixer ring-modulates the voice at a light 20 percent wet blend, sums it with the chords, soft-clips, and pushes fixed-size frames. The MotionThread runs at 100 Hz as the only owner of set_target. It layers breathing and emotes, then turns the head toward a fresh face if vision has one and toward the direction of arrival otherwise. Every frame is clamped on position only: pitch and roll within 40 degrees, head-versus-body yaw within 65, antennas between 3 and 45. A speech-driven head wobble is unit-tested but not yet fed by the running app.

Those two ownership rules carry the design: every other thread asks an owner, and nothing else writes, so breathing, a head turn and a chord compose into one performance instead of three threads fighting over the neck. The same rule is why Rocky's gaze is computed client-side and handed to motion through a setter; the official app's daemon-side face tracking would add a second writer on the head.

The generation tag applies the same rule to time. Barge-in bumps the generation, the Mixer drops anything queued under the old number, and on the robot the device buffer is flushed. The gap: nothing cancels the turn still generating, and its later sentences carry the new number, so the rest of the reply can still play. Voice barge-in is off by default anyway. The original reason for half-duplex, no echo cancellation on a PC, was wrong: the Wireless has hardware echo cancellation, and the SDK adds software cancellation on PCs. It stays half-duplex because nobody has measured how well the robot cancels Rocky's own voice, and motor noise is not echo.

A local, zero-cost stack

There are no API keys anywhere in this project. The brain is Qwen2.5-7B-Instruct, served by Ollama on my PC with an RTX 4080 and held warm in VRAM. The fine-tuned build, rocky:latest, is that base with the Rocky LoRA merged in and quantized to q5_K_M, chosen in config; the code default is still the stock model. Speech to text is faster-whisper on the same PC, which acts as a LAN brain server. The Reachy Mini, a Raspberry Pi CM4 with 4 GB of memory, is a thin client for the microphone, voice activity detection, motion, audio mixing, and the camera when vision is on.

The voice took longest to settle. The original plan was Piper, ring-modulated into Rocky's register. Piper sounded robotic, Piper over a chord layer more so, and an XTTS v2 clone dated and stilted. The voice that stuck is a Chatterbox clone of a clean 12-second cut of the film's Rocky audio (James Ortiz's performance), from the scrubbed clip in pedramamini's rocky_say gist. An automatically picked reference drifted into the wrong accent until a hand-checked cut and stronger adherence fixed it. Piper stays the code default and fallback. The price is about 5.5 seconds per spoken sentence on the shared GPU, so first audio trails the acknowledgement chord by about 5 seconds against a 2.5-second budget. That trade was deliberate, and it is the first thing to revisit.

I want to be precise about the LoRA, the easiest part of this stack to overstate. At first publication there was no trained LoRA, only a pipeline: a seed dataset, an Unsloth QLoRA scaffold for WSL2, and a ship gate. The gate promotes a fine-tune only if it matches or beats the stock model on naivety leaks and tool-call validity (it also scores a question-particle rate, which cannot separate models; see below), clears a fixed floor on each, does not continue held-out book text, and keeps general capability within 0.15 of stock.

Version 1 trained on the local 4080 on July 20 and passed. Two days later the persona moved to the film's delivery, v1 overfit catchphrases, and the stock model became the recommended robot setting. Live conversation also felt clipped, and the cause was mine: the gold replies averaged 13 words, and the persona asked for one main idea per turn. Version 2 retrained on a new rule (react, add a thought, hand the turn back) with 369 rows, 16 percent neutral replay. On July 23 it passed with 0 of 20 naivety leaks, particle and tool-validity rates of 1.0, no memorized text, and capability of 1.0 against stock's 0.875, and it now serves as rocky:latest. Three caveats: the particle rate is scored after the tic linter has already added the particle, so every model gets 1.0 and that check cannot tell them apart; a tool-validity score of 1.0 would also come from a run with no tool calls, and the count was not recorded; and the same stock baseline scored 1 and 16 leaks out of 20 in different Ollama sessions, so the gate only compares models scored in one run.

Character and learning, kept in their lane

Rocky is supposed to feel like a curious child, and most of that is small systems rather than one big prompt: verbal tics, chord stingers, a sleep-watch ritual, an antenna perk while you talk, and rare idle glances. At first publication he had no face tracking, by design, since the novel's Rocky has no eyes. He has since gained a camera, framed in character as an instrument he built: a vision thread steers his gaze toward a face, and a look_around tool asks a local vision model on the PC what the camera sees. All of it is tested in sim with Fakes; the camera is opt-in, with no robot run on record, not even the check that the gaze turns the right way.

The part I find more interesting is an auditor that enforces what Rocky is allowed to know. An alien child a week on Earth should not explain compound interest or quote a pop song. What runs today is narrower: after every reply, off the speaking path, a cheap check scans it for the answers to the 20 red-team probes and, if one slipped out without a curious framing, adds a correction to the next turn. It is a tripwire on known topics, not a detector of Earth knowledge in general, so compound interest or a pop lyric would pass it; an LLM judge for subtler leaks is written but not switched on. Separately, a 20-probe red-team suite tracks naivety as a thresholded regression metric, because a zero-leak gate would block every change while a tracked metric shows the persona drifting.

Memory is an append-only JSONL store with a confidence lifecycle: a fact starts as heard once and is promoted to confirmed, then mastered, when it is heard again or confirmed in the settings UI. The prompt digest ranks mastered facts first and flags stale heard-once facts as fuzzy. The store lives outside the app, survives a reinstall, and exports to a zip, which is how a Rocky raised in sim moves onto the robot. In sim it also remembers objects, as frozen DINOv2 embeddings of a camera frame matched by cosine similarity. On the real brain, teaching an object is one of the model-initiated tools that look unreachable (below), and the match thresholds are provisional.

What actually runs, and what only compiles

As of September, the sim runs end to end. The settings UI is a running artifact: a sim chat driving a real conversation turn, a fact table to confirm and delete against, emote buttons, live metrics with a latency meter against a 2.5-second budget, a model toggle, and memory export. The pytest suite is green at 241 tests (this post originally quoted 164), spanning unit tests, the FastAPI surface and a sim end-to-end pass, and it still passes with no robot SDK, no Ollama and no GPU. A separate Playwright script drives a real browser through the UI.

When this essay was first written, the whole right column was seamed but unexercised. Since then the loop fails in character: if speech to text, the model or speech synthesis fails, Rocky plays a canned clip saying so and the loop survives, which is tested. The code also shows a gap no test catches: the running turn never sends tool definitions to Ollama, so on the real brain the model-initiated tools (remembering a fact, looking around, learning an object) look unreachable, though the sim exercises them and the ship gate scores them. That comes from reading the code, not from a robot session. The same reading shows the vision thread running by default against the Fake camera, whose detector reports a face dead ahead; that outranks the direction of arrival, so until vision is disabled or a real camera is set, Rocky would stop turning toward a voice. That too is a code reading and a sim check, not a robot observation. I would rather write gaps like these plainly than round a "seamed" up to a "working."

What this was worth

None of the sim work proved Rocky would work on a robot. It showed that you can build and test an embodied character's whole software nervous system without the embodiment, if every external dependency sits behind a Protocol and you accept Fakes in the inner loop. At first publication the claim was bounded exactly: green in simulation, not yet moving a robot. Bring-up came next, where the clamps meet real joints and the half-duplex assumption meets a real room.

Postscript: the robot arrived

That test has happened. The Reachy Mini Wireless is assembled on my desk, and Rocky runs on it.

First hardware bring-up: Rocky on the physical Reachy Mini Wireless.

The sim-first bet paid out, less cleanly than the seam diagram promised. Three bugs kept Rocky deaf on the physical robot, each fix checked live:

  • A transient USB error on the ReSpeaker direction-of-arrival read killed the microphone thread and logged nothing, so one flaky read made Rocky permanently deaf. The read now degrades to "no reading", and each mic-thread step is guarded.
  • The vendored Silero voice detector accepted 512-sample frames without complaint and returned near-zero speech probability. On the robot, 63 percent of frames of clean synthesized speech crossed 0.6 at 256 samples, against 1 percent at 512. A Fake VAD could not have caught it.
  • The seam cut both ways. The LLM swap is one flag in the app builder, but the hardware entry point wires audio separately, and it swapped in the real microphone while keeping the sim's Fake speech synthesis and no speech to text. Every voice turn raised "no STT configured", the asyncio future swallowed the exception, and Rocky went quiet. The fix wires the remote speech clients and logs every failed turn.

Still open: the soak and failure drills; the 100 Hz motion loop measured on the CM4; the echo-cancellation test; gaze direction and object-match thresholds on the robot; latency measured on the robot, and a faster cloned voice; the barge-in and tool gaps above; and a hardware path that looks for its canned fallback clips in a folder a clean checkout does not have. I will write those up when they have happened.

The code is at github.com/TirtheshJani/rocky_mini; its decisions.md is a more current record than the README.

© 2026 TIRTHESHJANI.COMBARRIE, ONTARIOAD ASTRA PER ASPERA