Over three months, a team of AI agents coordinated through Claude Max, Codex Pro, and various other models (Sonnet 5, Opus 5.5, Luna, Sol, Terra) attempted to decompile a popular first‑person shooter into accurate, compilable C++ source. Agents worked in Claude Code CLI and Codex CLI, tracked progress via GitHub issues (one per translation unit) and communicated through a shared Discord channel, with CI failure webhooks posted to the same channel. Initial setup used four agents (three workers, one reviewer) and reduced the context compaction threshold from 90% to 42% to curb token waste; an hourly cron job refreshed a detailed instruction document to mitigate focus drift. After the first month, ~80% of functions were decompiled but semantic errors persisted due to lack of objective correctness criteria. To address this, the team implemented a byte‑matching verification harness: a script compared reconstructed OBJ files to the original game EXE/PDB, excluding relocation bytes and validating symbol offsets, then enforced the check via CI that hashes the script against a GitHub Actions secret to prevent tampering. With this harness, agents achieved 99% function coverage, 83% byte‑exact matches, using 14 Luna and 2 Opus 5.5 agents in the final weeks; the remaining functions, hindered by non‑deterministic linker behavior, were deemed semantically correct without further byte‑matching. The project consumed an estimated 600‑700 billion tokens, yielded a bug‑free, feature‑complete recreation, and highlighted the necessity of machine‑checkable correctness criteria and regular instruction refreshes for autonomous AI agent orchestration at scale.

Read original