A 12-year-old developer built KODA, an AI coding ecosystem optimized for low-resource devices, running on a $150 Poco C55 phone with 4GB RAM. The system comprises a 92KB single-file web app, a 22KB code editor extension with zero dependencies, and a Cloudflare Edge Worker harness routing between three AI models: OpenAI GPT-OSS-120B for general coding (SWE-bench Verified: 62.4%), GPT-OSS-20B for speed with minimal quality loss (SWE-bench High: 60.4%), and DeepSeek R1 Distill Qwen 3.8B for complex reasoning (AIME 2024: 86.0%). Benchmarks show KODA's custom harness achieves 52.2% on the Code Index at max effort, outperforming Claude Code (48.2%) and other agents, though Claude Code leads at medium effort (46.9% vs 46.3%). The extension delivers <500ms first-token latency via SSE streaming, includes a 4-layer regex prompt injection filter (~1µs defense), and enforces a 200ms constitutional AI safety budget. Context handling supports 6 tabs with a 20KB per-file cap. The multi-model architecture addresses distinct failure modes—general coding, latency sensitivity, and reasoning—demonstrating that architectural composition, not individual model performance, drives real-world efficacy.

Read original