Richael fine‑tuned the Qwen3‑14B language model using LoRA for the Soulor AI companion app, spending about $1.30 and completing training in roughly 30 minutes. The initial adapter failed to produce the required JSON envelope that the frontend expects, despite learning the persona correctly; rewriting the training corpus so every example matched that envelope fixed the issue, and the retrained adapter achieved perfect scores of 15/15 on both companion chat and feed‑comment evaluations at production prompt size. Latency analysis revealed that the original 7.9 second per‑token response was limited by GPU memory bandwidth rather than vLLM flags; migrating to an H100 reduced latency to 2.0 seconds, while disabling the thinking mode on the hosted Qwen fallback cut its latency from 14.6 seconds (with truncated output) to 5.7 seconds with complete replies. To keep costs low, the Modal endpoint scales to zero, but a 108‑second cold start would be unacceptable; a gateway probe checks whether a container is warm before each request, routing to the fine‑tuned model only when warm and using the probe itself as a warm‑up, so the first message of a conversation is served by a hosted model and subsequent turns by the fine‑tuned model. Two recurring bug patterns—unsaved writes due to missing awaits and hanging network calls without timeouts—were identified as the root cause of several apparent outages and were resolved by awaiting all writes and adding explicit timeouts to every provider call. The same inference engine powers both the companion and Simulation Mode, which presents five AI analysts (Optimist, Cynic, Mentor, Status Observer, Gossiper) that share the conversational context, allowing users to rehearse difficult conversations and seamlessly return to the main chat.

Read original