A conversational AI plush toy runs entirely on a local LAN, using Whisper for speech‑to‑text, Hermes 3 Llama 3.1 8B (Q8_0) as the language model, and Kokoro for text‑to‑speech. It incorporates per‑person memory via LangGraph, face identification with dlib, and emotion recognition via FER+ to adapt responses to the speaker’s identity and mood. The plush acts as a dumb terminal built around a Pi Zero WH with microphone, camera, speaker, and head servo, streaming audio and video to the local AI pipeline.
Read original
reddit/r/LocalLLaMA