SafeEra Inference

Local GPU stack. Requests stay on this machine.

Low-latency voice

Voxtral Realtime partials start Qwen3.8 as they settle. Qwen tokens stream into Voxtral TTS, and audio plays as PCM arrives. Omni stays a separate panel.
Ready
Context is on. The last N turns are sent to Qwen and kept in this browser. Thinking is off on this path. A new phrase stops playback.
Qwen accepts image and video. Nemotron accepts image, audio, and video. Omni accepts all three and can speak back.