Spotlight: Spatius — Rendering AI avatars on-device, not in the cloud
Most "talking avatar" products work the same way under the hood: a server renders video frames of a face, encodes them, and streams that video to your device. It looks fine until the network gets shaky, the server gets busy, or the bill comes in. Spatius is built on a different bet: send audio, not video, and let the device do the rendering.
The core engine takes an audio stream (via LiveKit WebRTC or plain WebSockets) and outputs real-time, lip-synced 3D facial animation directly on the device, no external animation pipeline required. Because the payload is audio-sized rather than video-sized, they claim it runs at 1080p and 25fps on genuinely low-end chipsets (the site names G88, S565, 8189, RK3576) without needing a dedicated GPU. That's the whole wedge: kiosks, phones, and budget hardware that could never sustain a smooth 2D video stream can apparently run a 3D avatar instead.
The smart part isn't the rendering trick itself (edge inference is not new), it's how plainly they've turned it into a cost argument. Spatius puts its cloud-edge hybrid pricing at $0.007 per minute of avatar time next to an "industry average" of about $0.15 per minute for cloud streaming, and calls it 'not a loss leader, just math.' Whether that industry comparison is apples-to-apples is something you'd want to verify against your own use case, but the framing is honest about what they're actually selling: an infrastructure cost advantage, not a flashier avatar.
They also don't lock you into their stock faces. The docs mention support for custom-trained 3DGS (3D Gaussian Splatting) avatar models alongside free stock ones, so you can bring a proprietary avatar and still get the on-device rendering benefits. SDKs are listed for Web, iOS, and Android, and the front page promises a 10-minute quick start.
Who should try this: teams building voice-driven products that need a face on constrained networks or cheap hardware, things like language tutors, in-car assistants, retail kiosks, or IoT devices where you can't assume a fat, stable connection or a beefy GPU on the other end. If your current avatar vendor bills you per minute of streamed video and your users are on mobile data or spotty wifi, the audio-in/render-local model is worth a serious look, and the pricing gap they're advertising ($0.007 vs ~$0.15/min) is large enough to matter at any real scale.
Who should skip it: if you're after ultra-photorealistic, film-grade avatars for polished marketing video, cloud rendering with more compute headroom still probably wins on visual fidelity. Same if you have no appetite for SDK integration work, streaming protocols, and device-side testing across chipsets, this is developer infrastructure, not a drag-and-drop widget. And the pricing page itself is a bit of a jumble right now ($0.42/hour, an $8,900 figure, $0.007/min, $0.15/min all appear without much labeling), so budget time to actually get a straight quote before you commit.
Where this goes next, in my read: the interesting test isn't the demo, it's whether the 25fps-on-a-cheap-chipset claim holds up across the full range of devices they list, and whether the custom avatar pipeline is as painless as the stock one. If both hold, Spatius is positioned less as "another avatar API" and more as the plumbing layer under a bunch of voice-AI products that currently can't afford to put a face on their assistant. That's a narrower, more defensible spot than trying to out-render the big cloud avatar players.
Try Spatius: spatius.ai
See the launch: Spatius on welaunch.sh
