AI Inference Guide
Edge & Mobile Inference
Edge & Mobile Deployment
Deploy supported models on mobile, IoT, and edge platforms. Compatibility and performance depend on the exact model, runtime, accelerator, memory budget, operating system, and thermal limits.
Key Features
Published 1.4B-2.7B variants
Mobile-focused evaluation
Snapdragon demo
CLIP-based vision
Platform: Android/Jetson/desktop research demos
Key Features
Early access
Deployment tooling
Heterogeneous compute focus
Verify supported runtimes
Platform: Consult the current compatibility list
Key Features
Model quantization
Hardware acceleration
Cross-platform
Optimized kernels
Platform: Mobile/Embedded
Performance Considerations
Memory Usage
Quantization can reduce memory use, but the saving and quality loss depend on model, precision, runtime, and evaluation set
Battery Life
Local inference reduces network use but can increase device compute and energy draw; profile the complete session
Hardware Acceleration
Benchmark available NPUs, GPUs, and CPU paths; the fastest backend varies by model and device