ENERZAi Runs a 1.58-bit Qwen Model on a Qualcomm IoT Chip
Korean startup ENERZAi says it successfully ran a 1.58-bit quantized version of the Qwen 3 1.7B model on a Qualcomm IoT chip's Hexagon NPU, using its proprietary Optimium inference engine. The company's pitch is straightforward: run the same model with fewer hardware resources, or deliver higher performance on identical hardware.
Edge AI — processing data with AI directly on devices such as smartphones, IoT sensors and vehicles — used to be limited to lightweight models for tasks like object recognition, voice recognition and noise suppression. As quantization techniques have matured, small language models are increasingly being pushed onto constrained hardware.
Extreme quantization, at 1.58 bits per weight, trades some accuracy for a large reduction in memory and compute. For sensor-class devices that must run on tiny power budgets, that trade-off can be the difference between running locally and not running at all.
ENERZAi said it has established partnerships with edge-AI chipmakers including Synaptics and Broadcom, and works with NXP, Qualcomm, MediaTek and Advantech. No independent benchmark of the claimed result was provided, so the demonstration should be treated as a vendor claim until third parties reproduce it.
Source: The Dong-a Ilbo. This article summarizes the linked reporting and distinguishes announced plans from demonstrated results.