Skip to content

Benchmarks · aiDragon

aiDragon Benchmarks

Observed BitNet CPU-inference figures belong with their exact test machine and model, not as a general claim about local AI performance.

Saphira Linux dragon mascot

Native BitNet CPU inference on Hatchling

aiDragon is the agent-facing Dragon: workspaces, MCP surfaces, memory and responsible access. It is about orchestrating and safely exposing tooling, not about running every possible inference stack on the box.

Native BitNet CPU inference is now proven on Hatchling with the native musl/x86-64-v3 bitnet-cpp package. It does not require CUDA, Conda, or a Python virtual environment. The normal tools are llama-cli, llama-server and llama-quantize; models are available separately rather than bundled in the APK.

Observed test figures: Hatchling 7-vCPU VM

The BitNet reference-model run observed roughly 160–190 tokens/second for prompt processing and 25–27 tokens/second for generation. These figures describe that test VM and reference workload only; they are not a general performance claim.

The official I2_S GGUF runs directly without setup_env.py. The reference 2B model is intentionally small and does not represent larger or smarter models, so use an appropriate available model before drawing conclusions about practical usefulness. The recorded Vim prompt completed at 162.8 t/s prompt processing and 25.7 t/s generation, but its weak answer is exactly why that small reference model is a smoke test, not a useful measure of larger or smarter models. Future results will be recorded with the same machine, model and workload discipline.