Skip to content

AI

Saphira Linux aiDragon

A practical Saphira workspace for AI agents: your code, services, databases and tools stay under your control, while the model can run wherever it works best.

Native BitNet CPU inference proven on Hatchling
Saphira Linux aiDragon, the AI agent workspace mascot

AI where it works best

aiDragon is not an attempt to turn every Saphira virtual server into a local GPU AI workstation. Saphira now stages NVIDIA Open kernel modules on the pure-musl host, but CUDA and NVIDIA proprietary userspace do not belong there. The planned CUDA route is a separate glibc systemd-nspawn container, which has not yet been proven. aiDragon therefore uses AI where it already works brilliantly today: through agents, APIs and MCP.

GoodThe model does not have to live on your server for an agent to administer, develop, diagnose or maintain it. Your source code, tools, infrastructure and decisions can remain yours.

A useful place for an agent to work

A coding or operations agent in a shell is more than a chat window. Given useful, controlled access, it can inspect files and source code, run builds and tests, explain unfamiliar software, maintain documentation and help diagnose a live system. Saphira already provides the practical environment those agents need.

  • Python, Node.js and PHP for applications and automation.
  • Git and SSH for ordinary development and remote administration.
  • Databases, networking and normal Unix tools.
  • A filesystem and services the agent can inspect and work with.

You have a Saphira server. You install an agent. The agent can work with your code and services. You decide which tools and permissions it receives. That is the starting point for aiDragon.

Native BitNet on Hatchling

Native BitNet is now shipped in the Hatchling repository as bitnet-cpp. It is a native musl/x86-64-v3 build, and CPU inference is proven on Hatchling. It does not require CUDA, Conda, or a Python virtual environment.

# Install the native package from Hatchling.
sudo apk add bitnet-cpp

# Confirm the staged package revisions.
apk list | grep bitnet
# bitnet-cpp-20260830-r0 x86_64 {bitnet-cpp} (MIT)
# bitnet-cpp-20260830-r1 x86_64 {bitnet-cpp} (MIT)
# bitnet-cpp-20260830-r2 x86_64 {bitnet-cpp} (MIT)
# bitnet-cpp-20260830-r3 x86_64 {bitnet-cpp} (MIT) [installed]

# The normal native tools.
command -v llama-cli llama-server llama-quantize
  • llama-cli, llama-server and llama-quantize are the normal inference tools.
  • BitNet helper Python tooling is packaged system-wide and works with Saphira Python 3.14.
  • Models are available separately, but are not bundled in the APK. Keep the GGUF path explicit.
  • Official I2_S GGUF models run directly; setup_env.py is not required.
  • The proven reference-model run used ~/models/BitNet-b1.58-2B-4T on Hatchling.
# Proven Hatchling reference-model invocation.
llama-cli \
  -m ggml-model-i2_s.gguf \
  -t 4 -tb 4 \
  -n 128 \
  --temp 0.6 \
  --top-p 0.9 \
  --repeat-penalty 1.15 \
  -p "Explain clearly how to use Vim to edit, save and quit a file."

# Observed result
[ Prompt: 162.8 t/s | Generation: 25.7 t/s ]
GoodObserved on Hatchling's 7-vCPU test VM with the reference model: roughly 160–190 tokens/second for prompt processing and 25–27 tokens/second for generation. These are recorded test figures for that machine and workload, not general performance claims.
NoteThe reference 2B model is intentionally small and is not representative of larger or smarter models. Choose an appropriate available model before judging BitNet's usefulness for a real task.

What aiDragon is not

  • A current promise of GPU-accelerated local LLM inference on Saphira.
  • An AI SaaS dashboard or proprietary agent platform.
  • A way to lock you into one model vendor.

aiDragon is the Saphira environment and guidance that make existing AI agents useful on infrastructure you control. AI is a tool: learn to use it. Used thoughtfully, it can remove enormous amounts of repetitive work while you remain in control of the machine and its permissions.