Aurora AI
Local models with llama.cpp, cloud providers, speech, search by meaning, and how privacy is kept.
The goal: an assistant that is private by default, useful where you already are, and never in your way. It is off until you switch it on, nothing is downloaded until you set up a feature, and every place it appears has its own switch.
Runtime: llama.cpp's llama-server, downloaded on demand
- Alternatives: Ollama, LocalAI, vLLM, a Python stack (transformers/PyTorch).
- Why: llama.cpp is the engine under most local AI tools. Its official release (build b11149, the Vulkan build) is a 30 MB download that runs on any GPU with Vulkan (Intel, AMD, NVIDIA) and falls back to the CPU. It serves an OpenAI-compatible API, so local and cloud providers share one client. Ollama wraps the same engine but its Linux bundle is over 1 GB (CUDA included), and PyTorch stacks are several GB. None of them is in Debian, so Aurora downloads the runtime only when you set it up, pinned by SHA-256 like everything in the AI catalog. A download only starts if the disk keeps enough room for the system afterwards (5% of the disk, at least 2 GB and at most 10 GB); a full disk mid-download removes the partial file. Models live in
~/.local/share/aurora/ai, outside the system snapshots, so a removed model frees its space at once. Settings โ AI shows the space used and removes each model; the shell warns when any disk gets that full. - On demand, not always on:
aurora/ai/server.pystarts the server on the first question, on127.0.0.1only, and a small reaper stops it after 10 idle minutes, so a model only uses memory while you use it.
Models: Qwen3 and Gemma 3 (GGUF, 4-bit)
- Choice: Qwen3 1.7B (1.1 GB, any computer), Qwen3 4B and Gemma 3 4B (2.5 GB, 8 GB of memory), Qwen3 8B (5 GB, 16 GB of memory), quantized to Q4_K_M.
- Why: they are the strongest openly licensed models at these sizes, they are multilingual (Qwen3 covers over 100 languages, Italian included), and they follow instructions well enough for commands and rewriting. Qwen3's "thinking" mode is turned off for speed (
enable_thinking: false), and any<think>text that slips through is filtered out of the stream. - Tested: on the build machine, Qwen3 1.7B answers a question in Italian in about 4 seconds, including loading the model, and turns "find files bigger than 1 GB" into
find ~/ -type f -size +1G.
Speech: faster-whisper and Piper, in a private venv
- Dictation: faster-whisper (Whisper small or base, CTranslate2, int8 on the CPU).
pw-recordcaptures the microphone, Whisper transcribes (detecting the language unless you fix one), andwtypetypes the text into the focused app through Wayland's virtual keyboard. Alternatives: whisper.cpp (no Linux release binaries), Vosk (lower accuracy), cloud speech (not private). - Read aloud: Piper, natural neural voices, one per language (19 voices in the catalog). Alternatives: espeak-ng (robotic), cloud TTS.
- Both are Python packages that aren't in Debian, so
pipinstalls pinned versions (faster-whisper==1.2.1,piper-tts==1.8.0) into~/.local/share/aurora/ai/venv, away from the system Python. A round trip on the build machine: Piper reads an Italian sentence, and Whisper writes it back nearly word for word.
Search by meaning: EmbeddingGemma + SQLite + numpy
- Why: EmbeddingGemma 300M is small (333 MB), multilingual and fast on a CPU. The indexer (
aurora-ai index, a user timer every hour, at idle priority) reads text, code, Markdown, PDF (pdftotext) and Word/LibreOffice documents in the folders you choose, and stores a vector per passage in SQLite. A query is one embedding and a matrix product. In a test, "quanto devo pagare di luce" finds the English electricity bill, and "where is my train seat" finds an Italian PDF ticket. - Alternatives: Tracker/LocalSearch full-text (not by meaning), vector databases (overkill for a desktop).
Cloud providers, keys and subscriptions
- Anthropic Claude (Messages API) and any OpenAI-compatible API, with presets for OpenAI, Google Gemini, Mistral, OpenRouter, Groq and a local Ollama. Keys are stored in the login keyring with libsecret, never in a file.
- Subscriptions: Claude and ChatGPT plans are only for the providers' own apps, so Aurora AI doesn't pretend to log in with them. Settings says clearly that API keys are billed per use, separately from any subscription. Dev Hub installs Claude Code and Codex, which do sign in with a plan.
Languages
- The system prompt tells the model to answer in the language of your latest message and to follow requests like "reply in Italian". It falls back to the system language, and you can fix a language in Settings. Dictation detects the spoken language, and any of the 19 reading voices can be downloaded.
Where it appears
- Spotlight (
?and, for longer questions, an "Ask Aurora" result); the Assistant (dock, top bar, <kbd>Super</kbd>+<kbd>Shift</kbd>+<kbd>Space</kbd>); Writing Tools on the selection (it pastes the result back withwl-copyand a simulated <kbd>Ctrl</kbd>+<kbd>V</kbd>); Files; screenshots (the OCR text goes to the Assistant, so any model can help, not only vision models); notification summaries;askandwhyin bash (aPROMPT_COMMANDhook remembers the last command and its exit status;askruns nothing without ay).
The Assistant as a layer surface
- Alternatives: an ordinary window kept on top by compositor rules (what it was).
- Why: a window takes the keyboard when it opens and can't place itself on Wayland. As a gtk4-layer-shell surface (top layer, anchored bottom-right, keyboard "on demand") it stays above windows on every workspace and only gets the keyboard when clicked. The dock writes how much of the screen's bottom it covers to a file in
$XDG_RUNTIME_DIR; the Assistant watches it and sits just above the dock, or at the bottom when there is none. Quick actions read the clipboard withwl-paste: GTK only sees the clipboard while its window has the keyboard. The model picker lists downloaded models, or what an OpenAI-compatible server (Ollama, LM Studio) answers at/models.