# Local AI bounty: friction log (RTX 3090 on Omarchy)

Tester: @cheatyyyy. Times are IST (UTC+5:30); UTC in brackets where the plugin log gives it.

## How this run was done

Claude Code (Opus, running on a Mac) drove most of the install over SSH, and I stepped in at the
keyboard for the installer screens, logins and password prompts. Rows say who acted. Steps 1–4 were
logged after the fact from the agent transcript and the plugin's own log
(`~/.local/state/omarchy/local-ai/log`). Rows from step 5 on are written live.

## Step 1: starting point

```
omarchy version: 4.0.4-1
kernel: 7.2.5-3-omarchy
GPU: NVIDIA GA102 [GeForce RTX 3090] (rev a1), 24 GB; driver 610.57.04, CUDA UMD 13.3
iGPU: Intel Raptor Lake-S UHD Graphics 770 (display is on the 3090)
RAM: 31 GiB, swap 62 GiB
Local AI plugin: 6.1.8 (96c8792); registry main c6e6f4c7
```

Prior local AI experience: heavy. I've used pretty much everything locally; my current daily local
models are ornith-1.5-35b and bonsai-2-27b (llama.cpp, mixing and matching).

Machine: my desktop, a Win11 + CachyOS dual boot before today. Omarchy went in as a third boot on
2026-09-28, on 131 GB freed by shrinking a Windows NTFS partition.

## Log

| Time (IST) | Step | Who | What happened | Impact / what fixed it |
|---|---|---|---|---|
| 14:12 | 0 | Claude Code | Read the plugin README: it needs Omarchy, Docker, and NVIDIA's container runtime; a 3090 recipe (Qwen3.8-27B, SGLang) exists. | — |
| 14:18–14:26 | 2 | Claude Code | Omarchy's ISO installer only offers whole-disk or free-space installs, so the agent shrank a Windows partition to make 131 GB of free space. The first shrink was refused by the agent's own permission check; a second attempt worked. | ~10 min. Not a plugin issue. |
| 14:39–14:47 | 2 | Claude Code | No USB stick: the agent put the ISO on disk and added a GRUB menu entry to boot it from the shared EFI partition. | Unusual path; worked first try. |
| 14:50–14:53 | 2 | me | Installer: picked the WD_BLACK NVMe, "Free space install" (131.1 GB), pressed Ctrl+C to switch to an unencrypted install so the box can be rebooted remotely. | I wasn't sure which drive to pick or how big the space should be, so I asked the agent. On that screen Ctrl+C switches to an unencrypted install; it doesn't cancel. |
| 14:55–15:00 | 2 | Claude Code | After install, Omarchy came up on a different LAN IP than the other boots, so the agent had to scan the LAN for the new SSH server. | ~5 min. Not a plugin issue. |
| ~15:01 | 2 | Claude Code | The MSI firmware drops UEFI boot entries that point at Omarchy's own EFI partition, so Omarchy wouldn't stay in the boot menu. Fixed by booting it through the shared EFI partition (Ubuntu shim → GRUB → chainload Limine). | Dual-boot specific; took a while to understand. |
| 15:06 | 3 | Claude Code | Ran the equivalent of `bin/omarchy-install-ai-local` as root over SSH (polkit policy + NVIDIA container runtime). `pacman -S nvidia-container-toolkit` failed: `warning: database file for 'omarchy' does not exist (use '-Sy' to download)` / `error: target not found: nvidia-container-toolkit`, then `nvidia-ctk: command not found`. | On a fresh Omarchy install the package databases had never been synced. Fixed with `pacman -Sy` first. **Suspected plugin bug: the installer should sync the databases or check that the package installed.** (The agent ran the script's commands by hand, not the script itself; confirm by rerunning the script on a fresh install.) |
| 15:07 | 3 | Claude Code | After the fix, a test container (`docker run --gpus all … nvidia-smi`) saw the RTX 3090 with 24 GB. | — |
| 15:08 | 4 | Claude Code | Panel's top pick for the 3090: Qwen3.8-27B, EXL3 3 bpw, SGLang, 13.8 GB, 200k context, tools and vision. | Clear. |
| 15:08–15:12 (09:38–09:42Z) | 4 | me | Started the model from the panel. Weights downloaded at ~63 MB/s, 13 GB in about 4 min, with progress in the panel. | Smooth. |
| 15:13–15:25 (09:43–09:55Z) | 4 | me/Claude Code | Status sat on "waiting for your password (the first start also downloads the engine)" for 12 minutes. I wasn't looking at the Omarchy screen, and from SSH there was no way to see or answer the prompt. Stop, then start again. | **12 min lost.** A start that needs a desktop password can't be driven remotely, and the state gives no hint that a popup is waiting on another screen. |
| ~15:10 | 4 | Claude Code | `docker ps` as my user showed no containers while the model was starting, because the plugin runs Docker through the password prompt. | Confusing when checking what's running. |
| 15:27–15:30 (09:57–10:00Z) | 4 | me | Entered the password on the desktop; the first start pulled the SGLang engine image (a few GB, ~2.5 min), then loading. | — |
| 15:31 (10:01:50Z) | 4 | — | Ready. Loading took ~1.5 min; VRAM 22.6 GB of 24. | — |
| 15:32 | 4 | Claude Code | First chat through the gateway (`127.0.0.1:12434/v1`, key from `gateway.key`): correct answer with reasoning; 800 tokens in 10.6 s ≈ **75 tok/s** including the prompt, card at ~278 W under my 280 W cap. The panel later showed 92 tok/s. | Success criterion for step 4 met. |
| 15:30 | 4 | Claude Code | While loading, fans ramped and the CPU hit 79 °C. Part of it was Omarchy's screensaver (`ttfx` in a `foot` terminal at 170% CPU) kicking in after ~2.5 min idle, on top of the model load. | Noise scare; not a plugin issue. |
| 16:23–16:25 | 4 | Claude Code | To stop the password prompts, added my user to the `docker` group (sudoless Docker; root-equivalent, my call). It only applies to new logins, so the panel kept asking for a password until the next reboot. | Worth a note in the README: the prompts go away with docker-group membership, after a re-login. |
| 16:24 (10:54Z) | 4 | Claude Code | After a reboot, `omarchy-local-ai run …` refused: `local-ai: nvidia:0 is in use`. The state file still said `ready` from before the reboot, though no container was running. | **Stale state after reboot.** `omarchy-local-ai stop <recipe>` first, then `run`, worked. Suspected plugin bug: state should be checked against the live containers. |
| 16:26 | 4 | Claude Code | `pi -p "Reply with exactly: local model ok"` from a plain shell answered "local model ok" once pi's default model was set to the local gateway in `~/.pi/agent/models.json` by hand. | The panel's "Open pi" wires this up; plain `pi` outside the panel doesn't. |
| 16:35 (11:05Z) | 5 | me | Opened pi from the panel on Qwen3.8-27B. | — |
| 20:05 | 4 | Claude Code | `omarchy-local-ai --help` and `omarchy-local-ai status` both print only the one-line usage (`snapshot \| run … \| stop … \| …`). Found the running recipe ID through `snapshot` JSON. | Minor: no `status` or `list` command. |
| 20:05–20:40 | — | Claude Code | Stopped the model with `omarchy-local-ai stop qwen38-27b-exl3-3bpw-rtx3090-sglang-tp1` for unrelated GPU benchmarks; restarted with `run … nvidia:0`. | Stop/run from the CLI is clean with sudoless Docker. |
| 20:44 | 6 prep | me | Ran `gh auth login` on Omarchy so the local agent can fork the registry and open the PR. `gh` was installed but not logged in on a fresh Omarchy. | The spec doesn't mention GitHub auth; a first-timer would hit it at the PR step. |
| 20:46 | 5 | Claude Code | Started pi in a tmux session (`bounty`) with the panel's own launch command (`PI_CODING_AGENT_DIR=…/local-ai/agents/pi pi --provider omarchy-local --model Qwen3.8-27B --thinking xhigh`, in `~/Work`) so it can be watched remotely. The panel itself only opens a desktop terminal. | Attaching from kitty over SSH failed with `missing or unsuitable terminal: xterm-kitty`; Omarchy has no kitty terminfo. `TERM=xterm-256color tmux attach -t bounty` works. |
| 20:46 | 5 | Claude Code | Asked: "which model are you, and what does nvidia-smi report?" pi ran two shell commands and answered "I'm Qwen3.8-27B … NVIDIA GeForce RTX 3090" (4.4k in / 262 out). | Correct. The xhigh reasoning prints in full above the answer, so the one-sentence reply appears twice. |
| 20:46:49 | 6 | Claude Code | Pasted the step 6 prompt word for word (GPU and port filled in). | — |
| 20:47 | 6 | local agent | Unprompted, pi queried the gateway for the served model (`turboderp-Qwen3.8-27B-exl3-3.00bpw`, SGLang, max_model_len 204800) and read `~/.local/state/omarchy/local-ai/deploy/` to find the running recipe. | Good start: it went looking for the running recipe before touching the registry. |
| 20:47–20:48 | 6 | local agent | Cloned the registry, found `registry/cards/nvidia/rtx-3090-24gb.json` and the existing recipe `registry/recipes/nvidia/rtx-3090-24gb/qwen3.8-27b.sglang.200k.json`, rendered it with `lab.py render`, and read the six gates in `lab.py` (the context gate fills ~174k of the 204,800 tokens). | Took item 2 (validate the running recipe) without help. |
| ~20:48 | 6 | local agent | Set a repo-local git identity on its own: `git config user.name "Safzan Pirani"`, `user.email "safzanpirani@gmail.com"` (the GitHub account's public email, which it looked up; it considered the noreply address and chose the public one), then created branch `rtx-3090-24gb-owner-run`. | Reasonable, but it made the privacy call itself without asking. |
| ~20:48 | 6 | local agent | Started `lab.py try` against the running endpoint with the gateway key from `gateway.key`. | Result pending. |
| ~20:49 | 6 | me | **Intervention:** interrupted pi mid-lab-run, resumed with "continue". The resume exposed a real problem: pi's request queued behind the lab's 174k-token context check (0 tokens from cache, about 5 min of prefill), the scheduler let the lab's later short requests jump ahead, and pi's client timed out ("Request timed out"). | See the SGLang rows below. |
| 20:49:28 | 6 | me | Resumed with the single word "continue". The interrupt had only aborted pi's `sleep 30 && cat lab-trial.log` poll; the `lab.py try` run itself kept going in the background (`--model qwen3.8-27b --engine sglang-qwen3.8-27b-exl3-3bpw-200k`, weights `turboderp/Qwen3.8-27B-exl3@6fe61ad6`). | — |
| 20:48:44–20:54 (15:18:44–15:24Z) | 6 | — | SGLang log: the lab's context gate prefilled ~166k tokens in about 5 min (~550 tok/s, `#cached-token: 0` on every chunk). pi's ~31k-token request joined the queue at 15:19:31Z; at 15:24:00Z a later 67-token lab request (next gate) ran before it. pi then printed "Error: Request timed out." and retried. | The agent that runs the lab shares the endpoint it is testing, so the long gates starve the agent until its client times out. Worth a note in the prompt or in `lab.py`: run the lab in the background and don't poll through the same model, or raise the agent's timeout for this task. |
| 20:54–20:56+ (15:24–15:26Z+) | 6 | — | The lab's speed gate ("Write a detailed 600-word story about a lighthouse keeper.") was still generating after 2.3 min (SGLang's token count for the request: 18.3k, ~130 tok/s). `lab.py` measures only the first 30 s but, by design ("the answer still runs", no max_tokens), keeps reading until the model stops. The server runs one request at a time, so pi's retry stayed queued the whole time. | Worst case the story runs to the 204,800-token limit, ~25 min. An unbounded speed prompt plus a single-slot server blocks the agent running the lab. Suggest closing the stream after the 30 s window. |
| 20:57:17 | 6 | local agent | The lab run ended by itself (the lab counted 17,010 streamed tokens in 231.5 s for the speed story): **all six gates passed**, decode 83.3 tok/s, prefill 584 tok/s. `lab.py` rewrote `registry/recipes/nvidia/rtx-3090-24gb/qwen3.8-27b.sglang.200k.json` (829 bytes). | No intervention needed on the long speed gate. |
| ~20:58 | 6 | local agent | pi got the GPU back, read `lab-trial.log`, ran `make` then `make check`. The recipe check passed, but the catalog was stale because the recipe changed; pi regenerated it and re-ran the check, which passed. | Handled the stale-catalog step without help. |
| ~20:58 | 6 | local agent | Tried to push the branch straight to `0xSero/local-ai-registry`: `fatal: unable to access … The requested URL returned error: 403`. | It went for the upstream remote first instead of forking. Recovered alone (next row). |
| 20:58:33 | 6 | local agent | Forked with `gh repo fork --clone=false`, added the fork as a remote, pushed, and opened [PR #137](https://github.com/0xSero/local-ai-registry/pull/137) with the exact lab command, a six-gate evidence table and the lab log. Diff: new proof in the existing 3090 recipe plus regenerated `dist/catalog.json` and `plugin/v2/recipes.json`. | Done without intervention. It noted in the PR that the new proof makes the 208k MTP+vision recipe the 3090's first pick in the catalog, so one owner run changes every 3090 user's default. CI hadn't reported on the fork PR yet. |
| 21:36–21:51 | 6 follow-up | Claude Code | Noticed the recipe on main already had proofs of 135.0 (owner) and 138.8 (vast) tok/s, against 83.3 from this run. Re-ran lab's own `stream_rate` speed gate with the model otherwise idle: **97.4 tok/s at 280 W** (98.7, 93.2, 100.2) and **104.9 tok/s at 350 W** (97.6, 115.9, 107.7, 103.7, 99.4). Locking the memory clock (`nvidia-smi -lmc 9876`) did nothing: CUDA stays in P2 at 9626 MHz. | The 83.3 proof was taken while pi's own request was queued on the same GPU. 350 W is ~8% faster than 280 W; both are still below the earlier proofs. |
| 21:52:34 | 6 follow-up | me | **Intervention:** set the power limit to 350 W and asked pi to drop the 83.3 proof, rerun the same `lab.py try` while waiting in one shell command with no model requests, then amend and force-push the PR and update its description. | Redoing the proof on a clean, uncapped run. |
| 21:53–22:01 | 6 follow-up | local agent | Restored the recipe from `origin/main`, started `lab.py try` in the background and waited in one `while kill -0 … sleep 30` loop with no model requests; SGLang showed 1 running, 0 queued during the speed gate. **All six gates passed: decode 110.0 tok/s, prefill 620 tok/s** (was 83.3 / 584). It briefly misread the log file's timestamp and re-checked before reporting. | The clean proof replaces the contaminated one. |
| 22:05 | 6 follow-up | local agent | Regenerated the catalog, `make check` passed, amended to one commit (`f18b575e`), force-pushed, and added a "Why this proof was redone" section to PR #137. Before that it spent a few minutes of xhigh reasoning on why its first `cat` of the log had failed, though the lab result was already on disk. | The new 110.0 proof still ranks below the 208k recipe's 129.0, so the PR still moves the 3090's first pick to the 208k recipe. |
