# Local AI bounty write-up: RTX 3090 on Omarchy

@cheatyyyy · 2026-09-28 · PR: [0xSero/local-ai-registry#137](https://github.com/0xSero/local-ai-registry/pull/137)

## Setup

My card is an RTX 3090 (24 GB) in a desktop that dual-booted Windows and CachyOS. I installed Omarchy
4.0.4 as a third boot on the same day, with Local AI plugin 6.1.8 and registry main `c6e6f4c7`. I
already run local models every day (ornith-1.5-35b and bonsai-2-27b on llama.cpp). Claude Code, a cloud
agent, drove most of the install over SSH. The local model did the registry work alone.

## What I expected

I expected to set up llama.cpp on a local endpoint myself and install my coding agents by hand: pi,
opencode, and juna, my own pi profile. juna uses jev and a few harness tricks to cut input tokens, so
prefill is faster. I was pleasantly surprised when one click in the panel started downloading the
model recommended for my GPU. There were a few issues, and the panel didn't add the model to my own pi
config, so plain `pi` outside the panel didn't use it.

## What happened

The panel recommended Qwen3.8-27B (EXL3 3 bpw, SGLang, 200k context) for the 3090. The weights
downloaded at about 63 MB/s, 13 GB in 4 minutes. The first start took about 20 minutes, most of it the
password stall described below. Later starts took about 90 seconds.

I opened pi from the panel and gave it the registry prompt word for word. It found the existing card
file and a matching recipe, read Local AI's state to identify the running deployment, and ran
`lab.py try` against the live endpoint. All six gates passed. It regenerated a stale catalog so
`make check` passed, forked the repo and opened PR #137 with the evidence. It took 12 minutes from
prompt to PR, and it spent about 9 of them waiting on the lab's context and speed gates.

That first proof measured 83.3 tok/s, well below the recipe's earlier proofs of 135.0 and 138.8. pi's
own request had been queued on the GPU during the speed gate, and my card was capped at 280 W. I raised
the limit to 350 W and asked pi to redo the proof without sending the model anything while the lab ran.
The clean run passed all six gates at 110.0 tok/s decode and 620 tok/s prefill, and pi amended the PR.
I stepped in twice in total: an interrupt during the first lab run, and the request to redo the proof.

## The three worst moments

1. **The password popup.** The first start sat on "waiting for your password" for 12 minutes. The
   prompt appears only on the desktop, and nothing over SSH or in the status shows that it is waiting.
2. **The agent timed out on the model it was testing.** The context gate prefilled about 166k tokens
   in roughly 5 minutes with nothing cached. pi's own request queued behind it on the single GPU,
   SGLang ran a later short lab request first, and pi reported "Request timed out".
3. **The speed test held the GPU.** The lab measures only the first 30 seconds of a streamed story, but
   it keeps reading until the model stops. The answer ran to 17,010 tokens and held the only GPU slot
   for almost 4 minutes while pi waited.

## What almost made me quit

A fresh install needed manual fixes. Installing the NVIDIA container runtime failed with "target not
found" because a fresh Omarchy had never synced its package databases. After a reboot, `run` refused
with "nvidia:0 is in use" because of stale state, and it worked only after a `stop`.

## What the agent got wrong on its own

- It pushed to `0xSero/local-ai-registry` first and got a 403. It forked the repo afterwards without
  help.
- It set a repo-local git identity with my GitHub account's public email. It considered the noreply
  address and chose the public one without asking me.
- It polled the lab through the same model it was testing. That polling request is the one that timed
  out, and it skewed the first speed proof.
- After the rerun it spent several minutes of xhigh reasoning on why its first `cat` of the log had
  failed, although the lab result was already on disk.
- Its new proof makes the 208k MTP+vision recipe the 3090's first pick. It reported this in the PR and
  did not question it, so one owner run changes the default for every 3090 user.

It validated the existing recipe instead of adding a duplicate, and every number in the PR comes from
the run.

## Speed, and how it felt

- The final proof, at a 350 W limit, measured 110.0 tok/s decode over the first 30 s of a streamed
  answer and 620 tok/s prefill on a 166k-token prompt.
- The lab's speed gate run alone averaged 97.4 tok/s at 280 W (3 runs) and 104.9 tok/s at 350 W (5
  runs). The first proof's 83.3 tok/s was measured with pi's request queued on the GPU.
- A plain 800-token chat answer took 10.6 s, about 75 tok/s including the prompt. The panel showed up
  to 92 tok/s.
- pi used about 1.7M input tokens and 31k output tokens across the task and the rerun.

I use local models a lot, so I know their limits. The main one is slow prefill, and it makes a local
model feel bad to use. I built juna (an efficient pi profile that uses 'jev' and some toolcall pruning tricks to lower input tokens) to cut that down. The bigger problem is that you can't dump context
on a local agent and expect it to work things out, the way you can with closed models like Claude Opus
or GPT-6.

When a turn needs little prefill, the local model answers almost instantly, and that helps a lot. For
simple tasks I can keep an agent running for as long as the VRAM is free. That gives me instant access
to a model that handles medium-level tasks. Most people who use AI this way know by feel which tasks
work locally and which need a cloud agent like Claude Code or Codex.

## One change that would save the most time

The lab should close its speed stream after the 30-second window. The answer ran to 17,010 tokens and
held the only GPU slot for almost 4 minutes. On a card this size, the agent running the lab shares that
slot.
