Building a personal AI assistant
For a year I kept a notes app, a reminders app and a chat window open all day. None of them talked to each other. So I started building one assistant that could hold all three, running on my own machine.
01Every assistant I tried wanted my data in its cloud.THE PROBLEMJUMP TO SECTION 02I wanted a small model on my laptop and a phone that syncs when it's home.THE PROBLEMJUMP TO SECTION 03If I can't unplug the router and still use it, it isn't mine.THE PROBLEMJUMP TO SECTION 04The backend does the work and streams SSE. The app only listens.THE SETUPSEE THE CODE 05Qwen 3.7 Flash tied on accuracy, invented half as many parameters and cost 4.4× less.THE SETUPSEE THE TABLE 06The same model doesn't win every task.THE SETUPSEE THE TABLE 07Typing was the bottleneck.TALKING TO ITJUMP TO SECTION 08Holding a button to talk made me use it five times more.TALKING TO ITWATCH THE DEMO 09The mic breathes while it listens, and that small detail made it feel alive.TALKING TO ITJUMP TO SECTION 10Local models cover 80% of daily tasks.WHAT I LEARNEDJUMP TO SECTION 11Latency matters more than intelligence.WHAT I LEARNEDJUMP TO SECTION 12Virso started on a HomeLab, then moved to Cloudflare Workers and OpenRouter.VIRSO.DEVOPEN SITE ↗
The problem
Every assistant I tried wanted my data in its cloud. I wanted the opposite: a small model on my laptop, a local database, and a phone app that syncs when it’s home. The first prototype was a Swift app talking to Ollama ↗ over localhost.
If I can’t unplug the router and still use it, it isn’t mine.
The setup
All the heavy work happens on the backend. It talks to the model, and I ran every test there. It streams the answer back with server-sent events (SSE). The app does very little: it opens the stream, reads each data: line and appends the token.
export default {
async fetch(req: Request) {
const { messages } = await req.json()
const upstream = await fetch("http://homelab:11434/api/chat", {
method: "POST",
body: JSON.stringify({ model: "granite4.1:8b", messages, stream: true }),
})
return new Response(toSSE(upstream.body!), {
headers: { "Content-Type": "text/event-stream" },
})
},
}var req = URLRequest(url: URL(string: "https://api.smoreb.me/chat")!)
req.httpMethod = "POST"
req.setValue("text/event-stream", forHTTPHeaderField: "Accept")
req.httpBody = try JSONEncoder().encode(ChatBody(messages: context))
for try await line in URLSession.shared.bytes(for: req).0.lines {
guard line.hasPrefix("data: ") else { continue }
reply.append(try decode(line.dropFirst(6)).token)
}When OpenRouter retired Granite 4.1 8B, the model Virso was using, I ran a new round on September 16. Same production prompt with 36 tools, the same 53 real phrases in Spanish and English (slang, voice, emoji and a few injection attempts), six models through OpenRouter.
In the full 208-case suite, Qwen 3.7 Flash tied gpt-oss-120b on intent (96.2%) but invented half as many parameters (19.2% vs 39.4%), with a p95 of 2.9 s against 11.2 s, at 4.4× less cost. In a chat, people feel the p95, not the median. The most expensive model, Llama 3.3 70B, was also the second worst on intent. Price didn’t buy accuracy.
One caveat: Qwen won for classifying, extracting and the weekly summary, not for writing chat replies. That one went to DeepSeek V4 Flash. The same model doesn’t win every task, so measure each one on its own. And none of them was deterministic: Qwen repeated the same answer in 92% of cases across passes.
Talking to it
Typing turned out to be the bottleneck. Holding a button and speaking made me use it five times more. Here’s the flow as it works today:
EXPORT AS H.264
The small detail that made it feel alive: the mic button breathes while it listens.
What I learned
- Local models are good enough for 80% of daily tasks.
- Latency matters more than intelligence for an assistant you talk to.
- Voice input changes how often you reach for it.
- Offline-first forces you to design a better data model.
Some of these lessons shaped how Virso works. Virso started exactly like this: a HomeLab, Ollama and a local model. When I saw that more people might want it, the architecture moved to the cloud: Cloudflare Workers for the backend and Qwen 3.7 served through OpenRouter. The SSE contract stayed the same, so the app barely changed.
This was the first system prompt we started testing with on the HomeLab, in case you want to try it.