# Building a personal AI assistant

_Stiven Morelo · Sep 12, 2026 · 6 min read_

For a year I kept a notes app, a reminders app and a chat window open all day.{' '}None of them talked to each other. So I started building one assistant that could hold all three, running on my own machine.

**In short**

- Every assistant I tried wanted my data in its cloud.
- I wanted a small model on my laptop and a phone that syncs when it's home.
- If I can't unplug the router and still use it, it isn't mine.
- The backend does the work and streams SSE. The app only listens.
- Qwen 3.7 Flash tied on accuracy, invented half as many parameters and cost 4.4× less.
- The same model doesn't win every task.
- Typing was the bottleneck.
- Holding a button to talk made me use it five times more.
- The mic breathes while it listens, and that small detail made it feel alive.
- Local models cover 80% of daily tasks.
- Latency matters more than intelligence.
- Virso started on a HomeLab, then moved to Cloudflare Workers and OpenRouter.

## The problem

Every assistant I tried wanted my data in its cloud. I wanted the opposite: a small model on my laptop, a local database, and a phone app that syncs when it's home. The first prototype was a Swift app talking to [Ollama](https://ollama.com) over localhost.

> If I can't unplug the router and still use it, it isn't mine.

_Image: First pass at the capture screen, before voice input._

## The setup

All the heavy work happens on the backend. It talks to the model, and I ran every test there. It streams the answer back with server-sent events (SSE). The app does very little: it opens the stream, reads each `data:` line and appends the token.

_server.ts_

```ts
export default {
  async fetch(req: Request) {
    const { messages } = await req.json()
    const upstream = await fetch("http://homelab:11434/api/chat", {
      method: "POST",
      body: JSON.stringify({ model: "granite4.1:8b", messages, stream: true }),
    })
    return new Response(toSSE(upstream.body!), {
      headers: { "Content-Type": "text/event-stream" },
    })
  },
}
```

_Assistant.swift_

```swift
var req = URLRequest(url: URL(string: "https://api.smoreb.me/chat")!)
req.httpMethod = "POST"
req.setValue("text/event-stream", forHTTPHeaderField: "Accept")
req.httpBody = try JSONEncoder().encode(ChatBody(messages: context))

for try await line in URLSession.shared.bytes(for: req).0.lines {
  guard line.hasPrefix("data: ") else { continue }
  reply.append(try decode(line.dropFirst(6)).token)
}
```

The stack behind the assistant

- [Swift](https://swift.org)
- [Ollama](https://ollama.com)
- [Granite](https://www.ibm.com/granite)
- [Qwen](https://qwenlm.github.io)
- [Gemma](https://ai.google.dev/gemma)

When OpenRouter retired Granite 4.1 8B, the model Virso was using, I ran a new round on September 16. Same production prompt with 36 tools, the same 53 real phrases in Spanish and English (slang, voice, emoji and a few injection attempts), six models through OpenRouter.

| Model | Intent | Invented | p50 | p95 | $ / 1K |
|---|---|---|---|---|---|
| Qwen 3.7 Flash | 94.3% | 22.6% | 2.1 s | 3.2 s | $0.082 |
| gpt-oss-120b | 94.3% | 47.2% | 4.1 s | 23.7 s | $0.396 |
| gpt-oss-20b | 92.5% | 41.5% | 12.6 s | 14.7 s | $0.254 |
| DeepSeek V4 Flash | 88.7% | 24.5% | 4.3 s | 10.2 s | $0.176 |
| Ling 3.0 Flash | 66.0% | 22.6% | 1.2 s | 2.4 s | $0.169 |
| Llama 3.3 70B | 58.5% | 17.0% | 2.2 s | 7.4 s | $1.757 |

53 cases × 1 pass, same pipeline, September 16, 2026. Invented = parameters the model made up.

In the full 208-case suite, Qwen 3.7 Flash tied gpt-oss-120b on intent (96.2%) but invented half as many parameters (19.2% vs 39.4%), with a p95 of 2.9 s against 11.2 s, at 4.4× less cost. In a chat, people feel the p95, not the median. The most expensive model, Llama 3.3 70B, was also the second worst on intent. Price didn't buy accuracy.

One caveat: Qwen won for classifying, extracting and the weekly summary, not for writing chat replies. That one went to DeepSeek V4 Flash. The same model doesn't win every task, so measure each one on its own. And none of them was deterministic: Qwen repeated the same answer in 92% of cases across passes.

## Talking to it

Typing turned out to be the bottleneck. Holding a button and speaking made me use it five times more. Here's the flow as it works today:

Video: https://media.smoreb.me/videos/virso-voice.mp4
_Hold to talk, release to send. Recorded on an iPhone 17 Pro simulator._

_Hover to try them. Both are live code, not recordings. (interactive, see the web version)_

The small detail that made it feel alive: the mic button breathes while it listens.

_Image: Mic idle → listening_

> ✦ Tip. Keep the system prompt under 300 tokens on small models. Anything longer and first-token latency doubles.

## What I learned

- Local models are good enough for 80% of daily tasks.
- Latency matters more than intelligence for an assistant you talk to.
- Voice input changes how often you reach for it.
- Offline-first forces you to design a better data model.

Some of these lessons shaped how [Virso](https://virso.dev) works. Virso started exactly like this: a HomeLab, Ollama and a local model. When I saw that more people might want it, the architecture moved to the cloud: Cloudflare Workers for the backend and Qwen 3.7 served through OpenRouter. The SSE contract stayed the same, so the app barely changed.

This was the first system prompt we started testing with on the HomeLab, in case you want to try it.

[assistant-prompt-v1.md](https://smoreb.me/downloads/assistant-prompt-v1.md) — The first prompt, before the cloud · Ollama config · 3 KB

---

smoreb.me · Stiven Morelo · https://smoreb.me/notes/building-a-personal-ai-assistant
