←Notes SMOREB
↓ .md
NOTE · AI

Building a personal AI assistant

Stiven Stiven Morelo·Sep 12, 2026 ·6 min read

For a year I kept a notes app, a reminders app and a chat window open all day. None of them talked to each other. So I started building one assistant that could hold all three, running on my own machine.

IN SHORT

01Every assistant I tried wanted my data in its cloud.THE PROBLEMJUMP TO SECTION 02I wanted a small model on my laptop and a phone that syncs when it's home.THE PROBLEMJUMP TO SECTION 03If I can't unplug the router and still use it, it isn't mine.THE PROBLEMJUMP TO SECTION 04The backend does the work and streams SSE. The app only listens.THE SETUPSEE THE CODE 05Qwen 3.7 Flash tied on accuracy, invented half as many parameters and cost 4.4× less.THE SETUPSEE THE TABLE 06The same model doesn't win every task.THE SETUPSEE THE TABLE 07Typing was the bottleneck.TALKING TO ITJUMP TO SECTION 08Holding a button to talk made me use it five times more.TALKING TO ITWATCH THE DEMO 09The mic breathes while it listens, and that small detail made it feel alive.TALKING TO ITJUMP TO SECTION 10Local models cover 80% of daily tasks.WHAT I LEARNEDJUMP TO SECTION 11Latency matters more than intelligence.WHAT I LEARNEDJUMP TO SECTION 12Virso started on a HomeLab, then moved to Cloudflare Workers and OpenRouter.VIRSO.DEVOPEN SITE ↗

The problem

Every assistant I tried wanted my data in its cloud. I wanted the opposite: a small model on my laptop, a local database, and a phone app that syncs when it’s home. The first prototype was a Swift app talking to Ollama ↗ over localhost.

If I can’t unplug the router and still use it, it isn’t mine.

First pass at the capture screen
First pass at the capture screen, before voice input.

The setup

All the heavy work happens on the backend. It talks to the model, and I ran every test there. It streams the answer back with server-sent events (SSE). The app does very little: it opens the stream, reads each data: line and appends the token.

export default {
  async fetch(req: Request) {
    const { messages } = await req.json()
    const upstream = await fetch("http://homelab:11434/api/chat", {
      method: "POST",
      body: JSON.stringify({ model: "granite4.1:8b", messages, stream: true }),
    })
    return new Response(toSSE(upstream.body!), {
      headers: { "Content-Type": "text/event-stream" },
    })
  },
}
The stack behind the assistant

When OpenRouter retired Granite 4.1 8B, the model Virso was using, I ran a new round on September 16. Same production prompt with 36 tools, the same 53 real phrases in Spanish and English (slang, voice, emoji and a few injection attempts), six models through OpenRouter.

ModelIntentInventedp95$ / 1K
Qwen 3.7 Flash94.3%22.6%3.2 s$0.082
gpt-oss-120b94.3%47.2%23.7 s$0.396
gpt-oss-20b92.5%41.5%14.7 s$0.254
DeepSeek V4 Flash88.7%24.5%10.2 s$0.176
Ling 3.0 Flash66.0%22.6%2.4 s$0.169
Llama 3.3 70B58.5%17.0%7.4 s$1.757
53 cases × 1 pass, same pipeline, September 16, 2026. Invented = parameters the model made up.

In the full 208-case suite, Qwen 3.7 Flash tied gpt-oss-120b on intent (96.2%) but invented half as many parameters (19.2% vs 39.4%), with a p95 of 2.9 s against 11.2 s, at 4.4× less cost. In a chat, people feel the p95, not the median. The most expensive model, Llama 3.3 70B, was also the second worst on intent. Price didn’t buy accuracy.

One caveat: Qwen won for classifying, extracting and the weekly summary, not for writing chat replies. That one went to DeepSeek V4 Flash. The same model doesn’t win every task, so measure each one on its own. And none of them was deterministic: Qwen repeated the same answer in 92% of cases across passes.

Talking to it

Typing turned out to be the bottleneck. Holding a button and speaking made me use it five times more. Here’s the flow as it works today:

0:00
Hold to talk, release to send. Recorded on an iPhone 17 Pro simulator.
Hover to try them. Both are live code, not recordings.

The small detail that made it feel alive: the mic button breathes while it listens.

Mic idle → listening
GIFMic idle → listening
Tip. Keep the system prompt under 300 tokens on small models. Anything longer and first-token latency doubles.

What I learned

Some of these lessons shaped how Virso works. Virso started exactly like this: a HomeLab, Ollama and a local model. When I saw that more people might want it, the architecture moved to the cloud: Cloudflare Workers for the backend and Qwen 3.7 served through OpenRouter. The SSE contract stayed the same, so the app barely changed.

This was the first system prompt we started testing with on the HomeLab, in case you want to try it.

.md assistant-prompt-v1.md The first prompt, before the cloud · Ollama config · 3 KB Download
Download as .md