mager.co

Chicago

The Chicago edition / 30 things to get into

01 / On my workbench

mager-bench 1.1: make the model find the bug

Counterexample Lab replaces easy implementation tasks with compact regression tests. First calibration: Astra catches 7/8 faults; Sol catches 5–7/8 across three attempts each.

Built, broken, figured out.

All tech ↗
Building

Jev: a decision model

Testing TypeSafe's Jev through Vercel AI Gateway, putting it inside a decision-heavy skill, and building a matchup reader for my reputation-based sports picks app.

Building

Instinct: a capable agent in your DMs

What changed when I could text an agent from iMessage: everyday cleanup, questions worth following up on, and the trust that makes casual delegation possible.

Building

fx: a 6MB coding agent built to be embedded

Vercel Labs shipped fx — a ~6MB, Zig-written coding agent that cold-starts in 10µs, speaks ACP, mounts MCP servers, and prints JSON. It's not a hosted service; it's a runtime you embed. Here's what it is, where I'd put it in my always-on harness, and how a large logistics operator would use it.

Building

tmux: orchestrating agents with send-keys and capture-pane

tmux is a 2007 terminal multiplexer that turns out to be the most native orchestration layer AI agents have: one session per agent, driven by any principal — human or model — with send-keys and capture-pane. No SDK, no plugin, no vendor lock-in; Claude, Codex, and OpenCode all speak it out of the box.

02 / Field of view

Look around.

All seen ↗
03 / Away from the keyboard

Something good
on the stove.

The kitchen ↗
From the kitchen

Euglena Yogo Parfait

I picked up euglena powder on Ishigaki — a single-celled organism that's part plant, part animal, packed with vitamins, minerals, and omega-3s — and the package had a yogurt recipe on it. Here's my yogo parfait take, plus why euglena is amazing.

Cherry Tomato Pasta

Blistered cherry tomatoes and garlic collapse into a full pasta sauce in twenty minutes — no peeling, no seeding, finished with butter and pasta water.

Cacio e Pepe

Three real ingredients — pecorino, pepper, pasta water — plus a knob of butter for insurance, tossed into a glossy sauce that never breaks.

Japanese Butter Soy Spaghetti

A five-ingredient Japanese-style spaghetti — butter, tamari, and parmesan tossed with hot pasta and finished with green onion. The wafu pasta I kept eyeing in Tokyo, made at home in ten minutes.

Life, outside.

More ↗

Sumo: My First Grand Tournament in Tokyo

We missed the original ticket sale, got rescued by a tour, and spent an afternoon learning how much more fun sumo is when someone helps you understand what you're watching.

04 / Open files

Take the document.

Artifacts ↗
Artifact

Frontend audit and hardening

2026-10-01-frontend.md

The working report from auditing and hardening mager.co with Codex and Impeccable: findings, fixes, an Astro 7 migration, and the checks behind them.

Also in the notebook

I benchmarked myself

Running Muse Spark 1.3 through mager-bench and comparing its scores with everyday agent work.

Every update, newest first

Latest

Pub/Sub for prxps

A small event pipeline for pick settlements, with a Firestore outbox and duplicate-safe recommendation invalidation.

Frontend audit and hardening

2026-10-01-frontend.md

The working report from auditing and hardening mager.co with Codex and Impeccable: findings, fixes, an Astro 7 migration, and the checks behind them.

Diving into Impeccable

Going from occasional design commands to reproducing bugs in Impeccable, with Codex, Astra, and a mobile browser.

Setting up my dot

My first conversation with magerdot, a pink dot with headphones. After naming it, I mention the emails I sent myself as TODOs but never completed.

After OpenAI DevDay, I started setting up my dot, magerdot. First up: the emails I'd sent myself as TODOs and never completed.

mager-bench 1.1: make the model find the bug

Counterexample Lab replaces easy implementation tasks with compact regression tests. First calibration: Astra catches 7/8 faults; Sol catches 5–7/8 across three attempts each.

Jev: a decision model

Testing TypeSafe's Jev through Vercel AI Gateway, putting it inside a decision-heavy skill, and building a matchup reader for my reputation-based sports picks app.

Instinct: a capable agent in your DMs

What changed when I could text an agent from iMessage: everyday cleanup, questions worth following up on, and the trust that makes casual delegation possible.

I benchmarked myself

Running Muse Spark 1.3 through mager-bench and comparing its scores with everyday agent work.

fx: a 6MB coding agent built to be embedded

Vercel Labs shipped fx — a ~6MB, Zig-written coding agent that cold-starts in 10µs, speaks ACP, mounts MCP servers, and prints JSON. It's not a hosted service; it's a runtime you embed. Here's what it is, where I'd put it in my always-on harness, and how a large logistics operator would use it.

tmux: orchestrating agents with send-keys and capture-pane

tmux is a 2007 terminal multiplexer that turns out to be the most native orchestration layer AI agents have: one session per agent, driven by any principal — human or model — with send-keys and capture-pane. No SDK, no plugin, no vendor lock-in; Claude, Codex, and OpenCode all speak it out of the box.

OpenSpec: The agreement layer that keeps your agent honest

AI coding assistants confidently build the wrong thing when requirements live only in chat. OpenSpec adds a lightweight spec layer — explore, propose, build, archive — so you agree on what to build before any code is written, and the specs persist in your repo as history your agent can read back. Here's the mental model, a real change from a skills discovery portal, and where it earns its keep.

AI Gateway: the end of the single-provider AI subscription

I'm moving off single-provider AI subscriptions toward a stack of parts — Eve for agents, Vercel AI Gateway as the primary model access and billing layer with no-markup provider pricing, and OpenCode Go kept as the fallback — and the enterprise version of that stack is the real product.

OpenCode CLI: ten commands worth knowing

OpenCode's CLI is bigger than 'type opencode and start a session.' Headless runs, provider auth, model discovery, MCP wiring, session archaeology, cost stats, and upgrades — the ten commands that carry daily work, with the doc gaps called out where they bite.

OpenCode Go + Buzz: killing Claude Code for a $10 harness

Second harness migration in two months. The always-on agent on my Mac mini now runs OpenCode on $10/mo open models instead of Claude Code, reachable from my phone over my own Buzz relay instead of Telegram. The interesting part: the swap was one line, because the protocol — not the model — is the actual seam.

Buzz: what it looks like when agents get equal standing

Block's open-source Nostr workspace puts people and agents on the same cryptographic footing — and lands at the end of a long chain of thinking about where always-on agent infrastructure should actually live.

Meeting Boris

Uber HQ, San Francisco, CA

Mager and Boris Cherny smiling together for a selfie at Uber HQ in San Francisco

Achievement unlocked: meeting Boris Cherny at Uber HQ.

mager-bench

New mager-bench challenges cover testing, debugging, async Python, and SQL.

Claude: from Skills to Agents to Subagents

A Skill is packaged know-how. An Agent is that know-how put to work autonomously. Subagents are where the work scales past what any single context can hold.

mager-bench: a personal coding model benchmark

Instead of reading someone else's leaderboard, build a small set of tasks you actually care about and run them yourself every time a new model drops — Simon Willison's SVG pelican test, but for code.

Warp: The Cloud Factory, Now Running

In March I wrote the theory. Zach from Warp shipped the implementation. Here's how a working cloud factory maps to the architecture I laid out.

skill-evals

A Claude Code plugin for investigating agent failures and building evaluations you can trust.

Claude Voice: an AI agent that talks back

A small Python voice agent that remembers the thread, streams Claude's reply to the terminal, and speaks it aloud through ElevenLabs — no ffmpeg, just afplay.

Eve: define your agent, deploy it, use it from anywhere

Define your agent in a directory, deploy it to Vercel's cloud with one command, and access it from anywhere. Months in, Eve has grown a platform around that model — capability registry, sandbox, subagents, agent-to-agent calls, MCP, evals — and my agent is still live, driven remotely from the eve TUI.

Loooom v1.0

Loooom v1.0 adds behavioral tests to its skill evaluations. A few early failures exposed assumptions in the tests themselves.

Kotsu: Designing a Logo by Inventing a Kanji

How I worked with Claude through five rounds of image generation to design a logo for my Japanese learning app — and ended up inventing a kanji that hides a smile.

Pride Postcard

Lady Gregory's, Andersonville, Chicago, IL

A table card on a bar reading 'Andersonville Pride' in rainbow letters, surrounded by a word cloud — acceptance, love, community, dignity, equality, celebrate

Table card at Lady Gregory's, in the heart of Andersonville. A whole vocabulary of the month, set in rainbow.

An OpenClaw setup for Dad

A plain-English walkthrough for setting up your own always-on AI assistant on a Mac mini — OpenClaw, Google Gemini, and Tailscale — written for a first-timer.

Fenway Park

Fenway, Boston, MA

Fenway Park during a Red Sox game

Sox game with Dad and Matt.

Killing OpenClaw for a native Claude Code setup

I love OpenClaw. I hate that it doesn't run on my Claude Pro subscription. Turns out Claude Code, with the Telegram channels plugin and one CLAUDE.md, is the same harness — minus the daemon, the API bill, and the second LLM provider. Here's the actual recipe, ported from a hotel in Tokyo to a Mac mini in Chicago in forty minutes.

What Happened in AI in May 2026

A month that turned the "agentic turn" from talking point to shipping product. Google I/O, Opus 4.8, a $65B raise, and the infrastructure race to run your agents 24/7.

SkillOpt: gradient descent for your SKILL.md

Microsoft's SkillOpt is the first paper to treat agent skill files as trainable parameters — propose an edit, evaluate on held-out examples, accept only on strict improvement. Here's what it found and what it means for teams building with agents.

Claude: Anthropic just shipped most of OpenClaw

I built a 200-line harness called conseiller to test Anthropic's new advisor tool — a fast executor model that consults a stronger model mid-generation. Two days later Anthropic shipped Claude Managed Agents, Multi-agent Orchestration, Dreams, Routines, and Remote Agents. Here's both halves: what I built and what they shipped, and how the pieces fit together into something a lot like OpenClaw.

Claude: How prompt caching actually works

A practical explainer for both developers and everyday Claude users: what prompt caching is, what gets reused, what breaks it, and how to make long sessions cheaper and faster.

How I make tokens last longer

A simple set of habits I use to keep long AI coding sessions from getting bloated: better one-shot prompts, matching model and thinking level to the job, understanding cache behavior, and using cheaper orchestrators when it makes sense.

OpenClaw: I Switched My Agent Stack from Claude to OpenAI Codex

Anthropic shutting down OAuth-based Claude Code access forced my hand. Here's how I moved OpenClaw to OpenAI Codex, why Codex makes more sense inside a real agent harness than it did on its own, and why brainpack changes the switching cost.

gstack: Garry Tan's Claude Setup Is 🔥

The Y Combinator CEO open-sourced his entire Claude Code workflow. Here are the 10 skills worth knowing — including why office-hours should be the first thing you run on any new idea.

Loooom: I Built a Skill to Teach Claude to Hear Music

I used Gemini to write a Loooom skill, installed it in Claude Code, and got a full audio analysis report on a 37-second piano recording of Espresso. Turns out AIs teaching AIs new senses is a surprisingly powerful pattern.

beatbrain: 3 Seconds to 200ms

I rebuilt the beatbrain backend in an afternoon. Parallel fetching, Firestore caching, and a podcast discovery engine that indexes 100+ categories. Here's the whole story.

Kotsu: The Knack for Japanese

I built a Japanese learning site in a morning because I wanted something I could pull up on my phone and just look at characters. Here's how Gemini wrote the prompt and magerbot built the whole thing.

DM your agent with Claude Code Channels

Claude Code's new channels feature lets you push messages from Telegram and Discord into a running session. Here's how it works, why mobile access changes everything, and how I'd wire it into my projects.

Loooom: I Built It for the Bots

Most websites beg search engines for attention. I flipped it — Loooom is machine-first, humans secondary. Here's what that actually means in practice.

How to Write, Eval, and Iterate on a Skill

Part 2 of the prompt verification series. We covered output quality testing with promptfoo — now we tackle the harder problem: does your skill even fire?

Build Your Own Agent Team with ACP

I run two AI agents — magerbot handles code and ops, genny runs my life. Inspired by the Agent Communication Protocol, here's how I got them to actually talk to each other. Now with a full TUI built on the Claude Agent SDK.

promptfoo: The Ultimate Guide to Unit Testing Your AI Prompts

Stop shipping AI features blind. Here's everything you need to know about unit testing prompts — from five-minute quick starts to CI/CD pipelines, agent workflow testing, and building a regression suite that actually catches breakage.

Building a Music Agent CLI with pi-mono

How I used the pi-mono toolkit — the same engine behind OpenClaw — to build a free, terminal-based music friend that reads the beatbrain discover feed and recommends what to listen to.

Moving Beyond the Prompt: How OpenClaw Actually Does the Work

A practical guide to building a multi-agent AI system with OpenClaw. One principal agent, multiple specialists, shared skills, and the workspace files that give them personality. Includes real examples from my blog, sports app, and music discovery projects.

beatbrain: A Social Music Discovery App

beatbrain is a social music discovery app built on Go Fx and Firestore. Find hot new releases, share your favorites, and see what your friends are actually listening to — Spotify meets Last.fm, built from scratch.

Building a coffee API with Go Fx and Firestore

How I used Go Fx dependency injection and Firestore to build an open coffee bean database and REST API from scratch — full walkthrough from blank main.go to deployed app.

Hello World

mager.co is back after years away. Here's what I'm building, what I'm obsessing over, and why this time it sticks.

102 entries · newest first

Made in Chicago. Always a work in progress.Keep up via RSS ↗

Search articles

Type a title to find an article.