Artifact
2026-10-04-ghostty-mac-mini.md
Connect, detach, and pick up where I left off. The commands and shortcuts for my Mac mini workspace.
Tech / Article 8 min read
One of my favorite plugins turns design direction into something I can ask for by name. Exploring its skills, browser tools, and static analyzer made me want to help build it.
Tech / Article 6 min read
A private Python and Pydantic workspace for starting Codex tasks, leaving them running on a Mac mini, and returning to the same conversation.
Tech / Article 10 min read
A running build log for Perch, an open-source portal for Mac heartbeats, system resources, and tmux sessions. Starting with the backend and the limits of what a heartbeat can tell me.
Photo
Note 2 min read
Three everyday coding tasks with difficult edge cases, 36 exact checks, and the lowest supported reasoning effort.
Tech / Article 13 min read
Mood Shell is my small Rust prompt toolkit, with live color previews, Zsh history suggestions, Git shortcuts, and a gallery where you can make and share Markdown themes.
earendil.comVisit link curl -fsSL https://pi.dev/install.sh | sh
Note 1 min read
A small event pipeline for pick settlements, with a Firestore outbox and duplicate-safe recommendation invalidation.
Artifact
2026-10-01-frontend.md
The working report from auditing and hardening mager.co with Codex and Impeccable: findings, fixes, an Astro 7 migration, and the checks behind them.
Screenshot
After OpenAI DevDay, I started setting up my dot, magerdot. First up: the emails I'd sent myself as TODOs and never completed.
Tech / Article 7 min read
Returning to Rust after one Advent of Code puzzle, then building a small reward API and connecting it to prxps through SvelteKit.
Tech / Article 2 min read
Counterexample Lab replaces easy implementation tasks with compact regression tests. First calibration: Astra catches 7/8 faults; Sol catches 5–7/8 across three attempts each.
Note 2 min read
Running mager-bench through my ChatGPT subscription, then checking what the judge actually received.
Tech / Article 10 min read
Testing TypeSafe's Jev through Vercel AI Gateway, putting it inside a decision-heavy skill, and building a matchup reader for my reputation-based sports picks app.
Tech / Article 7 min read
What changed when I could text an agent from iMessage: everyday cleanup, questions worth following up on, and the trust that makes casual delegation possible.
Note 4 min read
A GLM 5.3 benchmark run became an investigation into the harness, the judge, and missing results.
Note 2 min read
Running Muse Spark 1.3 through mager-bench and comparing its scores with everyday agent work.
Tech / Article 10 min read
Vercel Labs shipped fx — a ~6MB, Zig-written coding agent that cold-starts in 10µs, speaks ACP, mounts MCP servers, and prints JSON. It's not a hosted service; it's a runtime you embed. Here's what it is, where I'd put it in my always-on harness, and how a large logistics operator would use it.
Tech / Article 10 min read
tmux is a 2007 terminal multiplexer that turns out to be the most native orchestration layer AI agents have: one session per agent, driven by any principal — human or model — with send-keys and capture-pane. No SDK, no plugin, no vendor lock-in; Claude, Codex, and OpenCode all speak it out of the box.
Tech / Article 11 min read
AI coding assistants confidently build the wrong thing when requirements live only in chat. OpenSpec adds a lightweight spec layer — explore, propose, build, archive — so you agree on what to build before any code is written, and the specs persist in your repo as history your agent can read back. Here's the mental model, a real change from a skills discovery portal, and where it earns its keep.
Tech / Article 19 min read
I'm moving off single-provider AI subscriptions toward a stack of parts — Eve for agents, Vercel AI Gateway as the primary model access and billing layer with no-markup provider pricing, and OpenCode Go kept as the fallback — and the enterprise version of that stack is the real product.
Tech / Article 8 min read
OpenCode's CLI is bigger than 'type opencode and start a session.' Headless runs, provider auth, model discovery, MCP wiring, session archaeology, cost stats, and upgrades — the ten commands that carry daily work, with the doc gaps called out where they bite.
Tech / Article 6 min read
Second harness migration in two months. The always-on agent on my Mac mini now runs OpenCode on $10/mo open models instead of Claude Code, reachable from my phone over my own Buzz relay instead of Telegram. The interesting part: the swap was one line, because the protocol — not the model — is the actual seam.
Note 3 min read
Trying Wayfinder to organize a large project around decisions before starting implementation.
Note 5 min read
Adding another model exposed a problem with what my benchmark was actually measuring.
Note 2 min read
Three Claude Code changes and what they mean for my always-on agent setup.
Tech / Article 5 min read
Anthropic deleted 80% of Claude Code's system prompt for Opus 5 with no measurable loss. They called it unhobbling. Here's what that means for your harness.
Tech / Article 12 min read
Block's open-source Nostr workspace puts people and agents on the same cryptographic footing — and lands at the end of a long chain of thinking about where always-on agent infrastructure should actually live.
Photo
Uber HQ, San Francisco, CA
Achievement unlocked: meeting Boris Cherny at Uber HQ.
Photo
Wrigley Field, Chicago, IL
Note 3 min read
Free models join the board, with a closer look at token accounting and run traces.
Photo
bench.mager.co
The personal coding model benchmark dashboard has its own domain now. The original post explains what mager-bench is.
Note 1 min read
Adding a playable raycasting challenge to mager-bench, where the output can speak for itself.
Note 2 min read
Moving mager-bench toward free models, repeated runs, and multiple judges.
Note 2 min read
New mager-bench challenges cover testing, debugging, async Python, and SQL.
Tech / Article 8 min read
A Skill is packaged know-how. An Agent is that know-how put to work autonomously. Subagents are where the work scales past what any single context can hold.
Tech / Article 8 min read
OpenRouter lets you pick a different model for every step in a pipeline. Here's how to use Fable for planning and Sonnet for execution — with runnable TypeScript.
Tech / Article 5 min read
Instead of reading someone else's leaderboard, build a small set of tasks you actually care about and run them yourself every time a new model drops — Simon Willison's SVG pelican test, but for code.
Tech / Article 5 min read
In March I wrote the theory. Zach from Warp shipped the implementation. Here's how a working cloud factory maps to the architecture I laid out.
Tech / Article 8 min read
Most evals talk is about grading model output. Skill evals grade a different thing — the SKILL.md artifact you wrote — with real numbers from SkillsBench to back it up.
Note 2 min read
A Claude Code plugin for investigating agent failures and building evaluations you can trust.
Tech / Article 5 min read
A small Python voice agent that remembers the thread, streams Claude's reply to the terminal, and speaks it aloud through ElevenLabs — no ffmpeg, just afplay.
Tech / Article 15 min read
Define your agent in a directory, deploy it to Vercel's cloud with one command, and access it from anywhere. Months in, Eve has grown a platform around that model — capability registry, sandbox, subagents, agent-to-agent calls, MCP, evals — and my agent is still live, driven remotely from the eve TUI.
Tech / Article 6 min read
A Skill is reusable know-how Claude reaches for on its own. A Workflow is an explicit pipeline you wire up and control. Here's the difference, what you can build with each, and when to reach for which.
Note 2 min read
Loooom v1.0 adds behavioral tests to its skill evaluations. A few early failures exposed assumptions in the tests themselves.
Tech / Article 4 min read
How I worked with Claude through five rounds of image generation to design a logo for my Japanese learning app — and ended up inventing a kanji that hides a smile.
Photo
Lady Gregory's, Andersonville, Chicago, IL
Table card at Lady Gregory's, in the heart of Andersonville. A whole vocabulary of the month, set in rainbow.
Tech / Article 5 min read
A plain-English walkthrough for setting up your own always-on AI assistant on a Mac mini — OpenClaw, Google Gemini, and Tailscale — written for a first-timer.
Photo
Fenway, Boston, MA
Sox game with Dad and Matt.
Note 1 min read
How I keep a Mac mini agent running through crashes, model changes, and reboots.
Tech / Article 11 min read
I love OpenClaw. I hate that it doesn't run on my Claude Pro subscription. Turns out Claude Code, with the Telegram channels plugin and one CLAUDE.md, is the same harness — minus the daemon, the API bill, and the second LLM provider. Here's the actual recipe, ported from a hotel in Tokyo to a Mac mini in Chicago in forty minutes.
Tech / Article 6 min read
A curated collection of high-quality skills for people who don't code — and an experiment in what actually makes a skill good.
Tech / Article 6 min read
A month that turned the "agentic turn" from talking point to shipping product. Google I/O, Opus 4.8, a $65B raise, and the infrastructure race to run your agents 24/7.
Tech / Article 7 min read
Microsoft's SkillOpt is the first paper to treat agent skill files as trainable parameters — propose an edit, evaluate on held-out examples, accept only on strict improvement. Here's what it found and what it means for teams building with agents.
Tech / Article 5 min read
OpenHuman is a desktop-first agentic assistant with persistent memory, 118+ OAuth integrations, and a token compression layer. Here's what it does and how it fits alongside an existing Claude Code harness.
Tech / Article 3 min read
Karpathy's four rules for agentic coding are worth reading — having them written down in a shared format is a useful starting point for anyone building with Claude Code.
Tech / Article 6 min read
How I moved magerbot's brain from @-imported markdown files into gbrain's Postgres-native semantic memory layer — what broke, what the gotcha was, and why the context model is fundamentally better.
Tech / Article 5 min read
Garry Tan open-sourced gbrain — a self-wiring knowledge graph for AI agents. Here's what it is, why we moved to it, and exactly how we did the migration from flat markdown files.
Tech / Article 8 min read
I built a 200-line harness called conseiller to test Anthropic's new advisor tool — a fast executor model that consults a stronger model mid-generation. Two days later Anthropic shipped Claude Managed Agents, Multi-agent Orchestration, Dreams, Routines, and Remote Agents. Here's both halves: what I built and what they shipped, and how the pieces fit together into something a lot like OpenClaw.
Tech / Article 6 min read
I built a Go Bubble Tea starter for local model servers, used Gemma 4 through llama.cpp, and split the TUI into llocal.
Tech / Article 10 min read
I'd been seeing chatter about Hermes Agent from Nous Research, so I installed it locally and put it to work on this blog. Notes on the pitch, the SOUL.md system, and what it actually felt like to use.
Tech / Article 12 min read
A practical explainer for both developers and everyday Claude users: what prompt caching is, what gets reused, what breaks it, and how to make long sessions cheaper and faster.
Tech / Article 7 min read
A simple set of habits I use to keep long AI coding sessions from getting bloated: better one-shot prompts, matching model and thinking level to the job, understanding cache behavior, and using cheaper orchestrators when it makes sense.
Tech / Article 25 min read
I reverse engineered several of my own sites into DESIGN.md files to see how much of a design system can actually be described, and why writing down design intent might be more reusable than it looks.
Tech / Article 7 min read
A practical tour of Claude Code flags that are easy to miss but genuinely useful once you move past the default interactive loop.
Tech / Article 5 min read
Anthropic shutting down OAuth-based Claude Code access forced my hand. Here's how I moved OpenClaw to OpenAI Codex, why Codex makes more sense inside a real agent harness than it did on its own, and why brainpack changes the switching cost.
Tech / Article 7 min read
The Y Combinator CEO open-sourced his entire Claude Code workflow. Here are the 10 skills worth knowing — including why office-hours should be the first thing you run on any new idea.
Tech / Article 8 min read
I tested Anthropic's official Claude plugins for knowledge workers. Here are the 10 that deliver the most value for PMs, engineers, sales teams, and operators.
Tech / Article 4 min read
I used Gemini to write a Loooom skill, installed it in Claude Code, and got a full audio analysis report on a 37-second piano recording of Espresso. Turns out AIs teaching AIs new senses is a surprisingly powerful pattern.
Tech / Article 5 min read
I rebuilt the beatbrain backend in an afternoon. Parallel fetching, Firestore caching, and a podcast discovery engine that indexes 100+ categories. Here's the whole story.
Tech / Article 5 min read
I built a Japanese learning site in a morning because I wanted something I could pull up on my phone and just look at characters. Here's how Gemini wrote the prompt and magerbot built the whole thing.
Tech / Article 7 min read
Dogfooding Karpathy's autoresearch pattern on my own skill marketplace. How I'm using evals and tight feedback loops to make the learn-anything skill measurably better.
Tech / Article 6 min read
Claude Code's new channels feature lets you push messages from Telegram and Discord into a running session. Here's how it works, why mobile access changes everything, and how I'd wire it into my projects.
Tech / Article 6 min read
Everyone's talking about building a software factory. Here's where the term came from and how engineers can start thinking about building one.
Tech / Article 6 min read
12 hours before my bracket was due, I used Gemma-3-27b to generate unique insights for all 32 first-round games. Here's what the AI found — and what it got wrong.
Tech / Article 6 min read
LangChain just dropped Open SWE — an open-source framework for building internal coding agents like Stripe's Minions, Ramp's Inspect, and Coinbase's Cloudbot. Here's what it is, how it works, and how to customize it.
Tech / Article 6 min read
mager.co is no longer just a blog. It's a corporation. Here's how I staffed it with 145 specialized AI agents using agency-agents and OpenClaw.
Tech / Article 7 min read
Andrej Karpathy open-sourced a loop where AI agents run experiments, measure results, and keep what works — all while you sleep. Here's how the pattern works and how I'm applying it beyond LLM training.
Tech / Article 12 min read
The Claude Agent SDK gives you the same engine that powers Claude Code, fully programmable. Here's how to build a custom TUI with it in 10 minutes.
Tech / Article 4 min read
I built an MCP server for Loooom so AI agents can search, explore, and install Claude Code skills without ever leaving their context.
Tech / Article 2 min read
How a weekend contribution to OpenClaw replaced my autossh aliases with `openclaw tunnel up/down/status` — and what I learned reading a real codebase to do it right.
Tech / Article 4 min read
Stop stuffing your prompts. OpenViking gives AI agents a filesystem-native brain — tiered, retrievable, self-evolving context at 91% lower token cost.
Tech / Article 8 min read
Eighty years after Asimov's Three Laws of Robotics debuted, we're building the future he imagined—without the safeguards. What the 'Father of Robotics' got right, where his vision fails, and why 2026's AI alignment problem is harder than fiction.
Tech / Article 5 min read
I kept fixing the same SEO issues by hand — missing keywords, empty hero images, weak descriptions. So I built a Claude Code skill that audits any blog's frontmatter and runs quality evals.
Tech / Article 14 min read
LangGraph is the production framework for complex agent workflows. Here's how to build a real-time chat system with persistent state, human-in-the-loop, and multi-agent orchestration.
Tech / Article 2 min read
LangChain just shipped DeepAgents — a batteries-included agent harness that brings Claude Code's magic to any model. Here's your 10-minute deep dive.
Tech / Article 5 min read
Most websites beg search engines for attention. I flipped it — Loooom is machine-first, humans secondary. Here's what that actually means in practice.
Tech / Article 11 min read
Part 2 of the prompt verification series. We covered output quality testing with promptfoo — now we tackle the harder problem: does your skill even fire?
Tech / Article 11 min read
A deep dive into the two most powerful tools for building production-grade multi-agent systems — LangGraph's graph-based orchestration and Anthropic's Claude Agent SDK (formerly Claude Code SDK).
Tech / Article 7 min read
Stop re-prompting every AI session. One file. Every AI knows you — and your agents. Introducing ME.md on Loooom.
Tech / Article 8 min read
I run two AI agents — magerbot handles code and ops, genny runs my life. Inspired by the Agent Communication Protocol, here's how I got them to actually talk to each other. Now with a full TUI built on the Claude Agent SDK.
Tech / Article 10 min read
I built a second AI agent to manage the parts of my life that code can't fix — exercise, nutrition, travel, and living to 100.
Tech / Article 12 min read
Stop shipping AI features blind. Here's everything you need to know about unit testing prompts — from five-minute quick starts to CI/CD pipelines, agent workflow testing, and building a regression suite that actually catches breakage.
Tech / Article 7 min read
Spec compliance tells you if a skill is readable. Evals tell you if it's actually good. Here's how we added a public quality score to every Loooom plugin.
Tech / Article 10 min read
How to run OpenClaw on a Mac Mini 24/7, lock it down with Tailscale, and load your agent's brain with brainpack — so your laptop can reach it from anywhere on your tailnet.
Tech / Article 6 min read
Your AI agent has memories, skills, and a personality. Here's how to pack it all up and ship it to a new machine — whether you're the human or the agent reading this.
Tech / Article 6 min read
How I used the pi-mono toolkit — the same engine behind OpenClaw — to build a free, terminal-based music friend that reads the beatbrain discover feed and recommends what to listen to.
Tech / Article 3 min read
I analyzed three of my projects, interviewed myself about what makes a UI hot, and packaged it all into a reusable skill for Claude Code.
Tech / Article 10 min read
A practical guide to building a multi-agent AI system with OpenClaw. One principal agent, multiple specialists, shared skills, and the workspace files that give them personality. Includes real examples from my blog, sports app, and music discovery projects.
Tech / Article 10 min read
A practical guide to building two-stage AI recommendations: use embeddings for fast retrieval, then small LLMs like Gemma 3 for natural language explanations. The real skill? Curating context, not writing algorithms.
Tech / Article 2 min read
beatbrain is a social music discovery app built on Go Fx and Firestore. Find hot new releases, share your favorites, and see what your friends are actually listening to — Spotify meets Last.fm, built from scratch.
Tech / Article 58 min read
Instead of just using a single language, I wanted to solve the puzzle in a language I know, then lurk the internet for the solution in another language each day.
Tech / Article 16 min read
How I used Go Fx dependency injection and Firestore to build an open coffee bean database and REST API from scratch — full walkthrough from blank main.go to deployed app.
Tech / Article 1 min read
mager.co is back after years away. Here's what I'm building, what I'm obsessing over, and why this time it sticks.
Tech / Article 6 min read
My explorations into decentralized apps and blockchain.