Blog
Writing about software engineering, AI, and things I've built.
-
Learning LLMs · Part 3
๐ง What a Mixture of Experts Layer Actually CostsImplementing sparse Mixture of Experts routing: a one-line bug that silently froze the router, why sparse computation was not faster, and what an auxiliary load-balancing loss fixes.
-
Learning LLMs · Part 2
๐ง Learning What Transformer Architecture Choices Actually DoChanging model capacity, normalization, positional encoding, and attention structure one experiment at a time, and what each change actually did.
-
Learning LLMs · Part 1
๐ง Learning How GPTs Work by Building One from ScratchBuilding a small GPT first with scalar Python and hand-written autograd, then in PyTorch, and reflecting on what made backpropagation, attention, and batching click.
-
๐ญ AIware Observability: Monitoring and Understanding AI Systems
Traditional observability tells you that an agent was slow or expensive, but not why it decided what it did. This post works through what breaks when you point ordinary metrics, logs, and traces at LLM-powered agents, and what cognitive observability adds on top. Based on my Watson paper and the AIware Leadership Bootcamp slides, rewritten for a general audience.
-
AI Experience · Part 3
๐งญ My Experience Learning to Use AI for Software Development: Q2 2026Q2 was less about adopting a new tool and more about which habits survived once the novelty wore off. Codex became my daily driver, Claude Code stayed for the heavier work, and context management turned into the thing I think about constantly. Also covers comprehension debt, the code that ships without anyone really understanding it.
-
AI Slop · Part 2
๐งน Eleven Families of AI Code SlopThe practical companion to the first slop post: what AI code slop actually looks like, organized into eleven families by the kind of damage each one does. The taxonomy came from surveying 70+ AI code quality tools, analyzing 42 of them in depth, and cataloging 575 individual rules. The families run from hygiene debris and annotation noise to fake-done implementations and naming drift.
-
AI Slop · Part 1
๐งน When Code Is Cheaper to Produce Than to JustifyAfter a few months of researching how AI-generated code fails, the most important finding was not a new category of defect. It is that when code becomes cheaper to produce than to justify, old quality signals start pointing the wrong way and Chesterton's Fence stops holding. This post covers that shift and what teams seem to be converging on in response.
-
AI Experience · Part 2
๐งญ My Experience Learning to Use AI for Software Development: Q1 2026Q1 2026 was not about adopting anything new, it was about going deeper with Claude Code: running several sessions at once in tmux, building an observability tool for my own agent runs, and trying a few agent command centers. Most of the new tools did not stick, and the post is honest about which ones and why.
-
SWE-bench Architecture · Part 3
๐๏ธ From Checklists to Prose Verdicts: Building an Architectural Evaluation PipelineThe other two articles in this series describe what the architectural evaluation pipeline produced. This one is the development story: 12 experiments, what broke along the way, and how the design moved from brittle rubric checklists to prose verdicts. The question stayed the same throughout, which is whether an agent can infer what architecturally good means for a codebase and then judge a patch against it.
-
SWE-bench Architecture · Part 2
๐๏ธ What Does Claude Think Is Architecturally Important?I wanted to know what a frontier model actually values in code architecture when it is looking at a real codebase with a real problem, not answering a generic prompt about best practices. So I had Claude Code generate 500 architectural rubrics, one for every instance in SWE-bench Verified, then analyzed all of them to see which structural properties it reaches for and how much they vary by repo.
-
SWE-bench Architecture · Part 1
๐๏ธ Most of SWE-bench Verified Doesn't Require Deep Architectural UnderstandingHow architecturally demanding is SWE-bench Verified, really? I had Claude Code evaluate all 500 instances for how much codebase-specific architectural understanding a correct fix would require. 62% came back trivial or low, which says something about what the benchmark is actually measuring.
-
๐ค Inside 13 Coding Agents: Control Loops, Tools, and Tradeoffs
Coding agents all promise roughly the same thing, so I wanted to know whether they are all just an LLM wrapped in a ReAct loop. This analysis cloned 13 open-source agents and traced their control loops, tool registrations, state management, and context strategies through the actual source code, not the READMEs. They are not all ReAct loops, and the design space is wider than it looks from the outside.
-
AI Experience · Part 1
๐งญ My Experience Learning to Use AI for Software Development: Early 2023 to End of 2025The first post in this series, covering early 2023 through the end of 2025. It is a rough timeline of how I went from pasting stack traces into ChatGPT out of curiosity to making AI tools a core part of how I build software. Not a guide or a set of recommendations, just an honest account of how the habits shifted.