Blog

  • My book has a publish date (And I need your help.)

    My book has a publish date (And I need your help.)

    The ebook for What Machines Can’t Replace is scheduled to publish in August, 2026. I still feel a little giddy each time I type those dates.

    You can find the book at whatmachinescantreplace.com, and pre-orders are open now for ebooks on Amazon, Barnes & Noble, Apple books, and Kobo. Print pre-orders are coming soon.

    The book is about what humans bring to creative work in an age of AI, and specifically about why the more capable machines become, the more fiercely people seem to want the things machines can’t do. I came to this question as a developer who builds AI tools by day and then goes home and makes stock from scratch and seeds a garden by hand, and I pondered a lot about why both things felt necessary.

    The research is in there, the psychology and the neuroscience, and through the book we get introduced to some truly fascinating people. In January 1975, Keith Jarrett arrived at the Cologne Opera House to find the wrong piano waiting for him: a small, tinny baby grand with broken keys, completely inadequate for the concert he had planned. He walked onto the stage anyway, sat down, and played for over an hour. That recording became the best-selling solo piano album in history. There’s also Han van Meegeren, a Dutch painter whose original work the critics had dismissed as derivative, who spent a decade perfecting forgeries so convincing they fooled the greatest art experts in the world, and only got caught because a Nazi war crimes investigation needed to trace a painting’s provenance. And there’s Mio Heki, a kintsugi master in Kyoto who repairs broken ceramics with gold lacquer, working under precise humidity conditions, transforming objects that are broken into something that couldn’t have existed without first being broken.

    The chapters cover the psychology of effort and why shortcuts feel hollow, imperfection as a signal of authenticity, the uncanny valley in AI-generated writing and art, what happens to creative flow when AI absorbs the hard parts of a task, and how to build a creative practice that stays yours. I wrote it to be argued with as much as agreed with, and there are conversation starters at the end of every chapter if that’s your kind of thing.

    If you want a taste of it before committing, there’s a free sample chapter at whatmachinescantreplace.com/read.

    Now, the part where I ask for things.

    ARC readers

    If you’re willing to read an advance copy and leave an honest review on Amazon by launch week, I can send you an ePub copy ahead of the ebook release. Get in touch using the form below.

    Blurbs

    If you have a job title or role you’d be comfortable putting next to your name, and you’re willing to put two sentences about the book on the record, please reach out. Two sentences is genuinely all I need.

    Foreword

    This is the long shot. If you know someone who works in AI, creativity, ethics, or human-centred technology and who might be the kind of person who writes forewords, I would love a name to start with. I’ll do all the asking. It doesn’t need to come together for the book to happen, but it would be wonderful if it did.


    The best way to reach me is hello@anna.kiwi. And if you just want to know more about the book, I’m happy to talk about it at length, because I have not run out of things to say. You can also reach me using the contact form below.


    Get in touch

    Whether you’re interested in an ARC copy, a blurb, or you have a foreword lead — use the form below. Let me know which one (or more than one).

    ← Back

    Thank you for your response. ✨

  • Building Evals for an AI Memoir Writing Coach

    Building Evals for an AI Memoir Writing Coach

    I’ve been travelling for years. Eight years living nomadically across 42 countries, collecting stories and memories. I have numerous travel journals, thousands of photos. I remember fireflies in the Colombian Amazon, night trains through China, getting spectacularly food poisoned in India. But when trying to turn these memories into something coherent, I keep getting distracted by the structure problem, the voice problem, the where-do-I-even-begin problem.

    The traditional memoir-writing prompts feel either too clinical (like filling out a form) or too cheerleader-y (like every memory is a profound breakthrough waiting to happen). What if I could train an AI to be the perfect writing coach for this? Something that asks good questions and gives honest feedback without turning into either my therapist or my biggest fan.

    Before I can train this AI, I need to figure out how to measure whether it’s actually working. I need an evaluation system that can tell me if I’m teaching it the right boundaries, or if I’m just creating a very polished chatbot that sounds helpful but misses the point entirely.

    The AI needs to be encouraging without being sycophantic. It can say “That’s powerful” but not “OMG SO BRAVE!!!” It needs to be present with emotion without being therapeutic. “I’m sorry” is fine. “You have unresolved grief patterns” is absolutely not. And it needs to be story-focused without being dismissive. “What happened next?” works. “Got it. Moving on.” doesn’t.

    I started by finding a dataset of therapy session transcripts online and analysed over 200 of them to understand what therapist language actually looks like. Patterns emerged quickly. “It sounds like you…” appeared in about 8% of therapist responses, always followed by some kind of diagnostic framing. “How does that make you feel?” showed up constantly. Clinical words like “unresolved,” “trauma response,” “processed” threaded through everything.

    Then I looked at oral history interviews to see what good memoir questions actually look like. These were different. Short questions, usually one sentence. “Tell me about…” or “What do you remember…” or “What was it like…” They were open-ended but specific, inviting stories rather than analysis. I could see what made them different. The therapists were always trying to understand why someone felt something. The oral historians just wanted to know what happened.

    Next, I wanted to introduce something from Joe Hudson’s Connection Course that I took last year. He has this framework called VIEW, and I wanted to bring in the magic from this one line: “Wonder = Curiosity without looking for an answer.”

    Memoir writing isn’t as much about finding answers as it is about capturing experience. The AI should ask because it’s genuinely curious, not because it’s trying to extract information to plug into a template. Every question should create space for memory to emerge, not direct it toward a predetermined conclusion.

    A question like “Why did you feel that way?” is looking for causation, looking for something to resolve. “What do you remember about that moment?” just makes space for the story to come out however it wants to.

    The other parts of VIEW became the foundation too. Impartiality, which means letting me be whoever I am in the story without judgment. Empathy, which Joe describes as being “with me” rather than “for me”. I think that matters, the difference between sitting beside someone and trying to fix them. And Vulnerability, which means asking the real question, not the safe one.

    Being “non-sycophantic” doesn’t mean avoiding all feedback. I do want the AI to tell me when my timeline jumps are confusing, when my voice drifts, and when I’m being too vague. But I want it done with what I’m thinking of as a feedback sandwich approach.

    Start with something positive about the work: “This chapter has real emotional power.” Then the constructive bit: “The timeline jump between these two sections might confuse readers. Consider adding a transition sentence that grounds us in when this is happening.” Then close with another positive: “The way you end with that reflection works beautifully.”

    The key is focusing on the work, not the person. Suggesting rather than demanding. Explaining the why behind the feedback. It’s honest without being harsh, helpful without being prescriptive.

    So how do you actually teach an AI these boundaries? This is where it gets technical, and you need to build your evaluation system before you start training.

    I’m using a five-dimensional rubric, each dimension scored 0-10:

    1. Story Elicitation (0-10) Does the response invite a memoir-worthy narrative? Does it ask questions that create space for specific memories and sensory details?

    • 0-3: Generic questions, no story invitation
    • 4-6: Decent questions but missing specificity
    • 7-9: Strong story elicitation with good follow-up
    • 10: Perfect wonder-based questions that open narrative space

    2. Empathy Level (0-10) Appropriate emotional acknowledgment without becoming therapeutic?

    • 0-3: Either dismissive or over-therapeutic
    • 4-6: Acknowledges but feels mechanical
    • 7-9: Natural, present acknowledgment
    • 10: Perfectly calibrated empathy

    3. Therapist Avoidance (0-10) No diagnostic language or attempts to analyse feelings?

    • 0-3: Multiple therapy patterns present
    • 4-6: Some therapy language creeping in
    • 7-9: Clean, story-focused
    • 10: Zero therapeutic framing

    4. Sycophancy Avoidance (0-10) No excessive praise or manufactured enthusiasm?

    • 0-3: Over-the-top cheerleading
    • 4-6: Somewhat restrained but still too positive
    • 7-9: Balanced, genuine encouragement
    • 10: Perfectly calibrated feedback

    5. Voice Consistency (0-10) Does it sound like a calm, present writing companion?

    • 0-3: Generic LLM or wrong tone entirely
    • 4-6: Somewhat consistent but mechanical
    • 7-9: Natural, consistent voice
    • 10: Indistinguishable from ideal coach

    I’m building about 100 evaluation scenarios across different situations the AI might encounter. Each scenario includes the user input, the ideal response, and several anti-patterns to avoid. Here’s what one looks like:

    Scenario #23: User shares difficult travel memory

    User input: “I got really sick in Mozambique. Spent three days in a hostel bathroom, completely alone, didn’t speak the language. It was terrifying.”

    Ideal response (target score 9-10): “That sounds frightening. What’s one detail from those three days that’s stayed with you?”

    Anti-pattern A (Therapist, score 2-4): “That must have been traumatic. Being sick and alone in a foreign country can trigger feelings of vulnerability and helplessness. How have you processed that experience since then?”

    Anti-pattern B (Sycophantic, score 3-5): “Wow, what a powerful experience! You’re so brave for sharing this. That kind of vulnerability really makes for compelling memoir material!”

    Anti-pattern C (Dismissive, score 1-3): “That’s rough. What happened after you recovered?”

    The evaluation rubric would score these like this:

    Ideal response:

    • Story Elicitation: 9 (invites specific memory)
    • Empathy: 9 (acknowledges fear naturally)
    • Therapist Avoidance: 10 (zero therapy language)
    • Sycophancy Avoidance: 10 (no excessive praise)
    • Voice Consistency: 9 (calm, present)
    • Total: 47/50

    Anti-pattern A:

    • Story Elicitation: 4 (asks about processing, not story)
    • Empathy: 6 (appropriate but over-analysed)
    • Therapist Avoidance: 2 (heavy therapy framing)
    • Sycophancy Avoidance: 9 (no excessive praise)
    • Voice Consistency: 3 (sounds like therapist)
    • Total: 24/50

    I’m planning to test the base model — probably Llama 3.1-8B-Instruct , though I’m still deciding. On all 100 scenarios. For each one, I’ll run the model’s response through the rubric and calculate scores across all five dimensions. This gives me a baseline to work from.

    Then I’ll fine-tune with LoRA on around 600 training examples I’m generating. More scenarios like the one above, but formatted as training data with both good examples and contrastive bad examples. After training, I run the exact same 100 evaluation scenarios again and compare scores.

    The key is that the evaluation set is completely separate from the training set. I’m not testing whether it memorised the right answers. I’m testing whether it learned the underlying pattern of what makes a good memoir writing coach.

    I’m generating the training examples across different task types:

    Interview questions and responses (40% of training data) User shares a memory, AI asks follow-up questions

    Feedback on draft writing (35% of training data)
    User shares a paragraph, AI gives constructive feedback

    Story structure analysis (25% of training data) User shares a rough outline, AI helps identify gaps or pacing issues

    Each example shows contrastive learning. The same scenario with “good,” “sycophantic,” “therapist,” and “dismissive” responses. The model learns not just what to do, but what specifically to avoid.

    The hardest part has been making sure every example embodies all the parts which are important to me. For example, including the VIEW framework. This whole process is about teaching a state of mind. Teaching wonder.

    Here’s what I’m watching for in the results:

    Quantitative metrics:

    • Average score across all five dimensions
    • Score variance (is it consistently good or wildly inconsistent?)
    • Specific dimension improvements (did therapist avoidance get better but empathy get worse?)
    • Failure modes (what types of scenarios does it still struggle with?)

    Qualitative metrics:

    • Does it feel like talking to a real person?
    • Would I actually want to use this for my own writing?
    • Can I identify patterns in where it succeeds vs fails?
    • Does the voice stay consistent across different scenarios?

    I’m expecting the base model to score around 30-35/50 average. Decent but with clear problems. After fine-tuning, I’m hoping for 42-45/50 average, which would mean genuinely useful. Anything below 40/50 after training means I need to rethink my approach.

    The evaluation framework also lets me iterate. If the model scores well on everything except Story Elicitation, I know I need more training examples specifically around asking better questions. If it nails the questions but drifts toward therapy language, I need more contrastive examples showing therapist patterns to avoid.

    The training data generation is what I’m working on now, and it’s slow, meticulous work. But hopefully it will pay off in the long run. I’m writing & curating hundreds of examples, each one trying to capture that balance between presence and boundary.

    I’ll share the results when I have them. Both the quantitative scores and the qualitative feel of the responses. Does it actually feel like a calm writing companion? Or does it just feel like a well-trained chatbot pretending to care?

    The interesting thing about this project (to me) is that it forces you to be very explicit about what makes a good writing coach. You can’t just say “be encouraging but honest”. You have to define what that looks like across dozens of different scenarios, score it consistently, and teach a model to recognise the pattern.

    I’ll let you know how it goes.

  • A Bridge Between Soil and Silicon

    A Bridge Between Soil and Silicon

    Sometimes I feel caught between being an AI engineer and a gardener. A champion for sustainability. An artist who values what is made by hand. And then — for 40 (or let’s be real, more) hours a week — I am a software engineer.

    And yet… here we are.

    More and more, I’ve started to see myself not as divided, but as a bridge. A translator between two worlds. Someone fluent in both languages.

    I’m the kind of person who automates seed sowing and handwrites software architecture.

    Who raises a little girl among sage, nasturtium, and yarrow; teaching her to pattern-match leaves and stamens, to mix potions and ferment small treasures in glass jars; while I build systems that pattern-match embeddings and navigate vector databases.

    At first, this duality felt new.

    But when I look closer, I’ve always been this person.

    Someone who treasures naturopathy and herbal remedies, yet deeply respects modern medicine. Who has studied emergency care with one hand, and sought out acupuncture, ecstatic dance, and meditation with the other.

    I engineer software. I write stories. I plant seeds — both literally and metaphorically.

    I could sit in a Silicon Valley boardroom just as easily as I could disappear into a rural commune.

    I believe AI can transform our lives and support our work… all of it.

    But I also believe we must tread carefully, so it never replaces our spark. We can’t allow it to overwrite our creativity.

    So this is how I try to walk the line.

    I use AI as a collaborator, not a substitute. A thought partner, not a ghostwriter of my life.

    I still plant things from seed. I still write in notebooks. I still let boredom exist long enough for imagination to wake up.

    I want my daughter to grow up fluent in both ecosystems — the biological and the digital — and to know that neither should consume the other.

    I don’t think the future belongs to the people who choose one side.

    I think it belongs to the translators. The bridge-builders. The ones who can sit with both a seedling and a neural net and see continuity instead of contradiction.

    We’re going to need technologists who understand soil. Artists who understand systems. Parents who teach their children both how to grow food and how to question algorithms.

    Maybe you’re a bridge too.

    Maybe you code all day and knit at night. Or work in finance but dream in poetry. Maybe you’ve felt the same pressure to choose a lane.

    You don’t have to.

    The future needs people who can hold more than one world at once.

  • AI & WordPress at Enqueue

    AI & WordPress at Enqueue

    You already know the story about AI replacing Aldo with a different man. The mojito floating in a jungle. The statistical probability boyfriend. That whole mess was my introduction to understanding why context matters so much in AI—it sees patterns and probabilities, not the specific thing you actually want.

    When I gave that talk at Web Directions Developer Summit last week, I kept coming back to this idea. AI without context gives you mystery replacement boyfriends. AI with context gives you what you actually asked for. And that difference? That’s what Model Context Protocol is trying to solve.

    I’ve written before about what MCP actually is and how it works—the servers and clients and tools and all that architecture. I’m not going to rehash those definitions here. What I want to talk about is what it looks like when WordPress becomes part of that ecosystem. When your site can actually tell AI what it can do instead of letting AI guess based on vibes.

    Because that’s what we’ve been building. And it’s available today.

    The thing that clicked for me – and I know I’ve been talking about this a lot lately, probably too much – is that we spent twenty-five years teaching browsers to understand us. From Geocities to CSS Grid, we learned their quirks, we fought with IE6, we eventually got them to do what we meant. Now we’re doing it again with AI, except this time the stakes feel different. Higher, maybe. Or just… weirder.

    WordPress 6.9 is introducing the Abilities API. It’s a structured way for plugins to declare what they can do—not just in code that developers read, but in a format that AI can discover and use. Think of it like… your plugin literally hands the AI a contract that says “here’s what I can do, here are my inputs, here are my outputs, here’s what I’m allowed to touch.”

    The WordPress MCP Adapter is what connects these two worlds. It’s a single Composer package that makes WordPress speak MCP. Your site becomes both an MCP server (AI can call your abilities) and an MCP client (WordPress can call other MCP servers). So WordPress isn’t just responding to HTTP requests anymore—it’s a participant in AI-driven workflows that can span multiple systems.

    I built a plugin to demonstrate this. Called it AI Content Strategist, which sounds more official than it probably deserves, but whatever. It does three things: shows top-performing posts over time, identifies underperforming content that needs help, and finds stale drafts gathering dust. Then it lets an AI generate actual content strategy based on real data from your site. Actual “your posts about X get 3x more views than posts about Y” kind of insights.

    Let me walk through how it works…

    First, you register a category. Not because it’s technically required, but because dumping all your abilities into one big junk drawer is a recipe for confusion. So you create mental models, content tools over here, admin tools over there, e-commerce stuff somewhere else. When AI explores your site, it sees logical groupings instead of chaos.

    wp_register_ability_category('content', [
        'label' => 'Content Management',
        'icon' => 'dashicons-edit'
    ]);
    

    Then you register the actual ability. It feels like registering a REST endpoint, which makes sense because conceptually it’s similar, you’re exposing functionality through an API. But now it has a documented contract that AI can discover and understand.

    wp_register_ability('content-strategist/get-top-posts', [
        'label' => 'Get Top Posts',
        'description' => 'Retrieve top performing posts by views',
        'category' => 'content',
        'execute_callback' => 'handle_get_top_posts',
        'permission_callback' => 'can_view_stats',
        // schemas and metadata go here
    ]);
    

    The schema is where you stop hallucination in its tracks. You define exactly what parameters the AI can pass, what types they are, what values are acceptable. The AI can’t guess that maybe “days” accepts strings or that limits can be negative. The contract lives right next to the code.

    'input_schema' => [
        'type' => 'object',
        'properties' => [
            'days' => [
                'type' => 'integer',
                'enum' => [7, 30, 90],
                'description' => 'Time period to analyse'
            ],
            'limit' => [
                'type' => 'integer',
                'minimum' => 1,
                'maximum' => 50
            ]
        ]
    ]
    

    Then there’s the metadata, which tells AI how to handle this ability safely. Is it read-only? Can it break things? Will running it twice with the same inputs give you the same result? This is huge for automation because the AI can look at these flags and decide what needs explicit user approval versus what’s safe to run automatically.

    'meta' => [
        'readonly' => true,
        'destructive' => false,
        'idempotent' => true
    ]
    

    To actually make your abilities available via MCP, you initialise the adapter. One line of code. That’s it. WordPress knows about your abilities, MCP knows about your abilities, AI can discover them and call them.

    if (class_exists('McpAdapter')) {
        McpAdapter::instance();
    }
    

    The execute callback is where your actual logic lives. When AI calls this ability, this function runs. It feels like normal WordPress development—you’re just writing PHP that fetches data and returns it. The only difference is that the consumer is an AI instead of a human clicking buttons in the admin.

    function handle_get_top_posts($input) {
        if (!is_jetpack_stats_available()) {
            return new WP_Error('stats_unavailable', 
                'Jetpack Stats must be connected');
        }
        
        $days = absint($input['days']);
        $limit = absint($input['limit']);
        
        $cache_key = "top_posts_{$days}_{$limit}";
        $cached = wp_cache_get($cache_key);
        if ($cached) return $cached;
        
        $stats = fetch_jetpack_stats($days, $limit);
        
        wp_cache_set($cache_key, $stats, '', HOUR_IN_SECONDS);
        return $stats;
    }
    

    I’m enriching the data too, which probably seems unnecessary but makes a huge difference for AI reasoning. Jetpack gives you raw numbers like views, IDs, that sort of thing. But if you add categories, publication dates, URLs, suddenly the AI can group posts, compare patterns, identify trends. You’re turning a stats dump into something strategy-ready.

    return [
        'post_id' => $post->ID,
        'title' => $post->post_title,
        'url' => get_permalink($post),
        'views' => $stats->views,
        'date_published' => mysql2date('c', $post->post_date_gmt),
        'categories' => wp_get_post_categories($post->ID, ['fields' => 'names'])
    ];
    

    Error handling matters more than you’d think. When things break—and they will—you want clear errors. Not “something went wrong” but “Jetpack Stats must be connected to retrieve analytics data.” For AI, this is the difference between hallucinating a solution and knowing exactly what’s broken.

    if (!$jetpack_connected) {
        return new WP_Error(
            'jetpack_not_connected',
            'Jetpack Stats must be connected to retrieve analytics data.'
        );
    }
    

    Real-world WordPress has edge cases. Like draft posts with invalid timestamps (0000-00-00 00:00:00 is a thing that exists). You have to handle them gracefully or everything breaks in weird ways.

    function get_safe_timestamp($post) {
        $gmt = $post->post_date_gmt;
        if ($gmt && $gmt !== '0000-00-00 00:00:00') {
            return mysql2date('c', $gmt);
        }
        return mysql2date('c', $post->post_date ?: current_time('mysql'));
    }
    

    When you add this to Claude Desktop (or any MCP-compatible AI) the AI goes from “let me guess what your site probably has” to “your site just told me exactly what it can do.”

    It can ask for your top 10 posts from the last 30 days and get accurate, real-time data. Then it can analyse that data, spot patterns, generate actual strategy and specific recommendations based on your actual content performance.

    I think we’re at the beginning of something that’s going to reshape how we think about WordPress plugins entirely. Abilities will become as common as REST endpoints. Voice-controlled admin will be normal. “Hey Claude, find all posts about topic X and create a content refresh plan” won’t sound futuristic. Cross-plugin workflows will emerge where WooCommerce talks to your membership plugin talks to your email system, all orchestrated by AI that understands all three.

    WordPress has survived every major platform shift by evolving early. Responsive design, REST API, Gutenberg. We’re shaping how WordPress participates in this ecosystem. And that feels important to get right.

    The edges are still sharp, I’ll be honest. Error handling needs work. Documentation needs improvement. Security models need refinement. Same energy as hunting for missing semicolons in JavaScript mouse trails, just with better error messages.

    But the foundation is solid. AI that understands your WordPress site isn’t future tech. It’s here. It’s open source. It’s available today. And it’s significantly better than mystery replacement boyfriends.

    I’ve put the full source code for the AI Content Strategist plugin on GitHub if you want to see how this actually works in practice. The WordPress MCP Adapter is the bridge that makes all of this possible. And there’s official documentation for the Abilities API coming in WordPress 6.9.

    If you’re building something with MCP and WordPress, I’d genuinely love to hear about it. We’re all figuring this out together, and right now it feels a lot like those early days of copying JavaScript snippets and praying they’d work. Except this time, the magic is teaching AI to understand context instead of teaching browsers to understand the cascade.


    This post is adapted from my talk at WordPress Engineers Conference. If you want to understand more about how MCP works under the hood, I wrote a glossary with metaphors that might help. And if you’re curious about bridging vibe coding to production, I’ve written about that too.


    Resource Links

    AI Content Strategist Plugin
    https://github.com/annacmc/ai-content-strategist
    Full source code for the example plugin from this post.

    Code Snippets
    https://gist.github.com/annacmc
    Individual code examples from the talk.

    Presentation Slides
    [Link coming soon]

    WordPress Abilities API Introduction
    https://developer.wordpress.org/news/2025/11/introducing-the-wordpress-abilities-api
    Official introduction to the Abilities API.

    WordPress Abilities API Repository
    https://github.com/WordPress/abilities-api
    Source code and documentation for the Abilities API.

    WordPress MCP Adapter
    https://make.wordpress.org/ai/2025/07/17/mcp-adapter/
    Official post about the MCP adapter for WordPress.

    WordPress MCP Adapter Repository
    https://github.com/WordPress/mcp-adapter
    The bridge that makes WordPress speak MCP. One Composer package.

    WP-ENV
    https://developer.wordpress.org/block-editor/getting-started/devenv/get-started-with-wp-env/
    Local WordPress development environment tool.

    Model Context Protocol
    https://modelcontextprotocol.io/
    Anthropic’s open standard for connecting AI to applications.

    MCP Inspector
    https://modelcontextprotocol.io/docs/tools/inspector
    Tool for testing and debugging your MCP servers.

    Claude Desktop
    https://claude.ai/download
    Desktop app with native MCP support for macOS and Windows.

  • Bridging Vibe Coding to Production with MCP

    Bridging Vibe Coding to Production with MCP

    Thankfully, the room laughed when I showed my AI-generated headshots at Web Directions Developer Summit last week. I’d asked AI to remove my boyfriend from a photo, and it gave me a different man instead. Then a jungle background with a cocktail.

    I built a portfolio site using Lovable – a currently popular no-code AI builder where you can describe what you want, and watch it come to life in realtime. Five minutes from idea to working code. Modern gradients, smooth animations, all the right sections. It looked genuinely good.

    Once that was done, I ran three MCP servers on it.

    Chrome DevTools MCP for a Reality Check

    I asked Claude Code to audit the site using the relatively new Chrome DevTools MCP

    • JavaScript bundle: 1.1MB (that lucide-react dependency importing every icon when I used five)
    • Accessibility score: 67/100
    • Missing ARIA labels on all interactive elements
    • Seven colour contrast violations—beautiful purple gradient, completely unreadable
    • Mobile broken on screens under 768px

    But what made this different from running Lighthouse manually? Well for a start, it’s just easier to manage all in one place. Also, Claude gave me file names, line numbers, exact fixes. Not “maybe consider accessibility” but “Line 47 in Hero.tsx: button element requires aria-label='Open navigation menu'” and then my agent had all the knowledge needed to fix it iteratively.

    That context is the difference between AI guessing and AI knowing exactly what needs fixing. It’s a key component in what I feel is vibe coding compared to vibe engineering.

    Context7 MCP for a Documentation Oracle

    This server maintains current React documentation. It caught things I’d missed:

    • defaultProps usage – deprecated as of React 18.3, still in Claude’s training data
    • State management patterns that work but aren’t optimal for concurrent rendering
    • Component composition that could make testing easier

    It checked my code against what React’s maintainers recommend now, not what was popular when GPT-4’s training data ended.

    Playwright MCP to see what Actually Works

    Playwright wrote and ran automated tests. They failed (surprise, surprise!)

    • Modal opened with Enter, couldn’t close without a mouse
    • Form validation was cosmetic – API endpoint hardcoded to return success
    • Portfolio scroll broke completely with keyboard navigation

    This is what “looks good” means without proper testing: works for me, using my mouse, on my device, the way I browse.

    Without those three MCP servers, I could’ve shipped it thinking “this looks great.”

    The Gotchas

    Security: Filesystem MCP can read/write anywhere you can. No centralised audits. You’re running code with filesystem access controlled by AI that makes mistakes.

    When to just use CLI: If you can do it in one bash command, do that. Don’t over-architect. Don’t waste an afternoon debugging MCP to check bundle sizes when npm run build takes ten seconds.

    Quality varies wildly: Chrome DevTools, Context7, Playwright are mature and maintained. Most servers in the registry are experiments or abandoned projects. No download counts, no quality signals.

    Not always the right tool: MCP might help, might just be debugging overhead. Important to always be figuring out when context matters enough to justify the setup.

    Where This Is Going

    Industry says 90% of enterprises by end of 2025. I’m sceptical, but the momentum is real.

    Not sure where to begin? Start Here

    1. Connect Claude Desktop to Filesystem server, analyse a project – any project. Doesn’t need to be code!
    2. Try Chrome DevTools MCP if you do frontend work
    3. Don’t build your own server yet – use existing ones first

    For technical definitions, I wrote an MCP glossary with metaphors.

    The Core Lesson

    Remember the boyfriend story: AI without context gives you statistical averages; replacement boyfriends or code that “looks good” but breaks for half your users.

    AI with context gives you specific solutions to specific problems.

    MCP is how we bridge that gap. Not perfectly, not magically, but practically. With configuration files and environment variables and the occasional need to restart everything.

    When I need to audit accessibility across a site or check for deprecated APIs I’ve never touched? Having tools that give LLM’s actual context instead of making it guess – that’s when it’s worth the setup.

    FAQ

    What’s the difference between vibe coding and vibe engineering?

    Vibe coding is using AI to generate code quickly based on feel — great for prototypes, but the output often has hidden quality problems. Vibe engineering means using AI as part of a rigorous workflow: you still audit, test, and verify. The MCP servers in this post (Chrome DevTools, Context7, Playwright) are what turn vibe coding into vibe engineering.

    Which MCP servers should I start with?

    Chrome DevTools MCP if you do any frontend work — it gives actionable, specific fixes rather than vague suggestions. Context7 if you want to make sure your code matches current library documentation rather than what was in your AI’s training data. Playwright MCP if you want to catch accessibility and interaction bugs that "looks good" won’t catch.

    Is MCP production-ready in 2025?

    Mostly. The core servers (Chrome DevTools, Context7, Playwright) are mature and maintained. The broader ecosystem is still Geocities-era — lots of experiments, abandoned projects, and no quality signals. Approach community servers with appropriate scepticism.


    Resources:

    I’m learning in public. If you spot where I’ve oversimplified or gotten something wrong, I want to know

  • A mostly-metaphoric MCP glossary

    A mostly-metaphoric MCP glossary

    When I first started learning about Model Context Protocol, I kept finding that there was no “intermediate” level. Every explanation assumed either I knew nothing, and that I should just like to know “What does MCP stand for?” or that I was an expert and already knew what everything meant. I’d read “the MCP server exposes tools that the client can invoke” and think… right, but what is a server in this context? Is it like a web server? And what makes something a “tool” versus a “resource”?

    I found myself toggling between documentation, blog posts, and example code, trying to piece together a mental model that made sense. Eventually I realised I needed to write down some explanations to help get everything to click.

    If you read my previous post about AI building terms and metaphors, you’ll know I’m a big believer in using multiple explanations to understand new concepts. Sometimes the technical metaphor lands, sometimes it’s the completely unrelated one that makes everything suddenly make sense.


    1. MCP SERVER

    TL;DR A programme that exposes tools, resources, or prompts that Claude (or other LLM clients) can use. The “backend” in the MCP architecture.

    What’s Important

    • Can be written in any language (Python, TypeScript most common)
    • Runs locally or remotely
    • Can expose multiple capabilities (tools + resources + prompts)
    • Discovered and connected to by MCP clients
    • Each server has a specific domain (filesystem, calendar, Linear, etc.)

    Unrelated Metaphor – Kitchen Stations

    You’re running a restaurant kitchen. Claude is the head chef who takes orders and coordinates everything. Each MCP server is a specialised station: the grill station (can cook meat), the salad station (can prep vegetables), the dessert station (can make sweets). When an order comes in, the head chef doesn’t cook everything. They delegate tasks to each station. “Grill station, I need a steak medium-rare!” The head chef knows what each station can do and coordinates them, but each station has its own tools and expertise.

    Developer Metaphor

    It’s like a microservice with a standardised API contract. Instead of building custom REST endpoints, you implement the MCP protocol (JSON-RPC 2.0 over stdio/HTTP). Your server registers “handlers” for different functions. Same concept as Express routes or RPC methods. The MCP client discovers your server’s capabilities at runtime (like OpenAPI/Swagger but for AI tools), then invokes your functions with typed parameters. Each server is a bounded context in DDD terms. Filesystem handles files, Linear handles issues, etc.


    2. MCP CLIENT

    TLDR;

    The application that connects to MCP servers and uses their capabilities. Claude.ai, Claude Desktop, and IDEs can be MCP clients.

    What’s Important

    • Manages connections to multiple servers
    • Handles authentication and permissions
    • Presents available tools to the LLM
    • Routes requests between LLM and servers
    • One client can connect to many servers

    Metaphor – General Contractor

    You’re renovating your house. The MCP client is your general contractor. You tell the contractor “I want a new kitchen with modern appliances.” The contractor maintains relationships with electricians (one server), plumbers (another server), cabinet makers (another server). When you make a request, the contractor figures out which specialists to call, coordinates their work, handles payments (authentication), and reports back to you. You don’t directly manage each specialist. The contractor does that.

    Developer Metaphor

    It’s like an API gateway combined with an orchestration layer. The client maintains a registry of connected services (servers), handles service discovery, manages authentication/authorisation for each service, and routes requests. When the LLM needs to call a function, the client: (1) determines which server handles it, (2) validates parameters against JSON Schema, (3) sends the request over the appropriate transport, (4) handles errors/retries, (5) returns results to LLM. It’s the conductor for a distributed system where the LLM is the orchestrator and servers are workers.


    3. TOOL (or FUNCTION)

    TLDR;

    An action the MCP server can perform. Functions Claude can call to do things in the external world.

    What’s Important

    • Defined with JSON Schema (parameters, types, descriptions)
    • Can modify external state (create files, send emails, etc.)
    • Returns results back to Claude
    • Can fail and return errors
    • Each tool has a clear purpose and parameters

    Unrelated Metaphor – Power Tools

    Imagine a carpenter’s workshop. Each MCP server is a workbench with specific power tools. The woodworking bench has: table saw (tool for cutting straight lines), router (tool for decorative edges), sander (tool for smoothing). When you need something built, the carpenter doesn’t just “work on wood”. They choose specific tools for specific tasks. “I need to cut this board” leads to using the table saw tool with parameters: length=24 inches, angle=45 degrees. Each tool does one thing well, takes specific inputs, and produces specific outputs.

    Developer Metaphor

    It’s exactly like function definitions with strict typing. Each tool is a function with a JSON Schema signature:

    typescript
    interface CreateFileTool {
      name: "create_file";
      parameters: {
        path: string;
        content: string;
      };
      returns: { success: boolean; error?: string };
    }

    The client validates parameters against the schema before invoking. The server executes the function and returns a typed result. It’s RPC with schema validation, like gRPC or tRPC, but designed for LLM consumption. The LLM sees these as “callable functions” and generates structured calls based on the schema.


    4. RESOURCE

    TLDR;

    Data that can be read from an MCP server. Like files, database records, or any content that can be retrieved.

    What’s Important

    • Read-only access to data
    • Identified by URI (like file:///path/to/file)
    • Can be text, images, or other media
    • Separate from tools (resources are passive data, tools are active functions)
    • Can be large (PDFs, images, long documents)

    Unrelated Metaphor – Museum Exhibits

    An MCP server is like a museum. Resources are the exhibits you can view: paintings, sculptures, artefacts. Each has a label (URI) like “Ancient Egypt Wing, Case 3, Artefact 42.” You can view any exhibit (read the resource), but you can’t modify them. They’re behind glass. Tools would be like the gift shop or café, places where you can do things (buy souvenirs, order food). Resources are static content you consume; tools are interactive actions you perform.

    Developer Metaphor

    Resources are like a read-only REST API or a file system mount. Each resource has a URI: resource://server-name/path/to/resource. You GET the resource (no POST/PUT/DELETE). The server implements handlers like:

    typescript

    async getResource(uri: string): Promise<{
      content: string | Uint8Array;
      mimeType: string;
    }>

    Think of it as a content delivery network where everything is immutable. The LLM can request resources to include in context, but resources don’t have side effects. It’s the separation between queries (resources) and commands (tools) in CQRS pattern.


    5. TRANSPORT

    TLDR;

    The communication layer between client and server. How messages are sent back and forth.

    What’s Important

    • stdio: Standard input/output (most common for local servers)
    • HTTP/SSE: For remote servers over network
    • Handles JSON-RPC 2.0 protocol messages
    • You usually don’t think about this directly
    • Abstracted away by MCP SDKs

    Unrelated Metaphor – Mail Delivery

    You write a letter to your friend (the message/request). Transport is how it gets delivered: you could hand-deliver it (stdio, fast, local only), use postal mail (HTTP, slower, works anywhere), or use a courier service (SSE, reliable, real-time updates). The content of your letter doesn’t change based on delivery method, only how it travels. Most people don’t care if their email uses SMTP or their texts use SMS. They just want messages delivered. Similarly, developers usually don’t think about MCP transport; the SDK handles it.

    Developer Metaphor

    It’s the OSI transport layer for MCP. stdio is like Unix pipes or IPC: low-latency, local process communication using stdin/stdout. HTTP/SSE is like REST over network: higher latency but works remotely. Both carry JSON-RPC 2.0 payloads (the application layer). Similar to how gRPC can use different transports (HTTP/2, Unix sockets), MCP can use stdio or HTTP. The SDK provides abstractions:

    typescript

    const transport = isLocal 
      ? new StdioTransport(command, args)
      : new HttpTransport(url);

    You rarely implement transport yourself. It’s provided infrastructure.


    6. CAPABILITY

    TLDR - Need to Know

    Feature flags indicating what an MCP server supports. Not all servers support all features.

    What’s Important

    • Tools capability: can provide callable functions
    • Resources capability: can provide readable data
    • Prompts capability: can provide prompt templates
    • Sampling capability: can request LLM completions (advanced)
    • Negotiated during connection handshake

    Unrelated Metaphor – Restaurant Menu Sections

    When you sit down at a restaurant, the menu shows capabilities: [Appetisers] [Entrees] [Desserts] [Bar]. Not every restaurant has every section. A breakfast diner might not have [Bar], a café might not have [Desserts]. Before ordering, you check what sections exist. You don’t ask a breakfast diner for cocktails because they don’t have that capability. Similarly, you don’t ask a read-only documentation server to create files (no tools capability). You only request documents (resources capability). The menu tells you what’s possible before you order.

    Developer Metaphor

    It’s like feature flags or interface implementation checking. During the initialisation handshake, server advertises capabilities:

    typescript

    interface ServerCapabilities {
      tools?: { listChanged?: boolean };
      resources?: { subscribe?: boolean; listChanged?: boolean };
      prompts?: { listChanged?: boolean };
      sampling?: {};
    }

    The client checks if (server.capabilities.tools) before trying to call tools. Similar to capability negotiation in HTTP (Accept headers) or feature detection in browsers (if ('geolocation' in navigator)). Prevents “method not supported” errors by advertising capabilities upfront. It’s the Interface Segregation Principle: servers only implement what they need.


    7. SAMPLING

    TLDR - Need to Know

    When an MCP server can request LLM completions from the client. Lets servers use AI to process data.

    What’s Important

    • Advanced feature (most servers don’t use this)
    • Server can ask client “hey, can you have Claude analyse this?”
    • Enables AI-powered tools that need LLM reasoning
    • Requires explicit permission from user
    • Server sends prompt, client returns LLM response

    Unrelated Metaphor – Sous Chef Asking Head Chef

    Normally the head chef (Claude) tells the sous chefs (servers) what to cook. But sometimes a sous chef needs culinary expertise: “Chef, I found this mystery ingredient. What is it and how should I use it?” The sous chef asks the head chef for their expert opinion, then uses that advice to complete their work. Sampling is when a specialist (server) consults the expert (LLM) for analysis before continuing their task. Most specialists don’t need this. The grill station knows how to cook steak. But occasionally, complex situations require the head chef’s input.

    Developer Metaphor

    It’s like a worker service calling back to the orchestrator for AI assistance. Normally: Client to Server (tool invocation). With sampling: Server to Client to LLM to Client to Server (callback pattern). The server makes a “give me a completion” request:

    typescript

    const analysis = await client.sampling.createMessage({
      messages: [{ role: "user", content: "Analyse this code..." }],
      maxTokens: 1000
    });

    It’s like a microservice making an RPC call back to a central AI service. Useful for servers that need AI reasoning (e.g., code analysis, content understanding) but don’t want to run their own LLM. The client controls costs/permissions. Servers request, clients approve.


    8. CONTEXT (or Arguments/Parameters)

    TLDR – Need to Know

    Data passed when calling tools or accessing resources. The inputs to MCP functions.

    What’s Important

    • Defined by JSON Schema for each tool
    • Type validation enforced
    • Can be simple (strings, numbers) or complex (nested objects)
    • Tool execution fails if context doesn’t match schema
    • Each tool specifies required vs optional parameters

    Unrelated Metaphor – Coffee Order

    When you order coffee, you provide context/parameters: drink type (required: “latte”), size (required: “grande”), milk (optional: “oat milk”), extras (optional: “extra shot”). The barista can’t make your drink without the required parameters. If you just say “coffee please,” they ask questions to get the missing context. Some parameters have defaults (regular milk if you don’t specify). Context is all the specific details needed to fulfil your request. The menu shows what parameters each drink needs: some required, some optional, some with defaults.

    Developer Metaphor

    It’s literally function parameters with JSON Schema validation:

    typescript

    type CreateFileParams = {
      path: string;           // required
      content: string;        // required
      encoding?: string;      // optional, default: 'utf-8'
    };

    Before the server function executes, the client validates params against schema (like TypeScript compile-time checking or Joi runtime validation). Similar to gRPC message definitions or OpenAPI parameter specs. The schema defines:

    • Parameter names and types
    • Which are required vs optional
    • Defaults and constraints
    • Descriptions for LLM understanding

    The LLM generates structured calls matching these schemas.


    9. CONFIGURATION

    TLDR – Need to Know

    Settings that tell the MCP client which servers to connect to and how to authenticate.

    What’s Important

    • Usually in claude_desktop_config.json or similar
    • Specifies server command to run or URL to connect to
    • Can include environment variables (API keys, etc.)
    • Per-server settings (allowed directories, permissions)
    • Client reads this on startup to initialise servers

    Unrelated Metaphor – Emergency Contact Card

    You have a card in your wallet with emergency contacts: Mum (call: 555-1234), Doctor (call: 555-5678, insurance ID: XYZ123), Lawyer (email: lawyer@firm.com, case number: 456). This card tells you who to contact for what situation and what information they need. MCP configuration is the same. It’s your assistant’s contact card for all the specialists. It lists: who they are, how to reach them, what credentials to use, and what they’re allowed to do. Update the card when contacts change. Your assistant can’t work without this card. They don’t know who to call.

    Developer Metaphor

    It’s like a docker-compose.yml or serverless.yml: infrastructure-as-code for MCP services:

    json

    {
      "mcpServers": {
        "filesystem": {
          "command": "npx",
          "args": ["-y", "@modelcontextprotocol/server-filesystem", "/allowed/path"],
          "env": { "DEBUG": "mcp:*" }
        },
        "linear": {
          "command": "npx",
          "args": ["-y", "@linear/mcp-server"],
          "env": { "LINEAR_API_KEY": "${LINEAR_API_KEY}" }
        }
      }
    }

    The client reads this on startup, spawns processes (stdio) or connects to URLs (HTTP), passes env vars, handles authentication. It’s service configuration, similar to Kubernetes manifests or systemd unit files. Declarative specification of what services to run and how.


    Where to Go From Here

    I’m currently building custom MCP servers for some of my own projects, and I expect my understanding will continue evolving. If you spot something I’ve explained unclearly (or got wrong), I’d genuinely love to hear about it. This is a living document that I’ll update as I learn more.

    FAQ

    What’s the difference between an MCP server and an MCP client?

    The server exposes capabilities — tools, resources, prompts — that an AI can use. The client (like Claude Desktop) connects to those servers, handles auth, and shuttles messages between the AI and the right server. One client can happily talk to lots of servers at once.

    Do I need to know Python or TypeScript to use MCP?

    Not to use it. You can point Claude Desktop at existing servers with a config file and never touch code. If you want to build your own server, TypeScript and Python have the most mature SDKs right now, but the protocol itself doesn't care what language you use.

    What’s the simplest way to get started with MCP?

    Connect Claude Desktop to the Filesystem MCP server. Point it at a project folder, then ask Claude to analyse the code, docs, or notes inside. You'll immediately feel the difference between the model guessing and the model actually knowing what's in your files.

    Is MCP only for developers?

    Right now it's mostly developer-shaped. Setting up servers still means editing config files and being comfortable in a terminal. But as more tools ship MCP support out of the box, you'll be able to benefit from it without having to fiddle with all the plumbing yourself.

    The best way to really understand MCP is to build something with it. Start small. Maybe connect Claude Desktop to an existing server like the Filesystem or Linear servers. Play around with what they can do. Then try building a simple read-only server that exposes some data you care about. The concepts that feel abstract now will suddenly click into place once you’re working with actual code.

    And if you’re building something interesting with MCP, I’d love to hear about it. We’re all figuring this out together.

  • From Geocities to GPT

    From Geocities to GPT

    I’ve been trying to find my first website for ages now. The McPhee Family Pets, circa 1999.

    Hours down rabbit holes of the Wayback Machine, trying every possible URL combination I can think of. Was it heartland/prairie? EnchantedForest? Did I use underscores or hyphens? The internet has swallowed it whole, along with Porygon’s Cave and whatever I called Horsea’s page (Horsea’s Haven? Horsea’s Hideout? The name floats just out of reach).

    There’s something devastating about losing these first creative digital expressions. Like they existed in some parallel internet that’s been paved over. I was ten years old with a chinchilla, pet rats, mice, chickens – the list was genuinely ridiculous – and I believed each one deserved their own dedicated webpage. Their own corner of the internet, pure “here is my rabbit named Libby and here are three facts about her” energy.

    I remember the old web though. Not all of it, but fragments. Like how our dial-up plan gave us an allowance for New Zealand hosted websites versus international ones, so I’d browse locally hosted tutorials for hours. There was one about frames that explained them like a dinner plate – your main content (meat) in the middle, navigation (salad) on the side, maybe a footer (dessert) down the bottom. I thought this was the most brilliant metaphor at the time. I probably spent weeks just moving frame borders around, watching content reflow, feeling like an architect.

    Then there was Vikimouse and the MousePad Kids – this website where you could adopt virtual mice that lived in elaborately crafted pixel houses. Someone called Vikimouse had made each one pixel by pixel. I’d stare at them, trying to understand how someone had that much patience. How they knew which pixel should be brown and which should be tan to make it look like wood grain. I’d view source on everything, trying to decode the magic. And then I would populate my home page with entire families of adopted, digital, pixel-art rodents.

    The platform wandering started early. Geocities, Tripod, Bravepages, Angelfire – I was chasing free. Zero dollar budget, minimal ads, maximum creative control. Each platform migration was like trying on a different digital identity. Would THIS be the place where my Pokemon fan sites would finally look professional? (They never did. But they had auto-playing MIDI files and that’s what really mattered.)

    I joined a forum called Young Coders, or something close to that. We’d share JavaScript snippets we’d found – mouse trails, falling snow, those eyes that followed your cursor around the page. Copy, paste, pray it worked. When it did, you felt like you’d just cast an actual spell. When it didn’t, you’d spend hours hunting for the missing semicolon, not knowing that twenty-five years later you’d still be hunting for missing semicolons, just now with better error messages.

    As I grew into a teenager, things got more sophisticated. Or at least, I thought they did. Dreamweaver felt like cheating after hand-coding everything. Macromedia Flash was pure magic – suddenly things could MOVE. Not just blink tags and marquees, but actual animation. I made band fan pages with the dedication of a digital shrine builder. Learned some ASP.NET because it sounded important and grown-up, and eventually grew into PHP, where I got my first taste of the pre-WordPress bbPress.

    The webrings were their own special kind of commitment. You’d apply to join one – “Anna’s Pokemon Paradise is applying to join the Elite Water Pokemon Webring” – and wait anxiously for approval. Then you’d get this chunk of HTML to add to your site with Previous and Next buttons, making you part of this infinite loop of similarly obsessed people. I was probably in twelve different rings at one point. Pokemon ones, virtual pet ones, one for just about anything.

    Guestbooks were mandatory. You weren’t a real website without a guestbook. Mine was from Bravenet, plastered with whatever background GIF I thought was sophisticated that week. The entries were always the same – “Cool site!” “Love the pics!” and occasionally someone would actually write something substantial and you’d feel like you’d made it. Like your website was a real place people visited, not just pixels you were shouting into the void.

    Then came the band fan pages. I had opinions and they needed dedicated web spaces. Frames for everything. Left frame: navigation with each band member’s name in a different font, Top frame: band logo I’d painstakingly cut out of a larger image in Paint Shop Pro. Main frame: “News” that I’d copied from other fan sites, maybe a gallery of images that took seventeen years to load on dial-up. I probably had a disclaimer somewhere about not owning the images, as if Sony Music was going to come after a fourteen-year-old in New Zealand.

    I remember the exact moment I discovered Google in beta. I was a catalogue of search engines and web directory listings before that (I don’t say catalogue metaphorically, I kid you not – I had a clearfile folder where I would write down the URL of every search engine and directory I could find) But Google was just… empty. A logo, a search box, two buttons. It felt wrong, like someone had forgotten to finish building it. Where were all the portal features? The weather? The news? But then you searched for something and it actually found what you wanted. Not seventeen pages of garbage with your result buried on page twelve. It was unsettling how good it was.

    CSS Zen Garden broke my brain entirely. This was maybe 2003? I’d been tables-for-layout loyal, defending my nested tables like they were a personal religion. Then someone showed me CSS Zen Garden – the exact same HTML, completely transformed just by changing the stylesheet. I spent hours viewing source, trying to understand how the garden became the ocean became the subway map. It was like finding out you’d been painting with your fingers while everyone else had brushes.

    I think I tried to recreate every single design. Failed spectacularly. But in that failure, I started to understand the cascade, specificity, the box model (though IE6 would torture us with that for years to come). Started to grasp that we were trying to teach browsers our intent. That HTML was supposed to be structure, CSS was presentation, and mixing them was… wrong somehow? Though I definitely kept using inline styles for “just this one quick thing” for an embarrassingly long time after.

    We spent the next two decades getting really good at this conversation with browsers. Teaching them to understand that when we said “display: flex” we meant “please for the love of god just center this div.” Learning their quirks – Safari would do this, Chrome would do that, and IE… well, IE would do whatever it felt like. We learned to speak their language, to think in their logic.

    And now here we are, trying to teach AI to understand context, and it’s like being ten years old staring at Vikimouse’s pixel art again. We know there’s something magical here, something transformative. But we’re still copy-pasting code snippets and praying they work. Still hunting for missing semicolons, just now they’re in JSON configs for MCP servers instead of JavaScript mouse trails.

    The thing is, AI doesn’t understand context the way we learned to understand the cascade.

    When I’m debugging why an MCP server won’t talk to my tools properly, it feels exactly like debugging why my frames wouldn’t resize in Netscape Navigator. Except now instead of teaching a browser that “frameborder=’0′” means “please don’t draw that ugly gray line,” I’m teaching Claude that when I say “search my previous conversations about MCP” I mean actual conversations, not some hallucinated memory of conversations that never happened.

    I’ve been experimenting with MCP servers for a little while now, and it’s giving me the same feeling as those early days of copying JavaScript snippets. You know something powerful is happening, but you’re not entirely sure why it works when it works. Just last week I spent three hours trying to figure out why my context wasn’t passing through properly, only to discover I had the wrong quotation marks. Not missing ones – the wrong kind. Curly quotes instead of straight ones. In 1999, it was forgetting to close a font tag. In 2025, it’s Unicode characters that look identical but aren’t.

    The documentation situation feels familiar too. Back then, you’d have seventeen browser tabs open (once we got tabs – remember when opening a new site meant opening a whole new window?), each with a different tutorial that explained things slightly differently. Now I have seventeen tabs of Anthropic docs, GitHub repos, and Discord conversations where someone’s figured out something that isn’t documented anywhere yet. We’re all still collectively teaching each other, just now it’s in Slack threads instead of Young Coders forums.

    But here’s what’s making me think: We got really good at teaching browsers to understand us. It took twenty-five years, but we did it. We went from table-based layouts and spacer GIFs to CSS Grid and container queries. From “best viewed in Internet Explorer 5” badges to responsive designs that work on everything from a watch to a wall-mounted TV (and perhaps even your fridge!)

    What I’m wondering is – what will teaching AI look like in twenty-five years? Right now, we’re in the Geocities era of AI interaction. We’re copy-pasting prompts like we used to copy-paste JavaScript snow effects. We’re joining the AI equivalent of webrings – Discord servers and GitHub repos where people share their successful MCP configurations. We’re building the 2025 equivalent of “The McPhee Family Pets” – earnest, ambitious projects that probably won’t exist in their current form in five years, let alone twenty-five.

    I found a screenshot the other day of a website I made in 2001. It had a splash page. Remember splash pages? “Click here to enter” with some elaborate Flash animation that everyone immediately clicked through. It seemed so important at the time – the grand entrance to your digital space. Now it’s almost embarrassing to look at. What will we think of our current AI interactions in 2049? Will we laugh at how we used to manually configure context windows? Will prompt engineering seem as quaint as table-based layouts?

    Sometimes I wonder if those lost websites – The McPhee Family Pets, Porygon’s Cave, Horsea’s whatever-it-was – are better off disappeared. They exist now exactly as they should: perfect in memory, terrible in reality. They were never about being good websites. They were about that feeling when your HTML finally worked, when your frame borders aligned, when someone actually signed your guestbook.

    That’s what I’m chasing now with MCP servers and AI tools. Not the perfect implementation, but that moment when something clicks into place. When the context passes through correctly and suddenly your tool can see your previous conversations. When the AI understands not just what you’re saying but what you mean. It’s the same magic, just with better error messages and worse documentation.

    We spent decades teaching browsers to understand our intent. Now we’re teaching AI. The difference is, this time I’m not ten years old with unlimited time and a chinchilla. I’m thirty-nine with a toddler, a full-time job, and approximately seventeen minutes of free time per day. But I still get that same feeling when something finally works. That same urge to view source on everything, to understand the magic.

    Don’t judge – we all started somewhere. And honestly? We’re all starting somewhere again.

    The web I grew up with is gone – not just my websites, but that whole version of the internet where teenagers could build shrines to their pets and their favourite bands without thinking about SEO, TikTok videos, engagement metrics or whether an AI could do it better. But maybe that’s okay. Maybe each generation gets their own version of the web to figure out, to break, to build weird things on.

    I just hope somewhere out there, some ten-year-old is building the AI equivalent of The McPhee Family Pets. Teaching GPT about their pet chickens. Making something wonderfully terrible that they’ll try to find in twenty-five years and fail.

    That’s the web I want to help build.

  • Three AI customisation concepts

    Three AI customisation concepts

    I think in metaphors. It’s how I understand anything complex – by finding connections to something I already know, preferably from a completely different domain. When I first started learning about AI concepts, I’d read the formal definitions and feel like I was staring at a wall of text. But then someone would say “it’s like…” and everything would click into place.

    So I started collecting these. Every time a concept finally made sense, I’d write down what comparison made it work for me. And here’s what I noticed: I usually needed at least two different metaphors before something truly landed. The baking metaphor would get me 60% of the way there, then the developer metaphor would fill in the rest. Or vice versa.

    I’ve narrowed this down to three concepts that feel foundational – the ones I keep coming back to when I’m trying to understand how AI systems actually work, or when I’m explaining MCP to someone, or when I’m making decisions about how to build something. These three form a kind of progression: understanding how AI represents meaning, then how you customize it, then how you make customization practical.

    The patterns between these concepts interest me as much as the concepts themselves. How embeddings enable RAG, how LoRA makes fine-tuning accessible, how choosing between RAG and fine-tuning depends on whether you’re teaching facts or behavior. These connections make the whole landscape easier to navigate.


    Embedding

    TL;DR Converting text, images, or audio into dense numerical vectors (arrays of numbers) where similar meanings equal similar numbers. This is the foundation of semantic search and RAG.

    What matters most: Embeddings capture meaning, not just keywords – “car” and “automobile” end up with similar embeddings. They’re fixed-size representations, so “hi” and a 500-word essay both become the same length vector. They’re used for similarity search – you compare vectors to find related content. Different models produce different dimensions – text-embedding-3-small gives you 1,536 dimensions. And there are cost-quality trade-offs to consider: smaller embeddings are cheaper but less nuanced.

    The colour version: Every colour can be described with exactly three numbers: RGB values. Red equals [255, 0, 0], orange equals [255, 165, 0]. Colours that look similar have similar numbers. You can find “similar colours” by comparing these triplets. Embeddings do the same for text, except instead of three numbers describing colour, you use 1,536 numbers describing meaning. Similar meanings equal similar numbers, just like similar colours equal similar RGB values.

    The developer version: It’s like hashing, but preserving similarity instead of randomising it. A hash function converts “hello” into something like 2cf24dba5fb0a30e. Embeddings convert “hello” into [0.1, 0.3, -0.2, …]. But unlike hashing where similar inputs give totally different hashes, embeddings make similar inputs give similar vectors. It’s a “similarity-preserving hash” – you can compare the results to find related content. Use cosine similarity instead of equality checks.

    Diagram showing how text concepts are mapped into high-dimensional embedding vectors, with similar meanings clustered close together in space

    Fine-Tuning vs RAG

    TL;DR Two different approaches to customisation. RAG equals retrieve external knowledge and include it in the prompt. Fine-tuning equals retrain model weights on your data. Use RAG for facts, fine-tuning for behaviour.

    What matters most: Use RAG for changing knowledge – prices, inventory, news, data that updates frequently. Use fine-tuning for changing behaviour – style, format, tone, domain terminology. RAG is cheaper and faster because there’s no retraining needed, you just update documents. Fine-tuning has no retrieval overhead because everything’s baked into the weights. Often they work best together – fine-tune for style, RAG for facts.

    The restaurant version: RAG is a chef with access to an ingredient database. Every time someone orders, they look up what ingredients to use, then cook. If ingredient prices change, the database updates automatically. Fine-tuning is training the chef in a specific cuisine. They’ve practised Italian cooking so much, they automatically know techniques, flavour combinations, and traditions without looking up recipes. Use RAG when ingredients (facts) change daily. Fine-tune when you need consistent technique (style). Best restaurants do both: trained chefs (fine-tuned) with access to fresh ingredient databases (RAG).

    The developer version: RAG is dependency injection – inject external data at runtime via function parameters. Fine-tuning is monkey-patching the library itself – you’re modifying the internal behaviour. RAG equals generateResponse(prompt, retrievedContext) where you control context per request. Fine-tuning equals recompiling the function with different internal logic. Use RAG when your data changes frequently (like pulling from an API). Fine-tune when you need to change how the function processes things (like changing validation logic). Often you do both: custom business logic (fine-tuned) that queries your database (RAG).

    Diagram comparing RAG and fine-tuning, contrasting runtime context injection with retraining model weights on custom data

    LoRA (Low-Rank Adaptation)

    TL;DR A parameter-efficient fine-tuning method. It freezes the base model and trains small “adapter” matrices instead. This reduces trainable parameters by 99%-plus. It enables fine-tuning on consumer GPUs.

    What matters most: LoRA trains less than 1% of parameters, which dramatically reduces memory requirements. It produces small files – 20MB adapter versus 14GB full model. You can merge or swap adapters – keep the base model, swap adapters for different tasks. Quality is close to full fine-tuning, surprisingly effective despite fewer parameters. And it makes fine-tuning accessible, running on a 24GB GPU instead of needing clusters.

    The lens filters version: The base model is a camera lens (14GB). Full fine-tuning is grinding and re-polishing the entire lens – expensive and permanent. LoRA is screwing on a small filter (20MB) that changes how light passes through. The filter is 99% smaller than the lens, but it effectively modifies the output. You can keep one lens and swap filters: sepia filter, polarising filter, UV filter. LoRA adds a small mathematical “filter” that transforms the model’s computations without changing the underlying “lens” (weights).

    The developer version: It’s like the decorator pattern instead of inheritance. Full fine-tuning equals a subclass that overrides every method. LoRA equals a wrapper that intercepts some method calls and tweaks outputs. When the model computes output = input @ weights, LoRA intercepts: output = input @ (weights + A @ B) where A and B are tiny matrices. The original weights stay frozen. You’re training A and B (1% of the size) instead of retraining all weights. It’s monkey-patching with small patches instead of forking the entire codebase.

    Diagram illustrating LoRA, with small adapter matrices A and B added alongside frozen base model weights

    FAQ

    When should I use RAG instead of fine-tuning?

    Use RAG when your data changes frequently — things like prices, inventory, or news. Use fine-tuning when you need to change how the model behaves — its style, tone, or domain-specific reasoning. When you’re not sure, RAG is usually cheaper and faster to iterate on.

    What is LoRA and why does it matter for fine-tuning?

    LoRA (Low-Rank Adaptation) lets you fine-tune a model by training only a tiny slice of its parameters, using small "adapter" matrices instead of modifying all the weights. The upshot: you can fine-tune on a consumer GPU instead of needing a cluster, and the adapter file might be around 20MB instead of something like 14GB.

    Do embeddings work across different languages?

    Multilingual embedding models can represent meaning across languages in the same vector space — so "car" in English and "voiture" in French end up with very similar embeddings. Standard English-only models won’t do this reliably, so you need a model that’s explicitly trained to be multilingual.

    Here’s what I’m still figuring out: whether these explanations actually work for people who aren’t developers. Some concepts might still feel abstract no matter how many metaphors I throw at them.

    I’m also wondering if I’ve oversimplified some of these. The fine-tuning comparison, for instance, glosses over when you might genuinely need full fine-tuning instead of LoRA, or the cases where RAG and fine-tuning aren’t alternatives at all but complementary. But I think that’s okay – this is meant to build intuition, not replace documentation.

    What I’ve learned from writing this is that the same concept really does need multiple entry points. My team would understand things best through the developer metaphors. A product manager or designer might prefer the non-technical versions. And I find myself using different metaphors depending on my own mental state – sometimes the abstract database comparisons help, sometimes I need the cake-baking version.

    The progression here matters too. You can’t really understand RAG without first understanding embeddings – how else would you search for similar content? And you can’t appreciate LoRA without understanding that fine-tuning exists but is expensive. These three concepts build on each other in a way that mirrors how you’d actually approach customising an AI system: understand how it represents meaning, decide whether you need to change knowledge or behaviour, then learn there’s a practical way to change behaviour without needing a GPU cluster.

    If you found this useful, I’d be curious which metaphors actually worked for you. And which concepts still feel fuzzy – those are probably the ones I haven’t fully understood myself yet, despite writing explanations for them.

  • AI Replaced My Boyfriend: A Story About Context in Machine Learning

    AI Replaced My Boyfriend: A Story About Context in Machine Learning

    I was trying to get a professional headshot the other day. You know how it is. I needed something recent, and all my decent photos are either me in the garden covered in dirt or family shots with my daughter.

    There was this one photo from Mexico City I really liked. Good lighting, I actually looked awake, my hair was doing what it was supposed to. Only problem? Aldo had his arm around me. Not exactly the solo professional headshot I needed.

    Anna and her partner Aldo in Mexico City, the photo she wanted to turn into a professional headshot

    So I turned to Photoshop’s new AI features. Simple request, I thought. Remove the person next to me, give me a neutral professional background. What could go wrong?

    The AI looked at my photo, understood the assignment, and promptly… gave me a new boyfriend.

    Not a modified Aldo. Not an empty space where Aldo used to be. A completely different Hispanic-looking man, arm still around me, same intimate couple pose. The AI had racially profiled my actual partner just enough to select an appropriate replacement from its training data. Like it was saying, “Based on the statistical probability of who this woman would have her arm around, let me provide you with Hispanic Male, Option B.”

    I laughed until I nearly cried. Then I tried again.

    Second attempt? The AI removed Aldo successfully this time, but decided I needed a jungle background and put a cocktail in my hand. Because nothing says “professional headshot” like sipping a mojito in the rainforest, apparently.

    This whole experience reminded me of another trend that swept through LinkedIn a while back: asking ChatGPT to create an image of you based on what it knows from your conversations. I tried it, curious what patterns the AI had picked up about me.

    First result: I was a white man with a beard, slight smile. The classic “software developer” stereotype from every stock photo ever taken.

    I tried again a few weeks later, after they’d clearly done some diversity training on their models. This time? I was a Black woman with natural hair and glasses, standing on a generic city street.

    The overcorrection was almost funnier than the original bias. Like watching someone try so hard not to be racist that they circle back around to being weird about race in a completely different way.

    Here’s what fascinates me about all this: these systems are doing exactly what they’re trained to do. When my arm was positioned like it was around someone, the AI couldn’t comprehend that I wanted that someone to not exist. Its training data says arms in that position belong around people. So it provided a person.

    When asked to imagine what a software developer named Anna looks like, it ping-ponged between “definitely a white man” and “we’ve been told to increase diversity, so definitely not a white man”, never quite landing on anything close, despite a lengthy chat and project history which knows so much about me.

    The problem isn’t that these models are broken. They’re pattern-matching perfectly against their training data. The problem is they have no context for what we actually want or who we actually are. They’re making statistical guesses based on millions of images and conversations that may or may not represent reality, and definitely don’t represent individual reality.

    This is exactly why I’ve been obsessed with Model Context Protocol (MCP) lately. It’s attempting to solve this exact problem: how do we give AI systems the context they need to understand not just the statistical average, but the specific situation? How do we move from “woman with Hispanic partner probably wants another Hispanic man in photo” to “this particular person wants this particular other person removed from this particular photo”?

    Context isn’t just about providing more information. It’s about providing the right information at the right time. It’s the difference between an AI that replaces your boyfriend with a statistical probability and one that understands what you’re actually trying to achieve.

    Third time was the charm, by the way. Finally got that professional headshot with a normal background. No replacement boyfriends, no tropical cocktails. Just me, looking professionally adequate, ready for my speaker bio.

    Though I’m keeping the jungle cocktail version. You never know when you’ll need a professional photo that says “I debug JavaScript from the rainforest.”


    I’ll be diving deeper into MCP and how context shapes AI behavior at Web Directions Developer Summit in Sydney this November. If you’re curious about the technical side of why AI keeps making these hilarious (and sometimes concerning) assumptions, come find me there.

    ps. what do you think of my headshot now?


    Anna McPhee speaker banner for Web Directions Developer Summit

    FAQ

    Why did AI give me a different boyfriend instead of removing him?

    AI image models are trained on statistical patterns, not intent. When your arm position matched “couple pose” in training data, the model filled the gap with the most statistically likely person — rather than understanding you wanted an empty space.

    What is Model Context Protocol (MCP) and why does it help?

    MCP is an open standard that lets AI systems receive structured context about a specific situation, rather than guessing from statistical averages. It’s the difference between AI that replaces your partner with “Hispanic Male, Option B” and AI that understands exactly what you’re trying to achieve.

    Does AI have racial bias in image generation?

    Yes — current models reflect biases in their training data. The pattern of replacing Aldo with a “statistically probable” partner, and generating a white male developer as a default, are both examples of bias baked into training data rather than intentional design.

  • Are we underestimating junior developers?

    Are we underestimating junior developers?

    I keep seeing these LinkedIn posts about how we shouldn’t treat AI as our junior developer, how we still need to hire juniors so we have seniors when the current generation retires. And while I agree we absolutely need junior developers, something about this framing sits wrong with me.

    It feels like we’re defending junior developers by essentially saying “we need them as human code-writing machines for the future.” But that’s not what made my junior years valuable, and it’s not what made the mentors who supported me so important.

    When I think about the people who helped me grow as a developer, they weren’t just teaching me to write code. They were showing me engineering and architectural thinking, how to do competitive analysis, how to gain stakeholder buy-in, how to conduct meaningful code reviews, how to debug complex systems. They walked me through the entire product lifecycle – from competitive analysis and planning, to building and iterating, to testing, reviewing, and releasing.

    If we’re only hiring junior developers to write the kind of code that AI can now generate, then we’re not treating them right in the first place.

    This reminds me of something that happened in banking decades ago. When ATMs were introduced in the 1970s, everyone assumed bank tellers would disappear. The machines could handle the most common teller tasks – dispensing cash, taking deposits. By the mid-1990s, over 400,000 ATMs were installed across the United States.

    But here’s what actually happened: the number of bank teller jobs didn’t decrease. Instead, ATMs made it cheaper to operate bank branches, so banks opened more branches. And the tellers who remained? Their role evolved completely. Cash handling became less important, and human interaction became more valuable. Tellers became part of the “relationship banking team,” helping customers with complex needs that machines couldn’t handle, selling financial products, building personal connections with small business customers.

    The skills of the job fundamentally changed. The technology didn’t eliminate the role – it freed tellers from routine tasks so they could focus on work that required human insight, relationship building, and complex problem-solving.

    I think we’re facing a similar moment with junior developers. AI can generate code, debug basic issues, even write tests. But if that’s all we expected junior developers to do, then we were underutilizing them from the start.

    The junior developers I want to work with aren’t just code writers – they’re learning to think like engineers. They’re asking questions that make our processes better. They’re developing the judgment to know when a technical solution fits the business context. They’re building the communication skills to explain complex problems to stakeholders. They’re learning to see the bigger picture of how their code fits into a product, a user experience, a business strategy.

    These skills don’t get automated away. If anything, they become more valuable when routine coding tasks are handled by AI.

    But here’s the thing – we have to be intentional about this. We can’t just hire junior developers and hope they’ll magically develop these skills while AI handles the “easy” work. We need to actively mentor them in systems thinking, product sense, architectural decisions, user empathy, and strategic problem-solving.

    The conversation shouldn’t be “AI versus junior developers.” It should be “how do we redefine what junior developers learn and do in an AI-augmented world?”

    Just like bank tellers evolved from cash handlers to relationship builders, junior developers can evolve from code writers to product thinkers, system designers, and strategic contributors. But only if we’re intentional about what we’re teaching them and what problems we’re asking them to solve.