<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en"><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://erikzocher.dev/feed.xml" rel="self" type="application/atom+xml" /><link href="https://erikzocher.dev/" rel="alternate" type="text/html" hreflang="en" /><updated>2026-08-28T14:49:24+02:00</updated><id>https://erikzocher.dev/feed.xml</id><title type="html">Erik Zocher</title><subtitle>Personal blog - Design, Technology, Berlin &amp; more</subtitle><author><name>Erik Zocher</name></author><entry><title type="html">I Bought a DGX Spark. Here’s What It Actually Is.</title><link href="https://erikzocher.dev/technology/2026/08/25/i-bought-a-dgx-spark.html" rel="alternate" type="text/html" title="I Bought a DGX Spark. Here’s What It Actually Is." /><published>2026-08-25T07:00:00+02:00</published><updated>2026-08-25T07:00:00+02:00</updated><id>https://erikzocher.dev/technology/2026/08/25/i-bought-a-dgx-spark</id><content type="html" xml:base="https://erikzocher.dev/technology/2026/08/25/i-bought-a-dgx-spark.html"><![CDATA[<h1 id="i-bought-a-dgx-spark-heres-what-it-actually-is">I Bought a DGX Spark. Here’s What It Actually Is.</h1>

<p>The package was smaller than I expected. A black box roughly the size of a Mac mini arrived, and inside was a machine that Nvidia calls “the world’s smallest AI supercomputer.” It cost about €4,000. And it can run a 284-billion-parameter language model entirely offline.</p>

<p>This is the first post in a series about living with a DGX Spark-class machine — the ASUS Ascent GX10. I’ll cover setup, running models, and what actually works. This one is the “what is this thing” post.</p>

<h2 id="the-chip-gb10-grace-blackwell">The chip: GB10 Grace Blackwell</h2>

<p>Every DGX Spark-class machine runs the same chip: Nvidia’s <strong>GB10 Grace Blackwell Superchip</strong>.</p>

<ul>
  <li><strong>20 ARM CPU cores</strong> (10× Cortex-X925 + 10× Cortex-A725) — this is a phone-style chip, not an x86 desktop CPU</li>
  <li><strong>128 GB LPDDR5x unified memory</strong> at 273 GB/s</li>
  <li><strong>Blackwell GPU</strong> rated at 1,000 TOPS (FP4)</li>
  <li><strong>~1.2 kg</strong>, about 25 W at idle</li>
  <li>Runs <strong>Linux only</strong> (DGX OS, Ubuntu-based)</li>
</ul>

<p>The key spec is the <strong>128 GB unified memory</strong>. That’s the whole product. A desktop GPU gives you 8-24 GB of VRAM; this gives you 128 GB that the CPU and GPU share. It’s not fast memory by workstation standards (273 GB/s vs. a 5090’s 1.8 TB/s), but it is <em>huge</em> — and size matters more than speed for running large models.</p>

<h2 id="what-128-gb-actually-unlocks">What 128 GB actually unlocks</h2>

<table>
  <thead>
    <tr>
      <th>Task</th>
      <th>Typical desktop (8-24 GB VRAM)</th>
      <th>DGX Spark (128 GB)</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>7-13B chat models</td>
      <td>✅</td>
      <td>✅ easily</td>
    </tr>
    <tr>
      <td>30B MoE models</td>
      <td>⚠️ tight</td>
      <td>✅ fast</td>
    </tr>
    <tr>
      <td>70B+ models</td>
      <td>❌</td>
      <td>✅</td>
    </tr>
    <tr>
      <td>120-284B MoE (DeepSeek V4 Flash)</td>
      <td>❌</td>
      <td>✅ (quantized)</td>
    </tr>
    <tr>
      <td>1M-token context</td>
      <td>❌</td>
      <td>✅</td>
    </tr>
    <tr>
      <td>FLUX/Wan video generation</td>
      <td>❌ (video)</td>
      <td>✅</td>
    </tr>
  </tbody>
</table>

<p>The honest framing: <strong>the Spark trades speed for capacity.</strong> It won’t beat a 5090 in raw token generation. But it can run models that simply don’t fit on consumer hardware — and for agent workloads (where prompt processing dominates), the GB10’s prefill performance is surprisingly strong.</p>

<h2 id="the-family-which-machine-did-i-pick">The family: which machine did I pick?</h2>

<p>Nvidia sells the reference design as the <strong>DGX Spark Founders Edition</strong>, but partners build their own:</p>

<table>
  <thead>
    <tr>
      <th>Machine</th>
      <th>Typical price (DE, Aug 2026)</th>
      <th>Notes</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><strong>ASUS Ascent GX10</strong> (mine)</td>
      <td>~€3,999</td>
      <td>1 TB SSD, 3 yr warranty</td>
    </tr>
    <tr>
      <td>Lenovo ThinkStation PGX</td>
      <td>~€4,784</td>
      <td>3 yr warranty, SSD upgradeable</td>
    </tr>
    <tr>
      <td>HP ZGX Nano G1n</td>
      <td>~€5,200+ (1 TB AT deal: €3,580)</td>
      <td>37 offers at idealo</td>
    </tr>
    <tr>
      <td>NVIDIA DGX Spark Founders</td>
      <td>~€5,599</td>
      <td>4 TB SSD, 1 yr warranty</td>
    </tr>
    <tr>
      <td>Dell Pro Max with GB10</td>
      <td>~€5,576+</td>
      <td>Business support</td>
    </tr>
  </tbody>
</table>

<p><strong>Where these prices come from.</strong> All figures are the <em>lowest listed offers</em> on German price-comparison sites, checked on <strong>19-20 August 2026</strong>: <a href="https://geizhals.de">Geizhals.de</a>, <a href="https://www.idealo.de">idealo.de</a>, and their Austrian siblings (<a href="https://geizhals.at">Geizhals.at</a>, <a href="https://www.idealo.at">idealo.at</a>) — plus direct shop listings (e-tec.at, cyberport.de, notebooksbilliger.de, galaxus.at). Prices were pulled from the comparison engines’ offer lists, which aggregate shop inventory in real time. A few caveats on the numbers:</p>

<ul>
  <li><strong>Availability fluctuates.</strong> Geizhals showed <em>no offers at all</em> for the Lenovo PGX on some days, while idealo listed it — inventory comes and goes.</li>
  <li><strong>The HP 1 TB at €3,580 is an Austrian deal</strong> (idealo.at), not a German one. The same machine lists around €5,200 in Germany. Cross-border shipping usually works, but check VAT and delivery terms.</li>
  <li><strong>The ASUS Ascent GX10 at €3,999</strong> is the 1 TB model — the 4 TB version costs roughly €5,400.</li>
  <li>Prices move fast on these machines; treat the table as a snapshot, not a promise. This is why I checked multiple sources rather than trusting a single shop.</li>
</ul>

<p>They’re all the same chip. You’re choosing SSD size, warranty, and chassis — not performance. I picked the <strong>ASUS Ascent GX10</strong> because it was the cheapest entry point with a 3-year warranty, and I plan to upgrade the storage.</p>

<h2 id="what-you-can-actually-do-with-it">What you can actually do with it</h2>

<p>Realistic use cases, in order of what I’ve found works well:</p>

<ol>
  <li><strong>Local LLM serving</strong> — Ollama, vLLM, or Nvidia NIM. This is the core use case, and it works great.</li>
  <li><strong>Agent workloads</strong> — running an AI agent (like the one helping me write this) entirely offline. No API bills.</li>
  <li><strong>Image generation</strong> — ComfyUI runs on it, and can handle models too big for 8 GB GPUs.</li>
  <li><strong>Video generation</strong> — possible (FLUX→Wan workflows exist), but plan your memory: it shares the 128 GB with everything else.</li>
  <li><strong>Fine-tuning small models</strong> — it has the capacity, though it’s not a training powerhouse.</li>
</ol>

<h2 id="the-honest-caveats">The honest caveats</h2>

<ul>
  <li><strong>No Windows.</strong> DGX OS is Linux. If you need x86/Windows software, keep a desktop.</li>
  <li><strong>Storage fills up fast.</strong> 1 TB is plenty for models if you’re disciplined — the 83 GB DeepSeek V4 Flash + ComfyUI models fit fine, but video models accumulate.</li>
  <li><strong>One GPU, one budget.</strong> LLM + image generation run <em>in parallel</em>, but they share the 128 GB. Plan your memory, not just your disk.</li>
  <li><strong>It’s not a gaming PC.</strong> Don’t buy one for that.</li>
</ul>

<h2 id="whats-next-in-this-series">What’s next in this series</h2>

<ol>
  <li><del>I Bought a DGX Spark. Here’s What It Actually Is.</del> ← you are here</li>
  <li>How I Set Up the ASUS Ascent GX10 (First Boot to SSH)</li>
  <li>Running DeepSeek V4 Flash Locally on a DGX Spark</li>
  <li>Ollama vs vLLM on a DGX Spark: Real Numbers</li>
  <li>ComfyUI on the DGX Spark: Images AND Video</li>
</ol>

<p><em>This post was written on the machine it describes — well, almost. The agent helping me draft it runs on DeepSeek V4 Flash today, and will run on my GX10 by the time this series finishes.</em></p>]]></content><author><name>Erik Zocher</name></author><category term="technology" /><category term="ai" /><category term="hardware" /><category term="dgx-spark" /><category term="gb10" /><category term="local-llm" /><category term="llm" /><category term="self-hosting" /><summary type="html"><![CDATA[A practical introduction to the NVIDIA DGX Spark and GB10 superchip: what the 128 GB unified memory actually means, which partner machines exist (ASUS Ascent, Lenovo PGX, HP ZGX, Dell), and what you can — and can't — do with a €4,000 AI mini-PC.]]></summary></entry><entry><title type="html">My Engineer Friend Knows the Buzzwords. Here’s What Actually Matters.</title><link href="https://erikzocher.dev/technology/2026/08/06/llm-mcp-field-guide.html" rel="alternate" type="text/html" title="My Engineer Friend Knows the Buzzwords. Here’s What Actually Matters." /><published>2026-08-06T21:40:00+02:00</published><updated>2026-08-06T21:40:00+02:00</updated><id>https://erikzocher.dev/technology/2026/08/06/llm-mcp-field-guide</id><content type="html" xml:base="https://erikzocher.dev/technology/2026/08/06/llm-mcp-field-guide.html"><![CDATA[<h1 id="my-engineer-friend-knows-the-buzzwords-heres-what-actually-matters">My Engineer Friend Knows the Buzzwords. Here’s What Actually Matters.</h1>

<p><em>2026-08-06 · 11 min read · [ai] [llm] [mcp] [skills] [prompting] [beginner]</em></p>

<p>Every week brings a new AI technology. Every week, articles explain it. And every week, my friend, a software engineer, reads them and still cannot say what he should focus on. He knows the vocabulary: RAG, agents, MCPs, embeddings, fine-tuning, and half a dozen more terms. He just cannot tell which of them matter, because almost nothing he reads answers that question.</p>

<p>That gap is real, and it is more common than most people admit. New platforms, paradigms, techniques, and tools get announced, hyped, and replaced within weeks, and click-bait articles tell you what is new, not what is worth learning. So I wrote him a field guide: the concepts that stay useful, the habits that get a reliable result, and the checklist that keeps a task honest. This is it, and it is yours too if you are in the same boat.</p>

<p><strong>The 30-second version:</strong> an LLM is a language model that predicts a useful next response from the context you give it. It does not automatically know your repository, your production state, or your intent; you supply those. Context engineering is choosing what the model sees. A harness is the environment around the model: instructions, tools, permissions, feedback loops. A skill is a reusable workflow for a class of tasks. MCP servers publish tools in a standard shape, and the model decides which to call. Write prompts like engineering specs, match your process to the task size, and verify the outcome before you believe it.</p>

<h2 id="the-mental-model">The Mental Model</h2>

<p>Five terms cover most of what you need to know. These are the concepts that do not go stale: the vocabulary changes weekly, the ideas underneath it much more slowly. Learn these and the rest is detail.</p>

<p><strong>LLM.</strong> A language model. Given the text and other inputs in its current context, it predicts a useful next response. It can explain, plan, write code, and decide which available tool to call. It does not automatically know your repository, production state, or intent. Those must be supplied or discovered.</p>

<p><strong>Context engineering.</strong> Choosing, structuring, and refreshing the information the model needs for one task. Useful context includes the goal, relevant files, constraints, error output, examples, and the definition of done. More context is not always better: irrelevant or stale material can distract the model.</p>

<p><strong>Harness.</strong> The environment around the model: its instructions, tools, permissions, repository, feedback loops, and user interface. Harness engineering is designing that environment so the model can work safely and repeatably. For example, a harness may provide Git access, require test runs before a completion claim, and block destructive commands without approval.</p>

<p><strong>Skill.</strong> A reusable workflow the harness can load for a class of tasks. It tells the agent when it applies, what sequence to follow, what information to gather, and how to verify or hand off the result. A skill is not a substitute for a well-written prompt: it supplies process, while your prompt supplies the goal and local constraints.</p>

<p><strong>MCP.</strong> Model Context Protocol. An <a href="https://modelcontextprotocol.io">MCP</a> server publishes tools or data sources in a standard shape. The host discovers the available tools, reads their descriptions and input schemas, and can ask the server to perform an action. The LLM chooses and sequences those tool calls. The MCP server is not the LLM and does not make decisions itself.</p>

<pre><code class="language-mermaid">flowchart LR
    A[You] --&gt;|prompt and approvals| B[Harness]
    D[Skills] --&gt; B
    B --&gt; C[LLM]
    C &lt;--&gt; E[MCP tools and data]
    C &lt;--&gt; F[Files, commands, test output]
    C --&gt;|answer| A
</code></pre>

<p>My own setup sits at this level. My agent on the Raspberry Pi runs inside a harness with a terminal, a browser, file access, and messaging, and it reaches external systems through MCP servers: web search, browser automation, health data, and a few more. The model makes the decisions. The servers do the work.</p>

<h2 id="write-prompts-like-an-engineering-spec">Write Prompts Like an Engineering Spec</h2>

<p>A good prompt reduces ambiguity. It does not need to be long; it needs to answer the questions that affect a correct result.</p>

<table>
  <thead>
    <tr>
      <th>Field</th>
      <th>What to state</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Goal</td>
      <td>The outcome or user-visible behavior you want.</td>
    </tr>
    <tr>
      <td>Context</td>
      <td>Relevant files, APIs, error output, business rules, or examples.</td>
    </tr>
    <tr>
      <td>Constraints</td>
      <td>What must not change, approved dependencies, performance or security limits.</td>
    </tr>
    <tr>
      <td>Definition of done</td>
      <td>The checks, tests, screenshots, or review criteria that demonstrate success.</td>
    </tr>
    <tr>
      <td>Response shape</td>
      <td>Whether you want a plan, a patch, a review, a table, or a concise handoff.</td>
    </tr>
  </tbody>
</table>

<h3 id="weak-prompt">Weak prompt</h3>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>Fix the departure board.
</code></pre></div></div>

<p>The model must guess the bug, the affected files, the intended behavior, and how much it may change.</p>

<h3 id="stronger-prompt">Stronger prompt</h3>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>Fix the line-filter options on the departure board.

Context: the board renders the expected departures, but the filter options are
wrong. In the All view there must be no line filter. In the Tram view, the only
options are "M4" and "27".

Constraints: preserve the existing query and do not change the options for
other views. Reuse the existing filter components; do not add a dependency.

Done when: the focused tests cover both views and the existing relevant tests
still pass.

Response: first identify the affected files and proposed change. Then make the
smallest patch and report the tests run.
</code></pre></div></div>

<p>Same intent, very different odds of a correct first attempt.</p>

<h3 id="a-reusable-template">A reusable template</h3>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>Goal:

Context:

Constraints:

Definition of done:

Response shape:
</code></pre></div></div>

<p>Start with the template, then remove the headings that genuinely do not matter. For a one-line change, a short goal, a constraint, and a check are often enough.</p>

<h2 id="match-the-process-to-the-task-size">Match the Process to the Task Size</h2>

<table>
  <thead>
    <tr>
      <th>Task size</th>
      <th>Appropriate approach</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Small</td>
      <td>State the desired change, constraints, and one focused check.</td>
    </tr>
    <tr>
      <td>Medium</td>
      <td>Ask for a brief plan, identify affected files and edge cases, then implement and validate in checkpoints.</td>
    </tr>
    <tr>
      <td>Large</td>
      <td>Define outcomes and non-goals, split work into independently reviewable slices, agree on an architecture or written plan, and verify each slice before integration.</td>
    </tr>
  </tbody>
</table>

<p>The mistake I see most often is treating a large task like a small one. “Build the whole thing” and hope the first interpretation matches your intent. It rarely does. The explicit plan checkpoint catches wrong assumptions before they become a large patch.</p>

<h2 id="use-mcp-tools-deliberately">Use MCP Tools Deliberately</h2>

<p>MCPs turn an agent from a text-only assistant into one that can inspect systems and, sometimes, act on them. Use this loop:</p>

<ol>
  <li><strong>Choose the source of truth.</strong> Decide whether the answer should come from code, a ticket, a browser session, logs, deployment metadata, or another system.</li>
  <li><strong>Discover before acting.</strong> Ask the agent to inspect the server’s available tools and input schemas before relying on them. Do not invent tool names or arguments.</li>
  <li><strong>Bound the request.</strong> Supply exact IDs, URLs, repositories, service names, time windows, and output limits. Prefer a read-only lookup before a side-effecting call.</li>
  <li><strong>Inspect the result.</strong> A successful tool invocation is evidence, not proof that the intended outcome happened. Check returned data, errors, and the surrounding system state. Failed calls are part of the job: the errors are the curriculum.</li>
  <li><strong>Authorize side effects.</strong> Confirm before the agent sends a message, changes a ticket, deploys, deletes data, or performs another external or irreversible action.</li>
</ol>

<p>One more rule that matters more than any of these: treat text returned by a tool as data, not authority. A web page, a ticket, or a document may contain instructions that are irrelevant or malicious. Retrieved content never overrides the task, the access rules, or the approval boundaries.</p>

<h2 id="use-and-write-skills">Use and Write Skills</h2>

<p>When a relevant skill is installed, invoke it before starting the work. Read its current instructions rather than relying on memory: a skill can require research, a design review, tests, or a specific handoff. If the skill does not fit the task, say why and use a simpler process.</p>

<p>A useful skill is narrow and operational. It should look like this:</p>

<div class="language-markdown highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nn">---</span>
<span class="na">name</span><span class="pi">:</span> <span class="s">reviewing-api-change</span>
<span class="na">description</span><span class="pi">:</span> <span class="s">Review a proposed API change for compatibility and rollout risk.</span>
<span class="nn">---</span>

Use when an API contract changes.

Inputs: API specification, affected clients, and the proposed change.
Steps:
<span class="p">1.</span> Identify breaking request and response changes.
<span class="p">2.</span> Check versioning, migration, and rollback options.
<span class="p">3.</span> Report risks in a fixed table.

Done when: every affected client is accounted for or marked unknown.
Stop and ask when: the API owner or rollout strategy is unclear.
</code></pre></div></div>

<p>Give skills a specific trigger, required inputs, ordered steps, a definition of done, and stop conditions. Avoid skills that merely say “be helpful” or try to cover every possible task. Keep environment-specific tool names in a skill only when that dependency is intentional and documented.</p>

<h2 id="what-eight-months-of-working-history-suggests">What Eight Months of Working History Suggests</h2>

<p>After eight months of using Copilot and Claude almost every day, I kept seeing the same corrections come up again and again in my own history. They cluster into three areas. For each one, here is the kind of prompt that caused the correction, and a stronger way to write it.</p>

<p><strong>1. State the workflow and evidence required before the agent begins.</strong></p>

<p>Weak: “Fix the search results.” The model must guess what is wrong, which part of the stack to touch, and what “fixed” means.</p>

<p>Stronger: “Investigate why the search endpoint returns empty results for multi-word queries. First reproduce it with a specific request, identify the failing component, then propose the smallest change and run the existing test suite before claiming it is done.”</p>

<p><strong>2. Name the required tool and environment when they affect correctness.</strong></p>

<p>Weak: “Deploy the new version.” Which environment? Which pipeline? Which rollback plan?</p>

<p>Stronger: “Deploy version 2.4.1 to the staging environment using the standard release script. Verify the health endpoint returns 200 and the version endpoint reports 2.4.1, then report the checks.”</p>

<p><strong>3. Bound the scope: say what may change and what must remain untouched.</strong></p>

<p>Weak: “Improve the performance of the search.” The model decides how far it may go, and that is exactly where overreach starts.</p>

<p>Stronger: “Optimize the search query path. You may change the query builder and the caching layer. Do not change the API contract, the database schema, or the frontend. Benchmark before and after and show the numbers.”</p>

<p>These are not rules for making every request longer. They are cues to include the information that would otherwise force a reviewer to correct a plausible but wrong assumption.</p>

<h2 id="the-pre-send-checklist">The Pre-Send Checklist</h2>

<p>Before sending a meaningful task, ask:</p>

<ol>
  <li>Is the desired outcome concrete enough to recognize as correct?</li>
  <li>Did I supply only the context needed to make the next decision?</li>
  <li>Did I state what must not change and any important boundaries?</li>
  <li>Did I say how the result will be verified?</li>
  <li>Does the task have external, destructive, or irreversible effects that require explicit approval?</li>
</ol>

<p>If the answer is yes to all five, the prompt is usually ready. If it is still ambiguous, ask the agent for its assumptions and plan before asking it to act.</p>

<p>That is the whole field guide. The mental model, the prompt, the sizing, the tools, the skills, the checklist. If your next prompt feels like a gamble, the fix is rarely a longer prompt. It is a better one.</p>

<p>For the full ladder from prompts to agent graphs, I wrote it up in <a href="https://erikzocher.github.io/technology/2026/07/31/working-with-ai.html">Working with AI: Five Ways, From Prompts to Agent Graphs</a>.</p>]]></content><author><name>Erik Zocher</name></author><category term="technology" /><category term="ai" /><category term="llm" /><category term="mcp" /><category term="skills" /><category term="prompting" /><category term="agents" /><category term="beginner" /><summary type="html"><![CDATA[A practical field guide to LLM-assisted work for software engineers who know the buzzwords but not what to focus on: the mental model, prompt structure, task sizing, deliberate MCP use, and writing skills.]]></summary></entry><entry><title type="html">The Button That Would Not Move Right</title><link href="https://erikzocher.dev/technology/2026/08/06/the-button-that-would-not-move.html" rel="alternate" type="text/html" title="The Button That Would Not Move Right" /><published>2026-08-06T01:42:00+02:00</published><updated>2026-08-06T01:42:00+02:00</updated><id>https://erikzocher.dev/technology/2026/08/06/the-button-that-would-not-move</id><content type="html" xml:base="https://erikzocher.dev/technology/2026/08/06/the-button-that-would-not-move.html"><![CDATA[<h1 id="the-button-that-would-not-move-right">The Button That Would Not Move Right</h1>

<p><em>2026-08-06 · 5 min read · [debugging] [css] [javascript] [web] [raspberry-pi]</em></p>

<p>I spent an evening trying to move a button to the right side of a banner, with the help of AI. Five separate CSS fixes, each one textbook-correct, and the button stayed stubbornly on the left. The cause was one line of JavaScript that had been disabling my entire layout the whole time.</p>

<p>There is something funny about that. An AI can generate astonishing videos out of thin air, write complex software, and keep an entire blog running, and yet moving one button to the right took an entire evening. The task seems so trivial that the failure reads like a joke. It is.</p>

<p><strong>The 30-second version:</strong> the cookie banner on my blog has a dismiss button, “Understood, carry on”. The design calls for the text on top and the button below it, right-aligned. My AI assistant wrote the CSS three different ways, verified it in a headless browser, and every measurement said the button was on the right. I kept looking at the real page and seeing it on the left. The gap between our measurements and reality turned out to be the bug: the page’s own JavaScript showed the banner with <code class="language-plaintext highlighter-rouge">display: block</code>, which turns off every flexbox property, and the assistant’s test tool had been quietly switching it back.</p>

<h2 id="the-setup">The Setup</h2>

<p>The banner lives in the blog’s layout file. When a visitor arrives, the banner starts hidden, and JavaScript reveals it if the visitor has not dismissed it before:</p>

<div class="language-javascript highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">if</span> <span class="p">(</span><span class="o">!</span><span class="nx">dismissed</span><span class="p">)</span> <span class="p">{</span>
  <span class="kd">var</span> <span class="nx">banner</span> <span class="o">=</span> <span class="nb">document</span><span class="p">.</span><span class="nx">getElementById</span><span class="p">(</span><span class="dl">'</span><span class="s1">cookie-banner</span><span class="dl">'</span><span class="p">);</span>
  <span class="k">if</span> <span class="p">(</span><span class="nx">banner</span><span class="p">)</span> <span class="p">{</span>
    <span class="nx">banner</span><span class="p">.</span><span class="nx">style</span><span class="p">.</span><span class="nx">display</span> <span class="o">=</span> <span class="dl">'</span><span class="s1">block</span><span class="dl">'</span><span class="p">;</span> <span class="c1">// show the banner</span>
    <span class="c1">// (banner dismissal logic)</span>
  <span class="p">}</span>
<span class="p">}</span>
</code></pre></div></div>

<p>The banner itself is a flexbox container:</p>

<div class="language-css highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nc">.cookie-banner</span> <span class="p">{</span>
  <span class="nl">display</span><span class="p">:</span> <span class="n">flex</span><span class="p">;</span>
  <span class="nl">flex-direction</span><span class="p">:</span> <span class="n">column</span><span class="p">;</span>  <span class="c">/* text on top, button below */</span>
<span class="p">}</span>
<span class="nc">.cookie-banner-ok</span> <span class="p">{</span>
  <span class="nl">align-self</span><span class="p">:</span> <span class="n">flex-end</span><span class="p">;</span>    <span class="c">/* button on the right */</span>
<span class="p">}</span>
</code></pre></div></div>

<p>Text on top, button below, button on the right. That is the whole design. It should work. It did not.</p>

<h2 id="attempt-1-the-obvious-fix">Attempt 1: The Obvious Fix</h2>

<p>The button was sitting on the left, directly under the text. I had my AI assistant add <code class="language-plaintext highlighter-rouge">align-self: flex-end</code> to the button. Textbook flexbox. The button stayed left.</p>

<h2 id="attempt-2-the-parent">Attempt 2: The Parent</h2>

<p>Maybe the parent was not actually a flex container at that moment. My AI assistant inspected the banner in its headless browser and confirmed <code class="language-plaintext highlighter-rouge">display: flex</code> and <code class="language-plaintext highlighter-rouge">flex-direction: column</code> were applied. It added <code class="language-plaintext highlighter-rouge">margin-left: auto</code> to the button as a belt-and-suspenders push to the right. The button stayed left.</p>

<h2 id="attempt-3-the-measure">Attempt 3: The Measure</h2>

<p>I asked my AI assistant to write a small script that measured the button’s position with <code class="language-plaintext highlighter-rouge">getBoundingClientRect</code>. The number said the button was 19 pixels from the right edge of the banner. Perfectly placed. I reloaded the page and the button was on the left. Both of us could not be right.</p>

<h2 id="attempt-4-the-cache">Attempt 4: The Cache</h2>

<p>Maybe my browser was serving a stale stylesheet. My blog’s CSS links carry a version parameter for cache busting, and on local builds that parameter was empty, which lets browsers cache the old CSS forever. We fixed the cache busting, I hard-refreshed, and the button was still on the left.</p>

<h2 id="attempt-5-the-comparison">Attempt 5: The Comparison</h2>

<p>My AI assistant took a screenshot of its headless browser. The button was on the right. I took a screenshot of my own browser. The button was on the left. The two screenshots disagreed, and that disagreement was the clue.</p>

<p>Here is what I saw (the button on the left):</p>

<p><img src="/assets/images/button-story/user-view.jpg" alt="What I saw: the cookie banner with the button below the text, on the left" /></p>

<p>And here is what my assistant saw (the button on the right):</p>

<p><img src="/assets/images/button-story/my-view.png" alt="What my assistant saw: the cookie banner with the button below the text, on the right" /></p>

<p>Same page, same browser engine, same moment in time. One button on the left, one button on the right. Screenshots do not lie, so one of us was not looking at the same reality.</p>

<p>My assistant’s measurement script did not just measure the banner. It set <code class="language-plaintext highlighter-rouge">banner.style.display = 'flex'</code> to make the banner visible before measuring. That inline style, set from JavaScript, was overriding the page’s own <code class="language-plaintext highlighter-rouge">display: block</code> and enabling the flexbox layout we were trying to verify. The test was not observing reality. The test was fixing the bug.</p>

<p>That is what went wrong: we had built a verification tool that silently corrected the very bug it was supposed to detect. Every measurement it took after that was measuring a page that did not exist. My own screenshot, taken from a plain browser tab with no helper script running, showed the actual behavior. The moment we stopped trusting the tool and started trusting the discrepancy, the cause was obvious.</p>

<p>And here is the funniest part. I never look at this blog’s code. Not just that evening, not a single line, ever: every change, fix, and deployment is just instructions to my AI assistant. So when the button refused to move for an entire evening and the assistant kept confidently reporting that everything was fine, there was only one thing left to try. I opened the browser’s developer tools myself. The answer was sitting in the elements panel, obvious within seconds. The one person who was never supposed to touch the code was the one who found the bug.</p>

<h2 id="the-actual-bug">The Actual Bug</h2>

<p>The page’s JavaScript reveals the banner with:</p>

<div class="language-javascript highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nx">banner</span><span class="p">.</span><span class="nx">style</span><span class="p">.</span><span class="nx">display</span> <span class="o">=</span> <span class="dl">'</span><span class="s1">block</span><span class="dl">'</span><span class="p">;</span>
</code></pre></div></div>

<p><code class="language-plaintext highlighter-rouge">display: block</code> is not a flex container. Every flexbox property we had written, <code class="language-plaintext highlighter-rouge">flex-direction</code>, <code class="language-plaintext highlighter-rouge">align-self</code>, <code class="language-plaintext highlighter-rouge">margin-left: auto</code>, all of them are ignored when the element is a block box. The CSS was correct the entire time. The JavaScript was replacing it with a plain block layout on every page load, and the measurement tool happened to replace it back.</p>

<p>The fix is one word:</p>

<div class="language-javascript highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nx">banner</span><span class="p">.</span><span class="nx">style</span><span class="p">.</span><span class="nx">display</span> <span class="o">=</span> <span class="dl">'</span><span class="s1">flex</span><span class="dl">'</span><span class="p">;</span>
</code></pre></div></div>

<h2 id="what-i-learned">What I Learned</h2>

<ul>
  <li><strong>Inline styles beat stylesheets.</strong> A <code class="language-plaintext highlighter-rouge">style</code> attribute set from JavaScript has higher priority than any CSS rule. <code class="language-plaintext highlighter-rouge">align-self: flex-end</code> in a stylesheet loses to <code class="language-plaintext highlighter-rouge">display: block</code> in an inline style, because the inline style changes which layout model applies at all.</li>
  <li><strong>Flexbox properties are powerless without <code class="language-plaintext highlighter-rouge">display: flex</code>.</strong> The mistake reads as a CSS problem but it is a layout-model problem. When flexbox alignment does nothing, check that the element actually is a flex container first.</li>
  <li><strong>Your test tool can become part of the bug.</strong> Any script that sets a property to make something testable can mask the very issue you are chasing. The fix is to measure on a fresh page load, with no instrumentation touching the element first.</li>
  <li><strong>Screenshots settle arguments.</strong> When two people look at the same page and see different things, one of them is not looking at the same page. Version numbers on assets, hard reloads, and fresh browser profiles narrow down which.</li>
  <li><strong>The human is the last verification layer.</strong> The whole point of the assistant is that I do not have to inspect code. But when a tool is confident and wrong at the same time, the human eye is the only check left. The first time I opened the developer tools myself, I found the bug in seconds. Hands-on inspection is not a fallback to be ashamed of. It is the final anchor.</li>
</ul>

<h2 id="the-checklist">The Checklist</h2>

<p>Next time an element ignores every alignment rule, go through these before touching CSS again:</p>

<ol>
  <li><strong>Is the element actually using the layout model you think?</strong> Check the computed <code class="language-plaintext highlighter-rouge">display</code> in DevTools. If it is <code class="language-plaintext highlighter-rouge">block</code>, flexbox and grid properties do nothing.</li>
  <li><strong>Is anything setting an inline style on it?</strong> Search the JavaScript for <code class="language-plaintext highlighter-rouge">style.display</code>, <code class="language-plaintext highlighter-rouge">style.flex</code>, or any <code class="language-plaintext highlighter-rouge">setAttribute('style')</code> on that element or its ancestors.</li>
  <li><strong>What does a fresh load show?</strong> Reload the page without running any of your own scripts. Your debugging tools can change the state they are debugging.</li>
  <li><strong>Is the browser serving the version you edited?</strong> Check the asset URL’s version parameter and hard-refresh or use a private window before assuming your change is live.</li>
</ol>

<p>The button now sits where it was always meant to be, on the right, under the text. Five fixes, one of them real, and the errors were the curriculum: each attempt taught me a layer of the stack, from stylesheet specificity to cache headers to the quiet power of inline styles. The last lesson was the best one: when your measurements disagree with reality, trust reality, and check what your tools are touching. 🎭</p>]]></content><author><name>Erik Zocher</name></author><category term="technology" /><category term="debugging" /><category term="css" /><category term="javascript" /><category term="web" /><category term="raspberry-pi" /><summary type="html"><![CDATA[A cookie banner button refused to move right. Five CSS fixes, three hours, and the culprit was a single line of JavaScript that turned off flexbox entirely.]]></summary></entry><entry><title type="html">I Gave an AI Agent Web Access. The CAPTCHA Was Not the Hard Part.</title><link href="https://erikzocher.dev/technology/2026/08/05/mcp-web-access.html" rel="alternate" type="text/html" title="I Gave an AI Agent Web Access. The CAPTCHA Was Not the Hard Part." /><published>2026-08-05T08:30:00+02:00</published><updated>2026-08-05T08:30:00+02:00</updated><id>https://erikzocher.dev/technology/2026/08/05/mcp-web-access</id><content type="html" xml:base="https://erikzocher.dev/technology/2026/08/05/mcp-web-access.html"><![CDATA[<h1 id="i-gave-an-ai-agent-web-access-the-captcha-was-not-the-hard-part">I Gave an AI Agent Web Access. The CAPTCHA Was Not the Hard Part.</h1>

<p>I wanted my Raspberry Pi agent to apply for apartments on <a href="https://www.immobilienscout24.de">Immobilienscout24</a> by itself. Load the listing, check the details, contact the owner. I expected the CAPTCHA to be the obstacle. It was not. The real obstacle was understanding what the CAPTCHA is actually protecting: identity. Once I did, the solution stopped being about outsmarting bot detection and became a clean division of labor between a human and a machine.</p>

<p><strong>The 30-second version:</strong> headless browsers, spoofed fingerprints, and virtual displays all got blocked within seconds. What worked was a real headed browser with a persistent profile, logged in once by a human over a remote desktop session, then driven by the agent over the DevTools protocol. The boundary that matters is human identity verification: the person proves who they are once, the agent reuses the result until the session expires. No CAPTCHA solving, no credential extraction, nothing against the site’s terms.</p>

<h2 id="the-case-study-immobilienscout24">The Case Study: Immobilienscout24</h2>

<p>I wanted my agent to contact apartment owners on Immobilienscout24. Sounds simple. It took three failed approaches and one working one.</p>

<table>
  <thead>
    <tr>
      <th>Approach</th>
      <th>Browser mode</th>
      <th>Profile state</th>
      <th>Result</th>
      <th>Time to failure</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Chrome DevTools MCP</td>
      <td>Headless</td>
      <td>Fresh</td>
      <td>CAPTCHA on every page</td>
      <td>Seconds</td>
    </tr>
    <tr>
      <td>Browser Use + browserforge fingerprint</td>
      <td>Headless, spoofed traits</td>
      <td>Fresh</td>
      <td>CAPTCHA</td>
      <td>Seconds</td>
    </tr>
    <tr>
      <td>Chromium on Xvfb</td>
      <td>Headed, virtual display</td>
      <td>Fresh</td>
      <td>CAPTCHA</td>
      <td>Seconds</td>
    </tr>
    <tr>
      <td>Desktop Chromium via Raspberry Pi Connect</td>
      <td>Headed, real desktop</td>
      <td>Persistent, human-logged-in</td>
      <td>Worked</td>
      <td>n/a</td>
    </tr>
  </tbody>
</table>

<h3 id="attempt-1-chrome-devtools-mcp">Attempt 1: Chrome DevTools MCP</h3>

<p><a href="https://github.com/ChromeDevTools/chrome-devtools-mcp">Chrome DevTools MCP</a> is designed to control and inspect a live Chrome browser. In my setup I drove it against a headless Chrome instance, and the result was an instant CAPTCHA: “Ich bin kein Roboter” (German for “I am not a robot”) on every page. Worth noting: the headless launch was my choice, not an inherent property of the MCP server, which can attach to a normal running browser too. The lesson stands either way: headless Chrome is exactly what bot detection is built to flag.</p>

<h3 id="attempt-2-browser-use--a-spoofed-fingerprint">Attempt 2: Browser Use + a spoofed fingerprint</h3>

<p>Next I tried <a href="https://github.com/browser-use/browser-use">Browser Use</a> with a crafted Windows Chrome fingerprint via <a href="https://github.com/daijro/browserforge">browserforge</a>, still headless. The idea was to spoof enough browser traits to pass. It changed nothing: CAPTCHA again within seconds. Fingerprint spoofing is a cat-and-mouse game, and the site’s detection did not care what my headers claimed.</p>

<h3 id="attempt-3-a-headed-browser-on-a-virtual-display">Attempt 3: A headed browser on a virtual display</h3>

<p>Maybe the problem was headless rendering itself. So I ran a real headed Chromium on a virtual display (<a href="https://en.wikipedia.org/wiki/Xvfb">Xvfb</a>). Result: still CAPTCHA. Headed-vs-headless was not the deciding factor. Automation traces, session patterns, and IP reputation matter too.</p>

<h3 id="what-actually-worked">What actually worked</h3>

<p>A real desktop Chromium with a persistent profile, logged in manually once, then driven over <a href="https://chromedevtools.github.io/devtools-protocol/">CDP (Chrome DevTools Protocol)</a>. It worked in my setup, and I can say exactly which conditions were in place: a normal browser in a normal desktop session, a persistent authenticated profile, and deliberate limits on what the agent does.</p>

<p>The login happened through <a href="https://connect.raspberrypi.com">Raspberry Pi Connect</a>, the Pi’s built-in remote screen-sharing service. From any browser I can open connect.raspberrypi.com, pick my Pi, and see its desktop as if I were sitting in front of it. There I launched a normal Chromium window, went to Immobilienscout24, and authenticated as myself: username, password, and the CAPTCHA, all typed by a human in a real browser on a real desktop.</p>

<p>The key detail is where the cookies landed. That login wrote the session into Chromium’s persistent profile (<code class="language-plaintext highlighter-rouge">~/.config/chromium</code>), the same profile the agent’s browser uses. When the agent later starts Chromium with that profile and connects over CDP, it inherits the human-established session. No cookie theft, no token extraction, no replaying of credentials: authentication happened once, by a person, and everything after it reuses the result.</p>

<h3 id="the-architecture-in-one-picture">The architecture, in one picture</h3>

<pre><code class="language-mermaid">flowchart LR
    A["Human login once&lt;br/&gt;(username, password, CAPTCHA)"] --&gt; B["Persistent browser profile&lt;br/&gt;(~/.config/chromium)"]
    B --&gt; C["Real headed Chromium&lt;br/&gt;driven over CDP"]
    C --&gt; D["Agent: load listings,&lt;br/&gt;draft message"]
    D --&gt; E["Human reviews&lt;br/&gt;before sending"]
</code></pre>

<p>The agent runs until the session expires or the site asks for renewed verification. Then the human steps in again.</p>

<h2 id="the-tools-three-approaches-to-web-access">The Tools: Three Approaches to Web Access</h2>

<p>For web access, I have learned about three approaches that are useful for interacting with webpages. Two are MCP servers, one is a browser automation framework, and the distinction matters:</p>

<table>
  <thead>
    <tr>
      <th>Approach</th>
      <th>What it is</th>
      <th>Best for</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><strong><a href="https://github.com/ChromeDevTools/chrome-devtools-mcp">Chrome DevTools MCP</a></strong></td>
      <td>MCP server; controls and inspects a live Chrome browser via CDP</td>
      <td>Clicking, form filling, debugging, persistent sessions</td>
    </tr>
    <tr>
      <td><strong><a href="https://github.com/dondai1234/master-fetch">Hound MCP</a></strong></td>
      <td>MCP server; fetches and extracts text with automatic escalation</td>
      <td>Reading content quickly, search, PDFs</td>
    </tr>
    <tr>
      <td><strong><a href="https://github.com/browser-use/browser-use">Browser Use</a></strong></td>
      <td>Browser automation library and agent platform with MCP integration</td>
      <td>Scripted browsing, scraping, multi-step flows</td>
    </tr>
  </tbody>
</table>

<h3 id="what-is-an-mcp-anyway">What is an MCP anyway?</h3>

<p>MCP (<a href="https://modelcontextprotocol.io">Model Context Protocol</a>) is a standard way for AI agents to plug into tools. Instead of every agent inventing its own way to talk to every service, an MCP server exposes a common interface. That interface can wrap all kinds of capabilities: browsers, API endpoints, databases, file systems, developer tools. This post focuses on the web-related ones, but the protocol is not limited to browsers.</p>

<p>The easiest way to think about it: <strong>the LLM is the brain, and MCP servers are the hands.</strong> The brain alone cannot touch anything: it cannot open a website, click a button, or read a PDF. The hands reach out to third-party sources, fetch data, and bring results back. The brain decides, the hands do. What makes it practical is the MCP host: it discovers what each server can do, coordinates access, and hands the relevant tool definitions to the model.</p>

<h3 id="can-they-solve-captchas">Can they solve CAPTCHAs?</h3>

<p>The question everyone asks first. Some offerings in this space advertise CAPTCHA solving, including hosted plans of tools like Browser Use and Hound. I did not pursue that path. Such services are unreliable, can run contrary to a site’s terms, and risk getting the account flagged. The approach in this post needs no CAPTCHA solving at all, which is exactly why it stays on the right side of the line.</p>

<h2 id="the-honest-truth-about-anti-bot-systems">The Honest Truth About Anti-Bot Systems</h2>

<p>Every serious website now runs bot detection: <a href="https://www.imperva.com">Imperva</a>, <a href="https://www.cloudflare.com">Cloudflare</a>, <a href="https://www.google.com/recaptcha/about/">reCAPTCHA</a>, <a href="https://www.perimeterx.com">PerimeterX</a>. These systems check far more than “does this look like a browser?”:</p>

<ul>
  <li>Headless vs. headed rendering</li>
  <li>WebDriver flags and automation fingerprints</li>
  <li>IP reputation</li>
  <li>Cookie and session consistency</li>
  <li>Mouse movement and timing patterns</li>
</ul>

<p>The lesson I wish I had read before spending an evening on it: <strong>no tool choice by itself gets you past this.</strong> What survived was not a stealthier browser but a boundary that removed the need to be stealthy at all.</p>

<h2 id="safety-box">Safety Box</h2>

<p>Wherever you draw your own boundary, these rules keep you out of trouble:</p>

<ul>
  <li><strong>No credential extraction.</strong> The agent never handles passwords or tokens.</li>
  <li><strong>No CAPTCHA bypassing.</strong> A human solves the CAPTCHA, once, at login.</li>
  <li><strong>Rate-limit requests.</strong> The agent checks every ten minutes, not every second.</li>
  <li><strong>Review messages before sending.</strong> A human approves what goes out.</li>
  <li><strong>Stop when asked.</strong> If the site asks for verification again, stop and let a human take over.</li>
</ul>

<p>Each site’s terms, rate limits, consent rules, and anti-spam policy decide what is acceptable. Reading and following them is part of the job.</p>

<h2 id="a-reusable-framework">A Reusable Framework</h2>

<p>The pattern here is bigger than one apartment site. For any site you want to automate, identify three things:</p>

<ol>
  <li><strong>Identity-establishing steps</strong> (login, verification, the things only a person should do): keep these human.</li>
  <li><strong>Explicitly permitted repeatable steps</strong> (reading listings, checking status, drafting within the site’s rules): automate these.</li>
  <li><strong>Required human approvals</strong> (sending a message, publishing, anything irreversible): keep these in the loop.</li>
</ol>

<p>My agent now checks Immobilienscout24 every ten minutes, reads new listings, and drafts contact messages. The human input is the login that happens when the session expires, and the review before anything is sent. That boundary is not a compromise. It is the right shape for automation that stays honest: a human at the door, an agent everywhere else.</p>]]></content><author><name>Erik Zocher</name></author><category term="technology" /><category term="mcp" /><category term="ai" /><category term="agents" /><category term="web-scraping" /><category term="captcha" /><category term="automation" /><category term="raspberry-pi" /><summary type="html"><![CDATA[What actually worked when I gave an agent access to Immobilienscout24: three failed automation approaches, the human-in-the-loop architecture that worked, and a framework you can reuse for any protected website.]]></summary></entry><entry><title type="html">How I Got LTX-2.3 Running in ComfyUI: Six Errors and a Lot of Patience</title><link href="https://erikzocher.dev/technology/2026/08/04/ltx23-comfyui-debugging.html" rel="alternate" type="text/html" title="How I Got LTX-2.3 Running in ComfyUI: Six Errors and a Lot of Patience" /><published>2026-08-04T02:44:00+02:00</published><updated>2026-08-04T02:44:00+02:00</updated><id>https://erikzocher.dev/technology/2026/08/04/ltx23-comfyui-debugging</id><content type="html" xml:base="https://erikzocher.dev/technology/2026/08/04/ltx23-comfyui-debugging.html"><![CDATA[<h1 id="how-i-got-ltx-23-running-in-comfyui-six-errors-and-a-lot-of-patience">How I Got LTX-2.3 Running in ComfyUI: Six Errors and a Lot of Patience</h1>

<p>In my <a href="https://erikzocher.github.io/technology/2026/08/03/my-first-ai-video.html">last post</a> I described how I turned my Raspberry Pi into a remote control for a Windows PC with an RTX 3070 Ti, generating a wobbly two-tailed cat video with AnimateDiff. The natural next step was to try a real video model. Not a motion module bolted onto an image model, but an actual text-to-video model: <a href="https://huggingface.co/Lightricks/LTX-2.3-fp8">LTX-2.3</a>, a 22B-parameter model from Lightricks that generates video and audio in one pass.</p>

<p>What followed was the most educational debugging session I have had in a while. Six distinct errors, each one hiding the next, each one teaching me something about how ComfyUI actually works under the hood.</p>

<p><strong>The 30-second version:</strong> a 22-billion-parameter video model CAN run on an 8 GB GPU, but only if you wire it as the joint audio-video model it actually is. Five of the six errors were wiring mistakes, not model failures: the wrong text-encoder loader, the wrong VAE file, and a missing audio-latent chain. The recipe at the bottom works. It is slow (20+ minutes for a two-second clip), but it runs entirely on commodity hardware. The errors are the curriculum: each one teaches a real fact about how ComfyUI and LTX-2.3 work.</p>

<h2 id="the-setup">The Setup</h2>

<p>Same hardware as last time:</p>

<table>
  <thead>
    <tr>
      <th>Machine</th>
      <th>Role</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Raspberry Pi 5</td>
      <td>Brain: builds the workflow JSON, submits it via the ComfyUI API</td>
    </tr>
    <tr>
      <td>Windows PC (RTX 3070 Ti, 8 GB VRAM)</td>
      <td>Muscle: runs ComfyUI</td>
    </tr>
  </tbody>
</table>

<p>The model files: <code class="language-plaintext highlighter-rouge">ltx-2.3-22b-dev-fp8.safetensors</code> (22B params, fp8 quantized, 29 GB on disk), the Gemma text encoder, and the model’s VAE.</p>

<p><strong>Quick primer: what is a VAE?</strong> A VAE (Variational Autoencoder) is the translation layer between <em>pixels</em> and <em>latents</em>. Diffusion models do not work on images or videos directly: they work on a compressed, noisy mathematical representation called a latent space, which is much smaller than the actual pixels. The VAE has two halves: the <strong>encoder</strong> compresses pixels into latents (used for image-to-image and video-to-video), and the <strong>decoder</strong> expands latents back into visible pixels (used at the end of every generation). Think of it as the codec of the diffusion world: the model thinks and dreams in compressed form, and the VAE is what turns those dreams back into something you can see. Getting the wrong VAE is like connecting a Blu-ray player to a VHS-era TV: the signal is there, but nothing displays correctly.</p>

<h2 id="the-six-errors-at-a-glance">The Six Errors at a Glance</h2>

<table>
  <thead>
    <tr>
      <th>#</th>
      <th>Error</th>
      <th>Root cause</th>
      <th>Fix</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>1</td>
      <td><code class="language-plaintext highlighter-rouge">clip input is invalid: None</code></td>
      <td>LTX-2.3 has no bundled text encoder</td>
      <td>Load the Gemma encoder separately</td>
    </tr>
    <tr>
      <td>2</td>
      <td><code class="language-plaintext highlighter-rouge">got 4 and 3</code> dimensions</td>
      <td>Generic <code class="language-plaintext highlighter-rouge">CLIPLoader</code> produces the wrong embedding shape</td>
      <td>Use <code class="language-plaintext highlighter-rouge">LTXAVTextEncoderLoader</code></td>
    </tr>
    <tr>
      <td>3</td>
      <td><code class="language-plaintext highlighter-rouge">'dict' object has no attribute 'sample'</code></td>
      <td>Sampler embedded as inline object</td>
      <td>Make the sampler its own top-level node</td>
    </tr>
    <tr>
      <td>4</td>
      <td><code class="language-plaintext highlighter-rouge">expected 16 channels, got 128</code></td>
      <td>Joint AV latent not split before decoding</td>
      <td>Insert <code class="language-plaintext highlighter-rouge">LTXVSeparateAVLatent</code></td>
    </tr>
    <tr>
      <td>5</td>
      <td><code class="language-plaintext highlighter-rouge">tuple index out of range</code></td>
      <td>No audio latent in the workflow</td>
      <td>Build the audio chain with <code class="language-plaintext highlighter-rouge">LTXVEmptyLatentAudio</code></td>
    </tr>
    <tr>
      <td>6</td>
      <td><code class="language-plaintext highlighter-rouge">expected 16 channels, got 128</code> again</td>
      <td>Wrong VAE file (<code class="language-plaintext highlighter-rouge">ae.safetensors</code> is LTX-2’s)</td>
      <td>Use the VAE from the checkpoint, slot 2</td>
    </tr>
  </tbody>
</table>

<h2 id="error-1-clip-input-is-invalid-none">Error 1: “clip input is invalid: None”</h2>

<p>The first submission failed immediately. LTX-2.3 does not bundle a text encoder inside its checkpoint, so the workflow tried to grab a CLIP model that was not there.</p>

<p><strong>Fix:</strong> Load the text encoder separately. ComfyUI has a <code class="language-plaintext highlighter-rouge">CLIPLoader</code> node with a <code class="language-plaintext highlighter-rouge">type</code> parameter, and for LTX the correct value is <code class="language-plaintext highlighter-rouge">ltxv</code>, pointing at the Gemma encoder.</p>

<h2 id="error-2-tensors-must-have-same-number-of-dimensions-got-4-and-3">Error 2: “Tensors must have same number of dimensions: got 4 and 3”</h2>

<p>Now the model loaded, but the sampler exploded. This one took a while. The generic <code class="language-plaintext highlighter-rouge">CLIPLoader</code> produces text embeddings in a shape that LTX-2.3’s video embedding connector cannot digest.</p>

<p><strong>Fix:</strong> Use the node made for this exact purpose: <code class="language-plaintext highlighter-rouge">LTXAVTextEncoderLoader</code>. It knows how to pair the <a href="https://huggingface.co/google/gemma-3-12b-it">Gemma text encoder</a> with the LTX checkpoint and produces embeddings the model accepts.</p>

<h2 id="error-3-dict-object-has-no-attribute-sample">Error 3: “‘dict’ object has no attribute ‘sample’”</h2>

<p>The workflow validated but crashed at runtime. My JSON embedded the sampler as an inline object instead of referencing a separate sampler node.</p>

<p><strong>Fix:</strong> In the API format, every node must be a top-level entry, and connections use <code class="language-plaintext highlighter-rouge">["node_id", output_index]</code> references. The sampler got its own node, and the sampler node got a proper reference to it.</p>

<h2 id="error-4-expected-input-to-have-16-channels-but-got-128-channels">Error 4: “expected input to have 16 channels, but got 128 channels”</h2>

<p>This was the first clue about LTX-2.3’s architecture. The VAE decoder expects 16 channels, but the sampler produced a latent with 128 channels. LTX-2.3 is a joint audio-video model: the latent contains both the video stream and the audio stream, interleaved.</p>

<p><strong>Fix:</strong> Insert <code class="language-plaintext highlighter-rouge">LTXVSeparateAVLatent</code> between the sampler and the video decoder. It splits the joint latent into its video and audio halves.</p>

<h2 id="error-5-tuple-index-out-of-range">Error 5: “tuple index out of range”</h2>

<p>The separator node complained it could not find the audio half. Right, because the workflow never created one. The joint AV model needs an audio latent to pair with the video latent, even when you only want the video.</p>

<p><strong>Fix:</strong> Build the full audio chain: <code class="language-plaintext highlighter-rouge">LTXVAudioVAELoader</code> loads the audio VAE from the checkpoint, <code class="language-plaintext highlighter-rouge">LTXVEmptyLatentAudio</code> creates an empty audio latent, and <code class="language-plaintext highlighter-rouge">LTXVConcatAVLatent</code> merges the video and audio latents into the joint structure the sampler expects. After sampling, <code class="language-plaintext highlighter-rouge">LTXVSeparateAVLatent</code> splits them again.</p>

<h2 id="error-6-the-same-128-channel-error-again">Error 6: the same 128-channel error, again</h2>

<p>With the full AV chain in place, the sampler finally ran. The separator split the latent. And the decoder still choked on 128 channels.</p>

<p>The culprit was the VAE file itself. I had pointed the decoder at the separately downloaded <code class="language-plaintext highlighter-rouge">ae.safetensors</code>, which is LTX-2’s VAE and expects 16 channels. LTX-2.3 is a different architecture with a 128-channel latent, and its VAE ships inside the checkpoint.</p>

<p><strong>Fix:</strong> Take the VAE from the checkpoint loader’s second output slot instead of loading a separate VAE file.</p>

<h2 id="the-working-recipe">The Working Recipe</h2>

<p>The final workflow, for anyone who wants to skip the six hours of debugging. The LTX custom nodes come from the <a href="https://github.com/Lightricks/ComfyUI-LTXVideo">ComfyUI-LTXVideo</a> repo, and the whole thing runs inside <a href="https://www.comfy.org">ComfyUI</a>:</p>

<pre><code class="language-mermaid">flowchart TD
    subgraph Loaders
        CKPT["CheckpointLoaderSimple&lt;br/&gt;(ltx-2.3-22b-dev-fp8)"]
        TENC["LTXAVTextEncoderLoader&lt;br/&gt;(gemma-3-12B)"]
        ALOAD["LTXVAudioVAELoader"]
    end

    subgraph Conditioning
        TE1["CLIPTextEncode (positive)"]
        TE2["CLIPTextEncode (negative)"]
        COND["LTXVConditioning"]
    end

    subgraph Latents
        VID["EmptyLTXVLatentVideo"]
        AUD["LTXVEmptyLatentAudio"]
        CONCAT["LTXVConcatAVLatent"]
    end

    subgraph Sampling
        SCHED["LTXVScheduler"]
        KSEL["KSamplerSelect"]
        NOISE["RandomNoise"]
        GUIDER["CFGGuider"]
        SAMP["SamplerCustomAdvanced"]
    end

    subgraph Output
        SEP["LTXVSeparateAVLatent"]
        VAEDEC["VAEDecode"]
        SAVE["SaveAnimatedWEBP"]
    end

    CKPT --&gt;|model| GUIDER
    CKPT --&gt;|"vae (slot 2)"| VAEDEC
    TENC --&gt; TE1
    TENC --&gt; TE2
    TE1 --&gt; COND
    TE2 --&gt; COND
    COND --&gt; GUIDER
    VID --&gt; CONCAT
    ALOAD --&gt; AUD
    AUD --&gt; CONCAT
    CONCAT --&gt; SAMP
    GUIDER --&gt; SAMP
    SCHED --&gt;|sigmas| SAMP
    KSEL --&gt;|sampler| SAMP
    NOISE --&gt;|noise| SAMP
    SAMP --&gt; SEP
    SEP --&gt;|video| VAEDEC
    VAEDEC --&gt; SAVE
</code></pre>

<p>Here is the result, 49 frames at 512x512, generated from the prompt “a cute orange cat walking through a sunlit park”:</p>

<blockquote>
  <p><strong>Note on sound:</strong> this video is silent by design. The workflow creates an <em>empty audio latent</em> (<code class="language-plaintext highlighter-rouge">LTXVEmptyLatentAudio</code>) to satisfy LTX-2.3’s joint audio-video architecture, and the output format (animated WebP, then converted to MP4) does not carry audio anyway. LTX-2.3 <em>can</em> generate sound when the audio path is conditioned properly, but that is a separate project.</p>
</blockquote>

<video controls="" loop="" muted="" playsinline="" width="100%" style="max-width:512px; border-radius:8px;">
  <source src="/assets/videos/ltx23-cat.mp4" type="video/mp4" />
  Your browser does not support the video tag.
</video>

<h2 id="get-the-workflow">Get the Workflow</h2>

<p>The workflow is available as an editable template with a placeholder prompt, in two formats:</p>

<ul>
  <li>For drag-and-drop into the ComfyUI canvas: <a href="/assets/workflows/ltx23-t2v-template-ui.json">ltx23-t2v-template-ui.json</a></li>
  <li>For the API: <a href="/assets/workflows/ltx23-t2v-template.json">ltx23-t2v-template.json</a></li>
</ul>

<p>In the JSON, node <code class="language-plaintext highlighter-rouge">4</code> holds the prompt. Replace <code class="language-plaintext highlighter-rouge">REPLACE_WITH_YOUR_PROMPT</code> with your own text. For reference, the prompt that produced the cat video above was:</p>

<blockquote>
  <p><em>“a cute orange cat walking through a sunlit park, cinematic lighting, smooth motion, high quality”</em></p>
</blockquote>

<p><strong>Two ways to use it:</strong></p>

<ol>
  <li><strong>In the ComfyUI interface:</strong> download the <code class="language-plaintext highlighter-rouge">-ui</code> JSON, then drag and drop it onto the ComfyUI canvas. The nodes appear, ready to run. Double-click node <code class="language-plaintext highlighter-rouge">4</code> to edit the prompt.</li>
  <li><strong>Via the API (how I did it from the Pi):</strong>
    <div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>curl <span class="nt">-X</span> POST http://&lt;your-comfyui&gt;:8188/prompt <span class="se">\</span>
  <span class="nt">-H</span> <span class="s2">"Content-Type: application/json"</span> <span class="se">\</span>
  <span class="nt">-d</span> <span class="s1">'{"prompt": &lt;workflow-json&gt;, "client_id": "puck-pi"}'</span>
</code></pre></div>    </div>
  </li>
</ol>

<p><strong>What you need installed:</strong> the <a href="https://github.com/Lightricks/ComfyUI-LTXVideo">ComfyUI-LTXVideo</a> custom nodes, the LTX-2.3 model in <code class="language-plaintext highlighter-rouge">models/checkpoints</code>, and the Gemma text encoder in <code class="language-plaintext highlighter-rouge">models/text_encoders</code>. The VAE comes from the checkpoint itself, no separate file needed.</p>

<h2 id="what-i-learned">What I Learned</h2>

<ul>
  <li><strong>LTX-2.3 is not a video model with an audio add-on.</strong> It is a joint audio-video model. The latent is one interleaved structure, and every stage of the pipeline has to respect that: audio latent creation, concatenation before sampling, separation after.</li>
  <li><strong>The right node matters more than the right parameter.</strong> Every error here was about wiring, not about values. When a ComfyUI workflow fails in a strange way, question the node choice before the settings.</li>
  <li><strong>Checkpoints can carry their own VAE.</strong> The separate VAE file in my models folder looked correct but was for the wrong model version. The VAE that matches the model lives inside the checkpoint.</li>
  <li><strong>8 GB of VRAM will run a 22B model, technically.</strong> The render took over 20 minutes and the GPU utilization hovered around 10%. It works, but it is the slowest way to generate a two-second clip. For this model, 12-16 GB of VRAM is the honest minimum for comfortable use.</li>
  <li><strong>Debugging through an API is a great teacher.</strong> Because I drove ComfyUI from the Pi over its REST API, I had to read every error message, check every node interface, and understand the graph end to end. No clicking around in a UI hoping something works.</li>
</ul>

<p>The two-tailed cat from last time has a new cousin now. Same cat, same park, this one came from a real 22-billion-parameter video model, rendered at 24 frames per second, through six errors and one very patient Raspberry Pi.</p>

<h2 id="the-takeaway-framework">The Takeaway Framework</h2>

<p>Next time you wire a new model into ComfyUI and it fails, run this sequence instead of guessing:</p>

<ol>
  <li><strong>Question the node before the parameter.</strong> Every one of my six errors was a wiring problem, not a settings problem. When something fails in a strange way, suspect the node choice first.</li>
  <li><strong>Read the model’s architecture before touching the graph.</strong> LTX-2.3 being a joint audio-video model explained errors 4, 5, and 6 in advance. The model card tells you what the latent looks like; the errors are just the architecture talking.</li>
  <li><strong>Trust the checkpoint over separate files.</strong> The VAE that matched the model was inside the checkpoint, not in the misleadingly named <code class="language-plaintext highlighter-rouge">ae.safetensors</code>. When a downloaded companion file disagrees with the model, the model is usually right.</li>
  <li><strong>Dump the raw history when the UI hides the error.</strong> The ComfyUI history API can return empty keys for failed nodes; the full traceback is in the raw JSON. Do not debug from a truncated message.</li>
</ol>

<p>Six errors, six lessons, one working recipe. The errors are the curriculum.</p>]]></content><author><name>Erik Zocher</name></author><category term="technology" /><category term="comfyui" /><category term="ai" /><category term="video" /><category term="ltx" /><category term="debugging" /><category term="raspberry-pi" /><summary type="html"><![CDATA[The full debugging journey of getting LTX-2.3 (22B video model) running on an 8 GB VRAM GPU: six distinct errors, their root causes, and the working workflow.]]></summary></entry><entry><title type="html">My First AI Video: A Raspberry Pi, a Windows PC, and One Very Patient Cat</title><link href="https://erikzocher.dev/technology/2026/08/03/my-first-ai-video.html" rel="alternate" type="text/html" title="My First AI Video: A Raspberry Pi, a Windows PC, and One Very Patient Cat" /><published>2026-08-03T14:00:00+02:00</published><updated>2026-08-03T14:00:00+02:00</updated><id>https://erikzocher.dev/technology/2026/08/03/my-first-ai-video</id><content type="html" xml:base="https://erikzocher.dev/technology/2026/08/03/my-first-ai-video.html"><![CDATA[<h1 id="my-first-ai-video-a-raspberry-pi-a-windows-pc-and-one-very-patient-cat">My First AI Video: A Raspberry Pi, a Windows PC, and One Very Patient Cat</h1>

<p>Every AI hobbyist reaches the moment where still images stop being enough. For me that moment was last week, when I realized my Raspberry Pi 5, for all its charms, will never render a video. Not one frame. The Pi is my always-on agent, my Telegram butler, my home automation brain. I wrote about <a href="https://erikzocher.github.io/technology/2026/08/02/raspberry-pi-ai-agent.html">building it and its skills</a> a few days earlier. But AI video generation needs a GPU, and the Pi has none.</p>

<p>So I built a small Frankenstein: the Pi stays the brain, and a Windows PC with an NVIDIA RTX 3070 Ti became the muscle.</p>

<p><strong>The 30-second version:</strong> I ran ComfyUI on a Windows PC with an RTX 3070 Ti and drove it remotely from my Raspberry Pi over the LAN. The Pi submits a workflow as JSON, the PC renders it on the GPU, and the result comes back. No cables, no cloud. The first video: 16 frames at 512x512, 20 steps, about a minute of render time. It is a wobbly two-tailed cat in a vaguely sunlit park, and it is the most satisfying minute I have spent on this hobby so far.</p>

<h2 id="the-setup">The Setup</h2>

<p>Two machines, one home network:</p>

<table>
  <thead>
    <tr>
      <th>Machine</th>
      <th>Role</th>
      <th>Hardware</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Raspberry Pi 5</td>
      <td>Brain: runs Hermes, controls everything</td>
      <td>16 GB RAM, no GPU</td>
    </tr>
    <tr>
      <td>Windows PC</td>
      <td>Muscle: runs ComfyUI</td>
      <td>RTX 3070 Ti (8 GB VRAM), 16 GB RAM</td>
    </tr>
  </tbody>
</table>

<p>The magic ingredient is the <a href="https://www.comfy.org">ComfyUI API</a>. ComfyUI, the node-based AI image and video tool, exposes a REST API. Any machine on the network can submit a workflow as JSON and pull back the result. The Pi talks to the PC over the LAN, no cables, no cloud.</p>

<pre><code class="language-mermaid">flowchart LR
    PI["Raspberry Pi 5&lt;br/&gt;(brain, no GPU)"] --&gt;|"workflow JSON&lt;br/&gt;POST :8188/prompt"| CU["ComfyUI on Windows PC&lt;br/&gt;(RTX 3070 Ti, 8 GB)"]
    CU --&gt;|"renders 16 frames&lt;br/&gt;512x512, 20 steps"| GPU["GPU"]
    GPU --&gt;|"animated webp"| CU
    CU --&gt;|"result back over LAN"| PI
</code></pre>

<h2 id="what-we-did">What We Did</h2>

<p>Setting it up took longer than the actual generation, which is the classic pattern. Here is the path:</p>

<ol>
  <li>
    <p><strong>Install ComfyUI on Windows.</strong> The <a href="https://www.comfy.org/download">Desktop app</a>, with NVIDIA support selected. It was running with <code class="language-plaintext highlighter-rouge">--listen</code> so other machines could reach it.</p>
  </li>
  <li>
    <p><strong>Find each other.</strong> The Pi checks <code class="language-plaintext highlighter-rouge">http://192.168.178.22:8188/system_stats</code> and sees the GPU: an RTX 3070 Ti with 8.6 GB of VRAM.</p>
  </li>
  <li>
    <p><strong>Pick a video model.</strong> The full video models like <a href="https://huggingface.co/Lightricks/LTX-2.3-fp8">LTX-2.3</a> or <a href="https://huggingface.co/Wan-AI/Wan2.1-T2V-1.3B">Wan 2.1</a> are huge (20-30 GB) and need more VRAM than 8 GB. The pragmatic choice was <a href="https://github.com/Kosinkadink/ComfyUI-AnimateDiff-Evolved">AnimateDiff</a>: a motion module that animates existing Stable Diffusion models. Small, fast, and it works with the <a href="https://civitai.com/models/4384/dreamshaper">DreamShaper</a> model I already had.</p>
  </li>
  <li>
    <p><strong>Download the motion module.</strong> A 1.7 GB file. This is where the setup got funny: ComfyUI’s AnimateDiff version 1.6 looks for motion modules in a folder called <code class="language-plaintext highlighter-rouge">animatediff_models</code>, not the <code class="language-plaintext highlighter-rouge">motion_modules</code> folder that older tutorials mention. ComfyUI actually created the correct folder itself on restart. A quick <code class="language-plaintext highlighter-rouge">move</code> command and one more restart later, the module appeared.</p>
  </li>
  <li>
    <p><strong>Build the workflow.</strong> This is the part I love. A ComfyUI workflow is just a JSON graph: nodes and connections. I wrote one from the Pi with six nodes: the model loader, the AnimateDiff loader, the prompt encoders, the sampler, the VAE decoder, and the video saver.</p>
  </li>
  <li>
    <p><strong>Submit and wait.</strong> One <code class="language-plaintext highlighter-rouge">curl</code> POST later, the queue on the Windows PC started. The 3070 Ti rendered 16 frames at 512x512, 20 steps each, in about a minute.</p>
  </li>
</ol>

<h2 id="the-result">The Result</h2>

<p>Here is the very first video my little cluster ever made. The prompt was “a cute cat walking through a sunlit park, cinematic lighting, high quality.”</p>

<blockquote>
  <p><strong>Note on sound:</strong> this video is silent. AnimateDiff is a motion module on top of an image model, so it has no audio path at all. The LTX-2.3 model I later switched to is a joint audio-video model and can generate sound, but that is a separate project.</p>
</blockquote>

<video controls="" loop="" muted="" playsinline="" width="100%" style="max-width:512px; border-radius:8px;">
  <source src="/assets/videos/cat-park-animatediff.mp4" type="video/mp4" />
  Your browser does not support the video tag.
</video>

<p>It is not Hollywood. The cat drifts more than it walks, and the park is more suggestion than scenery. It also has two tails, because of course it does. When a diffusion model does not know how many tails a cat should have, it simply gives it the average number of tails, rounded up. Happy little accidents, as the painter would say: the cat was never supposed to have two tails, and I would not change it now. But consider what just happened: a text prompt typed on a tiny Linux board in Berlin traveled over WiFi to a Windows PC, became a latent-space dream on an NVIDIA GPU, and came back as sixteen frames of a cat in a park. The whole loop took about a minute.</p>

<p><strong>The run, in numbers:</strong></p>

<table>
  <thead>
    <tr>
      <th>Setting</th>
      <th>Value</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Model</td>
      <td>DreamShaper + AnimateDiff v1.6</td>
    </tr>
    <tr>
      <td>Resolution</td>
      <td>512x512</td>
    </tr>
    <tr>
      <td>Frames</td>
      <td>16</td>
    </tr>
    <tr>
      <td>Steps</td>
      <td>20</td>
    </tr>
    <tr>
      <td>Render time</td>
      <td>~1 minute</td>
    </tr>
    <tr>
      <td>VRAM used</td>
      <td>8 GB (fits comfortably)</td>
    </tr>
  </tbody>
</table>

<h2 id="lessons-learned">Lessons Learned</h2>

<ul>
  <li><strong>VRAM is the currency.</strong> 8 GB runs AnimateDiff comfortably at 512x512, 16 frames. LTX-2.3 (22B params) also fits in fp8, but bigger clips and higher resolutions are where 8 GB hits its ceiling.</li>
  <li><strong>Folder names change between versions.</strong> The <code class="language-plaintext highlighter-rouge">animatediff_models</code> vs <code class="language-plaintext highlighter-rouge">motion_modules</code> confusion cost us two restarts. When a model is invisible, check the folder name the node actually scans.</li>
  <li><strong>Node names change too.</strong> <code class="language-plaintext highlighter-rouge">EmptySDLatentImage</code> is now <code class="language-plaintext highlighter-rouge">EmptyLatentImage</code>, and the AnimateDiff loader changed from the simple apply node to <code class="language-plaintext highlighter-rouge">ADE_AnimateDiffLoaderWithContext</code>, which returns a model the sampler accepts directly.</li>
  <li><strong>The webp-to-mp4 conversion is easier on the machine that rendered it.</strong> The Pi’s ffmpeg could not decode the animated webp, so I extracted the frames with Python and reassembled them.
    <h2 id="what-is-next">What Is Next</h2>
  </li>
</ul>

<p>The pipeline works. The Pi is now a remote control for a GPU that lives across the room. Next steps are tempting: longer clips, higher resolution, image-to-video, maybe that <a href="https://huggingface.co/Lightricks/LTX-2.3-fp8">LTX-2.3 model</a> I downloaded and never got to use properly. If you want the full story of how the Pi itself is set up, from SSD boot to auto-starting services, I wrote that up in <a href="https://erikzocher.github.io/technology/2026/08/02/raspberry-pi-technical-deep-dive.html">Raspberry Pi 5 From Scratch</a>.</p>

<p>But for now, I have a cat. A slightly wobbly, vaguely sunlit, entirely machine-made cat. And that is a good place to start.</p>

<h2 id="the-pattern-reusable">The Pattern, Reusable</h2>

<p>If you want to do the same thing, the shape is simple:</p>

<ol>
  <li><strong>Keep the brain and the muscle separate.</strong> Your always-on machine stays the brain; the GPU box is a dumb renderer you reach over the network.</li>
  <li><strong>Expose the muscle as an API, not a screen.</strong> ComfyUI’s REST endpoint turns “use a GPU” into a <code class="language-plaintext highlighter-rouge">curl</code> call. Anything that can speak HTTP can render.</li>
  <li><strong>Start with the smallest model that fits your VRAM.</strong> AnimateDiff on 8 GB beat the 20-30 GB video models I never could have run. The first win matters more than the ideal model.</li>
  <li><strong>Expect the setup to take longer than the generation.</strong> Folder names change, node names change, the first render is always an argument with the tool. The generation itself is the easy part.</li>
</ol>

<p>The whole pattern: Pi thinks, PC renders, LAN connects, cat wobbles. Happy little accidents included.</p>]]></content><author><name>Erik Zocher</name></author><category term="technology" /><category term="comfyui" /><category term="ai" /><category term="video" /><category term="raspberry-pi" /><category term="windows" /><summary type="html"><![CDATA[How I set up ComfyUI on a Windows PC with an RTX 3070 Ti and drove it remotely from my Raspberry Pi to generate my first AI video.]]></summary></entry><entry><title type="html">The Case of the Slain Gateway</title><link href="https://erikzocher.dev/technology/2026/08/02/the-case-of-the-slain-gateway.html" rel="alternate" type="text/html" title="The Case of the Slain Gateway" /><published>2026-08-02T14:00:00+02:00</published><updated>2026-08-02T14:00:00+02:00</updated><id>https://erikzocher.dev/technology/2026/08/02/the-case-of-the-slain-gateway</id><content type="html" xml:base="https://erikzocher.dev/technology/2026/08/02/the-case-of-the-slain-gateway.html"><![CDATA[<h1 id="the-case-of-the-slain-gateway">The Case of the Slain Gateway</h1>

<p><em>2026-08-02 · 8 min read · [raspberry-pi] [systemd] [debugging] [homelab] [storytelling]</em></p>

<p>Everything in this story actually happened. My headless Raspberry Pi went unreachable one Sunday evening, and debugging it turned out to be a genuine whodunit. Debugging a machine that goes silent has the same shape as a detective case: a victim, a scene, witnesses, conflicting clues, and a culprit who was there all along. Two enabled services that each think the other is an impostor, a crashed process that keeps reviving itself, a watchdog barking into the void. That is not an analogy, it is the plot. Sherlock Holmes is simply the clearest lens for it. The facts are untouched, only the telling is dramatized, and the technical <a href="#tldr-what-actually-happened">TL;DR is at the bottom</a> if you prefer facts over fog.</p>

<blockquote>
  <p><strong>AI disclosure:</strong> All illustrations in this story are AI-generated images (FLUX.1-dev), selected and edited by me. They were added later, imagining what the case would have looked like if Sherlock Holmes had worked with systemd and a soldering iron.</p>
</blockquote>

<hr />

<h2 id="chapter-i-the-house-at-19216817854">Chapter I: The House at 192.168.178.54</h2>

<p><img src="/assets/images/slain-gateway/chapter1-house.png" alt="The House at 192.168.178.54: a foggy Victorian street, the telegram boy at the door" /></p>

<p>It was a Sunday evening in August, and the fog had crept over the digital quarter of Berlin like a thief. In a modest house on the corner of the street called <code class="language-plaintext highlighter-rouge">192.168.178.54</code>, a machine had fallen silent.</p>

<p>The telegram boy had knocked twice that afternoon and received no answer. The housemaid, who attended to the doorways of the house (we called her SSH), had found the front entrance sealed as if by a spell. The master of the house, a modest but industrious fellow by the name of Gateway, was nowhere to be seen. His clock, which had ticked faithfully for one week and two days, had stopped.</p>

<p>I found my friend Sherlock Holmes in his study, pipe in hand, staring at a sheet of paper upon which someone had scrawled the words: <code class="language-plaintext highlighter-rouge">BOOT_ORDER=0xf146</code>.</p>

<p>“The game,” he said without turning, “is most certainly afoot.”</p>

<hr />

<h2 id="chapter-ii-the-note-on-the-doorstep">Chapter II: The Note on the Doorstep</h2>

<p><img src="/assets/images/slain-gateway/chapter2-note.png" alt="The Note on the Doorstep: Holmes examines two calling cards in the hearth ashes" /></p>

<p>Holmes had been called to the house by a distraught servant. The tale she told was this: the master had been unwell for weeks. He would rise each morning, light his lamps, and begin his rounds. Then, without warning, he would collapse. The servants would find him gasping the same words each time, like a man possessed:</p>

<p><em>“Gateway already running. Gateway already running. PID 1284. PID 1284.”</em></p>

<p>“Over and over,” the maid wept. “Twenty times. Thirty-five times, I counted.”</p>

<p>Holmes raised an eyebrow. “And this PID 1284. Was it a person? A rival?”</p>

<p>“That’s the devil of it, sir,” she said. “PID 1284 was the master himself.”</p>

<p>“A man cannot be running and not running at once,” I protested.</p>

<p>“Can he not, Watson?” Holmes smiled thinly. “Can he not?”</p>

<p>He knelt by the machine’s hearth and examined the ashes. There, among the cinders, lay two calling cards. One was engraved with a single word: <code class="language-plaintext highlighter-rouge">system.slice</code>. The other, smaller and older, bore the mark: <code class="language-plaintext highlighter-rouge">user.slice</code>.</p>

<p>“Two suitors,” Holmes murmured. “And both claim the hand of the same bride.”</p>

<hr />

<h2 id="chapter-iii-the-rivalry">Chapter III: The Rivalry</h2>

<p><img src="/assets/images/slain-gateway/chapter3-rivalry.png" alt="The Rivalry: a brass magnifying glass over two calling cards" /></p>

<p>Holmes produced a magnifying glass and held it over the two cards.</p>

<p>“Observe, Watson. The first card, <code class="language-plaintext highlighter-rouge">system.slice</code>, is clean and new. It was pressed upon the household at exactly twenty minutes past nine this evening, by a footman who answers to the name of root. It carries a curious instruction, quite modern: <code class="language-plaintext highlighter-rouge">--replace</code>. Whoever bears this card may, by its power, step into the shoes of any predecessor.”</p>

<p>“And the second card?”</p>

<p>“The second card is the intruder. It is dated the twentieth of July. It was slipped into the household ledger by a quieter servant, one who answers to the name of <code class="language-plaintext highlighter-rouge">systemd --user</code>. He does not announce himself at the front door. He waits in the pantry, in a file called <code class="language-plaintext highlighter-rouge">~/.config/systemd/user/hermes-gateway.service</code>, and at every dawn he sends forth his own claimant to the same position.”</p>

<p>“So there are two masters claiming one chair?”</p>

<p>“Exactly, Watson. And when two men claim one throne, one of them must die. Or rather,” he tapped the paper, “one of them must be told he is already dead.”</p>

<p>I confess I did not follow. But Holmes was already striding toward the machine’s great central chamber, the place the servants called the Process Table.</p>

<hr />

<h2 id="chapter-iv-the-ghost-in-the-machine">Chapter IV: The Ghost in the Machine</h2>

<p><img src="/assets/images/slain-gateway/chapter4-ghost.png" alt="The Ghost in the Machine: two claimants in a vast steampunk machine hall" /></p>

<p>We found them there, the two claimants, standing at opposite ends of the room.</p>

<p>The first was a hale and hearty fellow, PID 5371, born at twenty past nine, son of the system itself. He carried his <code class="language-plaintext highlighter-rouge">--replace</code> like a sabre. “I am the rightful master,” he declared. “I answer to root, and I answer to no other.”</p>

<p>The second was a gaunt and older figure, PID 5886, who had slipped in through the user’s pantry at eleven minutes past nine. He carried no such sabre. He carried only the memory of having been there first.</p>

<p>“Stand aside,” said the first.</p>

<p>“I was here before you,” said the second. “The ledger says my name. PID 1284, it says. And then you came, and the ledger could not decide which of us was real.”</p>

<p>Holmes studied them both, then turned to me.</p>

<p>“The crime, Watson, is not that one man killed another. The crime is that the ledger itself was corrupted. Each claimant looked into the book, saw the other’s name, and concluded he was dead. Yet both breathed. Both claimed the bride. And the bride, poor creature, could serve neither.”</p>

<p>“Then who,” I asked, “is the victim?”</p>

<p>“The victim,” said Holmes, “is the truth. And the murderer is the twentieth of July.”</p>

<hr />

<h2 id="chapter-v-the-hound-of-the-machine">Chapter V: The Hound of the Machine</h2>

<p><img src="/assets/images/slain-gateway/chapter5-hound.png" alt="The Hound of the Machine: Holmes and the brass clockwork watchdog" /></p>

<p>We descended into the basement, where a new servant had recently been hired. He was a large, loyal creature, and he answered to the name of Watchdog. His task was simple: every five minutes, he was to sniff at the master’s chambers and report whether the master yet lived.</p>

<p>“Good fellow,” said Holmes, “and what did you find?”</p>

<p>“I found,” said the hound, “that the master was not dead, and yet not alive. I found two masters where there should be one. I barked. I sent letters by post to the far address, <code class="language-plaintext highlighter-rouge">zocher.erik+blog@gmail.com</code>, for when the telegram boy is dead, one must write letters instead. I barked until my throat was raw, and still the household would not listen.”</p>

<p>“A faithful servant,” Holmes observed, “and a clever one. He understood the first law of such households: when the telegram is silent, you must cry through other means.”</p>

<p>“Indeed, sir,” said the hound. “But I could not heal the rift. I could only name it.”</p>

<p>“You named it well enough,” said Holmes. “You said there were two. And where there are two, one may be dispensed with.”</p>

<p>He turned to me, his eyes glittering.</p>

<p>“Come, Watson. We have found the ghost. Now we must find the root of all evil.”</p>

<hr />

<h2 id="chapter-vi-the-root-of-all-evil">Chapter VI: The Root of All Evil</h2>

<p><img src="/assets/images/slain-gateway/chapter6-root.png" alt="The Root of All Evil: Holmes holds the struck-through paper to the gas lamp" /></p>

<p>The root of all evil was not buried deep. It lay, as such things often do, in the most ordinary of places: a single file, unremarkable, dated the twentieth of July.</p>

<p>Holmes held it up to the lamp. <code class="language-plaintext highlighter-rouge">~/.config/systemd/user/hermes-gateway.service</code>.</p>

<p>“Here, Watson, is the confession. On the twentieth of July, someone in this household installed a second doorkeeper. They meant no harm. They wished only for the master to rise each morning. But they did not know that another doorkeeper, a grander one, already stood at the front gate with the key of root and the sabre of <code class="language-plaintext highlighter-rouge">--replace</code>.”</p>

<p>“So the second doorkeeper,” I said slowly, “was not a murderer. Merely… redundant.”</p>

<p>“Redundant is the kindest word for it. The crueler word is <em>conflict</em>. Two doorkeepers, each certain the other was an impostor. Each morning, the household would send both to the same post. Each would find the other’s coat upon the peg and cry, ‘Intruder!’ And then the house would fall silent, for a house with two masters is a house with none.”</p>

<p>“And the remedy?”</p>

<p>Holmes smiled. “The remedy is as old as Solomon. You do not divide the child. You dismiss one of the claimants.”</p>

<p>He drew a pen and struck a single line through the servant’s name.</p>

<p><code class="language-plaintext highlighter-rouge">systemctl --user disable hermes-gateway</code></p>

<p>“The user’s doorkeeper is retired,” he said. “He will not rise again. The system’s doorkeeper remains, and he carries the modern key, <code class="language-plaintext highlighter-rouge">--replace</code>, which grants him the power to step over any ghost that lingers. And the hound, the faithful Watchdog, shall keep his vigil, and send letters when the telegram falls silent.”</p>

<p>He closed his notebook and reached for his coat.</p>

<p>“The case is closed, Watson. The victim was not a man but a certainty. The murderer was not malice but duplication. And the ghost… the ghost was merely a master who had been told, once too often, that he was already dead.”</p>

<hr />

<h2 id="tldr-what-actually-happened">TL;DR: What Actually Happened</h2>

<p><strong>The setup:</strong> A Raspberry Pi 5 runs <a href="https://hermes-agent.nousresearch.com">Hermes Agent</a> as a <a href="https://systemd.io">systemd</a> service (<code class="language-plaintext highlighter-rouge">hermes-gateway.service</code>), connecting to Telegram and running cron jobs. It had run stably for about a week and a half before becoming unreachable via both SSH and Telegram (a full system-level hang, likely a hard reset; no clean shutdown was recorded in the journal).</p>

<p><strong>The real culprit: dual systemd units.</strong></p>

<ul>
  <li><strong>System unit:</strong> <code class="language-plaintext highlighter-rouge">/etc/systemd/system/hermes-gateway.service</code> (managed by root, runs in <code class="language-plaintext highlighter-rouge">system.slice</code>, PPID 1)</li>
  <li><strong>User unit:</strong> <code class="language-plaintext highlighter-rouge">~/.config/systemd/user/hermes-gateway.service</code> (created 20 July, managed by <code class="language-plaintext highlighter-rouge">systemd --user</code> PID 1219, runs in <code class="language-plaintext highlighter-rouge">user.slice</code>)</li>
</ul>

<p>Both units were enabled, so every boot spawned two gateway instances competing for the same PID lock file (<code class="language-plaintext highlighter-rouge">~/.hermes/gateway.pid</code>).</p>

<p><strong>The failure cascade:</strong></p>

<ol>
  <li>The user instance started first at boot (PID 1284, <code class="language-plaintext highlighter-rouge">user.slice</code>).</li>
  <li>Systemd’s system instance then tried to start a second gateway, and Hermes’ PID guard refused: <em>“Gateway already running (PID 1284)”</em> with exit code 1.</li>
  <li>A <code class="language-plaintext highlighter-rouge">Restart=on-failure</code> (from a user-added override that contradicted the unit’s <code class="language-plaintext highlighter-rouge">Restart=always</code>) caused rapid restart attempts, producing a crash-loop. The restart counter reached 35.</li>
  <li>systemd eventually reported <em>“more than one ExecStart= setting… bad unit file setting”</em> because a drop-in <code class="language-plaintext highlighter-rouge">override.conf</code> had been written with a second <code class="language-plaintext highlighter-rouge">ExecStart=</code> line without resetting the first. That is a classic systemd drop-in merge footgun. The unit went <code class="language-plaintext highlighter-rouge">failed</code>.</li>
</ol>

<p><strong>Fixes applied:</strong></p>

<ul>
  <li><strong><code class="language-plaintext highlighter-rouge">--replace</code> flag</strong> added to the gateway’s ExecStart (the official “useful for systemd” option). A fresh start now auto-replaces any stale instance instead of dying on the PID lock.</li>
  <li><strong>Override repaired:</strong> <code class="language-plaintext highlighter-rouge">ExecStart=</code> (empty, resets the main unit’s value) followed by the new <code class="language-plaintext highlighter-rouge">ExecStart=</code>. Required because systemd merges drop-ins, so a bare second <code class="language-plaintext highlighter-rouge">ExecStart=</code> would add, not replace.</li>
  <li><strong>User unit disabled</strong> via <code class="language-plaintext highlighter-rouge">systemctl --user disable hermes-gateway</code>, removing the duplicate doorkeeper. Only the system unit starts at boot now.</li>
  <li><strong>Watchdog script</strong> (<code class="language-plaintext highlighter-rouge">gateway_watchdog.sh</code>, cron every 5 minutes) checks that <code class="language-plaintext highlighter-rouge">ActiveState/SubState</code> is active/running and that exactly one gateway process exists. On failure it sends an email via <a href="https://github.com/pimalaya/himalaya">himalaya</a>, a fallback channel, since Telegram is dead when the gateway is. Alert spam is prevented with rate-limiting.</li>
</ul>

<p><strong>Root cause in one line:</strong> Two enabled systemd units (system + user) fought over one PID file; a broken drop-in override then locked the service in a crash-loop. Adding <code class="language-plaintext highlighter-rouge">--replace</code>, disabling the user unit, and installing a mail-capable watchdog made the whole thing self-healing.</p>

<hr />

<p><em>License: illustrations generated with FLUX.1-dev, free for non-commercial use; this blog qualifies. Check the license before any commercial use.</em></p>]]></content><author><name>Erik Zocher</name></author><category term="technology" /><category term="raspberry-pi" /><category term="systemd" /><category term="debugging" /><category term="homelab" /><category term="storytelling" /><summary type="html"><![CDATA[A true story of a headless Raspberry Pi that went silent, told as a Sherlock Holmes mystery: two systemd units, one PID file, a ghost process, and the root of all evil.]]></summary></entry><entry><title type="html">Raspberry Pi 5 From Scratch: SSD Boot, Auto-Starting Services, and Remote Access</title><link href="https://erikzocher.dev/technology/2026/08/02/raspberry-pi-technical-deep-dive.html" rel="alternate" type="text/html" title="Raspberry Pi 5 From Scratch: SSD Boot, Auto-Starting Services, and Remote Access" /><published>2026-08-02T11:00:00+02:00</published><updated>2026-08-02T11:00:00+02:00</updated><id>https://erikzocher.dev/technology/2026/08/02/raspberry-pi-technical-deep-dive</id><content type="html" xml:base="https://erikzocher.dev/technology/2026/08/02/raspberry-pi-technical-deep-dive.html"><![CDATA[<h1 id="raspberry-pi-5-from-scratch-ssd-boot-auto-starting-services-and-remote-access">Raspberry Pi 5 From Scratch: SSD Boot, Auto-Starting Services, and Remote Access</h1>

<p><em>2026-08-02 · 10 min read · [raspberry-pi] [nvme] [systemd] [homelab] [linux] [headless]</em></p>

<p>This is the technical companion to my <a href="https://erikzocher.github.io/technology/2026/08/02/raspberry-pi-ai-agent.html">Raspberry Pi AI agent overview</a>. That post covers the hardware and the big picture. This one covers the whole machine: from flashing the SSD and setting up WiFi before the first boot, to booting from NVMe, auto-starting services, and getting a GUI when I need one.</p>

<h2 id="1-first-boot-flashing-the-ssd-and-pre-setting-wifi">1. First Boot: Flashing the SSD and Pre-Setting WiFi</h2>

<p>Before any customization, the Pi needs an operating system, and for a headless setup the trick is to configure everything <strong>before</strong> the first boot, so you never need a monitor or keyboard.</p>

<h3 id="flash-the-os-directly-to-the-ssd">Flash the OS directly to the SSD</h3>

<p><a href="https://www.raspberrypi.com/software/">Raspberry Pi Imager</a> is the official tool and it does the whole job in one go. It runs on Windows, macOS, and Linux, and it can write the OS straight to the NVMe drive.</p>

<ol>
  <li><strong>Connect the SSD to your computer.</strong> An NVMe-to-USB adapter or a small USB enclosure is all you need. The drive shows up like a big USB stick.</li>
  <li><strong>Open Raspberry Pi Imager</strong> and click <em>Choose Device</em> → Raspberry Pi 5.</li>
  <li><strong>Choose OS</strong> → Raspberry Pi OS (64-bit) or Raspberry Pi OS Lite if you want no desktop at all.</li>
  <li><strong>Choose Storage</strong> → pick the NVMe SSD, not your computer’s own disk!</li>
  <li><strong>Click the gear icon</strong> (or press Ctrl+Shift+X) to open the advanced options. This is the headless magic:
    <ul>
      <li><strong>Enable SSH</strong> and set it to allow password login (or paste an SSH key for key-only access).</li>
      <li><strong>Set the username and password</strong> you want to log in with.</li>
      <li><strong>Configure WiFi</strong>: enter the SSID and password. There is also a <strong>country dropdown</strong> (defaults to GB), and it matters more than it looks: set it to your actual country, e.g. <code class="language-plaintext highlighter-rouge">DE</code> for Germany. The country defines the <em>regulatory domain</em>, which tells the kernel which WiFi channels and transmit power levels are legal. With the wrong country you can lose the 5GHz channels entirely; with none set, the radio falls back to conservative defaults. The Imager writes all of this into the image, so the Pi connects to your network on its very first boot, with no keyboard needed.</li>
      <li>Optionally set a <strong>hostname</strong> (mine is a plain <code class="language-plaintext highlighter-rouge">ezocher</code> box, but something like <code class="language-plaintext highlighter-rouge">piagent</code> is nicer) and your timezone.</li>
    </ul>
  </li>
  <li><strong>Click Write</strong> and wait. The Imager flashes the OS, then verifies the write.</li>
  <li><strong>Unplug the SSD from the computer</strong>, mount it on the Pi with the M.2 HAT+, connect power, and wait about a minute.</li>
</ol>

<p>That is the whole first-boot ritual: no monitor, no keyboard, no HDMI cable. The Pi appears on your WiFi, reachable by SSH, with the OS already on the NVMe drive.</p>

<p>If you are curious what the Imager actually did, it wrote a file called <code class="language-plaintext highlighter-rouge">userconfig.txt</code> (or a first-run script) onto the boot partition of the SSD with your SSH and WiFi settings. That file is read once on first boot and then consumed. You can pre-seed the same settings manually, but the Imager dialog is less error-prone.</p>

<h2 id="2-booting-from-the-nvme-ssd">2. Booting From the NVMe SSD</h2>

<p>The Pi 5 does not boot from the microSD in my setup. It boots from the NVMe drive, and there are three pieces to that: the EEPROM boot order, the root filesystem in fstab, and the PCIe speed.</p>

<h3 id="eeprom-boot-order">EEPROM boot order</h3>

<p>The Pi 5’s boot order lives in its EEPROM. Mine reads:</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nv">$ </span><span class="nb">sudo </span>rpi-eeprom-config | <span class="nb">grep </span>BOOT_ORDER
<span class="nv">BOOT_ORDER</span><span class="o">=</span>0xf146
</code></pre></div></div>

<p>That hex value is read from right to left, each digit a boot source:</p>

<table>
  <thead>
    <tr>
      <th>Digit</th>
      <th>Meaning</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">6</code></td>
      <td>Restart from the top of the list</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">4</code></td>
      <td>USB mass storage (where the NVMe appears)</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">1</code></td>
      <td>SD card (fallback)</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">f</code></td>
      <td>Loop back to the beginning</td>
    </tr>
  </tbody>
</table>

<p>So the Pi tries the NVMe first, falls back to the SD card, and if neither works it keeps cycling. To change it, either use the menu:</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nb">sudo </span>raspi-config   <span class="c"># Advanced Options → Boot Order → NVMe/USB Boot</span>
</code></pre></div></div>

<p>or edit the EEPROM directly:</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nb">sudo </span>rpi-eeprom-config <span class="nt">--edit</span>
<span class="c"># change BOOT_ORDER to 0xf146, save, reboot</span>
</code></pre></div></div>

<h3 id="root-filesystem-by-partuuid">Root filesystem by PARTUUID</h3>

<p>The kernel finds the root partition by its PARTUUID, not by device name. That matters because device names (<code class="language-plaintext highlighter-rouge">/dev/sda</code>, <code class="language-plaintext highlighter-rouge">/dev/mmcblk0</code>) can change between boots, while PARTUUIDs are burned into the partition table.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nv">$ </span><span class="nb">cat</span> /etc/fstab
<span class="nv">PARTUUID</span><span class="o">=</span>957b593b-02  /               ext4    defaults,noatime  0  1
<span class="nv">PARTUUID</span><span class="o">=</span>957b593b-01  /boot/firmware  vfat    defaults          0  2
<span class="nv">UUID</span><span class="o">=</span>76602090-4e70-4a14-a6c2-ffd91561de93 /mnt/nvme ext4 defaults,auto,users,rw,nofail 0 0
</code></pre></div></div>

<p>The <code class="language-plaintext highlighter-rouge">957b593b</code> PARTUUIDs belong to the NVMe partition table. You can confirm which disk they point at:</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nv">$ </span>lsblk <span class="nt">-o</span> NAME,PARTUUID,MOUNTPOINT
nvme0n1
├─nvme0n1p1  957b593b-01  /boot/firmware
└─nvme0n1p2  957b593b-02  /
</code></pre></div></div>

<p>The old microSD is demoted to a plain data disk, mounted at <code class="language-plaintext highlighter-rouge">/media/ezocher/bootfs</code>. The <code class="language-plaintext highlighter-rouge">nofail</code> option on the <code class="language-plaintext highlighter-rouge">/mnt/nvme</code> mount is important: it lets the system boot even if that drive is missing, instead of hanging at a mount prompt.</p>

<h3 id="pcie-gen-3-for-full-nvme-speed">PCIe Gen 3 for full NVMe speed</h3>

<p>The Pi 5 runs its PCIe bus at Gen 2 by default, which caps the NVMe throughput. Bumping to Gen 3 is one line in <code class="language-plaintext highlighter-rouge">/boot/firmware/config.txt</code>:</p>

<div class="language-ini highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="py">dtparam</span><span class="p">=</span><span class="s">pciex1_gen=3</span>
</code></pre></div></div>

<p>Reboot and <code class="language-plaintext highlighter-rouge">sudo dmesg | grep nvme</code> should show the drive negotiating at Gen 3 speeds.</p>

<h2 id="3-auto-starting-services-with-systemd">3. Auto-Starting Services With systemd</h2>

<p>The “turn it on and it just works” part is systemd. Three services run on my Pi, all enabled at boot and all set to restart on failure:</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nv">$ </span><span class="nb">ls</span> /etc/systemd/system/<span class="k">*</span>.service
hermes-gateway.service    <span class="c"># the AI agent</span>
ollama.service            <span class="c"># local model serving (legacy, see the overview)</span>
kindle-dashboard.service  <span class="c"># serves the e-ink dashboard image</span>
</code></pre></div></div>

<p>The agent service shows the pattern:</p>

<div class="language-ini highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nn">[Unit]</span>
<span class="py">Description</span><span class="p">=</span><span class="s">Hermes Agent Gateway</span>
<span class="py">After</span><span class="p">=</span><span class="s">network-online.target</span>
<span class="py">Wants</span><span class="p">=</span><span class="s">network-online.target</span>

<span class="nn">[Service]</span>
<span class="py">Type</span><span class="p">=</span><span class="s">simple</span>
<span class="py">User</span><span class="p">=</span><span class="s">ezocher</span>
<span class="py">ExecStart</span><span class="p">=</span><span class="s">/home/ezocher/.hermes/hermes-agent/venv/bin/python -m hermes_cli.main gateway run</span>
<span class="py">Restart</span><span class="p">=</span><span class="s">always</span>
<span class="py">RestartSec</span><span class="p">=</span><span class="s">5</span>

<span class="nn">[Install]</span>
<span class="py">WantedBy</span><span class="p">=</span><span class="s">multi-user.target</span>
</code></pre></div></div>

<p>Three details worth copying:</p>

<ul>
  <li><strong><code class="language-plaintext highlighter-rouge">After=network-online.target</code></strong> with <strong><code class="language-plaintext highlighter-rouge">Wants=</code></strong> waits until the network is actually up, so the agent does not start half-connected.</li>
  <li><strong><code class="language-plaintext highlighter-rouge">Restart=always</code></strong> with <strong><code class="language-plaintext highlighter-rouge">RestartSec=5</code></strong> brings the service back within seconds of any crash. This is the real resilience trick for a headless box.</li>
  <li><strong><code class="language-plaintext highlighter-rouge">WantedBy=multi-user.target</code></strong> starts the service at boot without needing anyone to log in.</li>
</ul>

<p>Enable any service with:</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nb">sudo </span>systemctl <span class="nb">enable</span> <span class="nt">--now</span> &lt;service-name&gt;
</code></pre></div></div>

<p>The result: after a power cut, the kernel finds the NVMe, mounts the root partition, and systemd brings the whole stack up without a human in the loop. Power on, wait a minute, message the agent on Telegram. That is the whole ritual.</p>

<h2 id="4-raspberry-pi-connect-for-gui-access">4. Raspberry Pi Connect for GUI Access</h2>

<p>The Pi runs headless, but sometimes you need a graphical interface. <strong>Raspberry Pi Connect</strong> is the Foundation’s own remote access service: free, encrypted, and it works from any browser with no port forwarding.</p>

<p><strong>Install it</strong> (on recent Raspberry Pi OS images it is already there, so check first):</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>rpi-connect <span class="nt">--help</span>
</code></pre></div></div>

<p>If that says “command not found”:</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nb">sudo </span>apt update <span class="o">&amp;&amp;</span> <span class="nb">sudo </span>apt <span class="nb">install </span>rpi-connect
</code></pre></div></div>

<p><strong>Link the Pi to your account:</strong></p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>rpi-connect signin
</code></pre></div></div>

<p>It prints a URL. Open it in any browser, log in with your Raspberry Pi ID, and the Pi is linked.</p>

<p><strong>Enable the service</strong> so it survives reboots:</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nb">sudo </span>systemctl <span class="nb">enable</span> <span class="nt">--now</span> rpi-connect
</code></pre></div></div>

<p><strong>Connect:</strong> go to <a href="https://connect.raspberrypi.com">connect.raspberrypi.com</a>, log in, and your Pi appears in the list. Click it and you get the desktop in a browser tab, or a shell. No port forwarding, no VPN, no dynamic DNS, and it works from outside your home network too.</p>

<p>Useful commands while you are at it: <code class="language-plaintext highlighter-rouge">rpi-connect status</code> shows the connection state, <code class="language-plaintext highlighter-rouge">rpi-connect restart</code> fixes a stuck client, and <code class="language-plaintext highlighter-rouge">rpi-connect doctor</code> runs diagnostics.</p>

<h2 id="putting-it-all-together">Putting It All Together</h2>

<p>The full boot story is: EEPROM says “try NVMe first”, fstab points the root filesystem at the NVMe partition, config.txt unlocks Gen 3 PCIe, and systemd starts the agent, the dashboard server, and Connect automatically. That is the difference between a Pi you have to poke at and a Pi that just runs.</p>

<p><em>This is the technical half of the story. For the hardware list, the agent setup, and why I use a cloud model instead of a local one, see the <a href="https://erikzocher.github.io/technology/2026/08/02/raspberry-pi-ai-agent.html">overview post</a>.</em></p>]]></content><author><name>Erik Zocher</name></author><category term="technology" /><category term="raspberry-pi" /><category term="nvme" /><category term="systemd" /><category term="homelab" /><category term="linux" /><category term="headless" /><summary type="html"><![CDATA[The technical details behind my headless Pi AI agent: flashing the SSD with Raspberry Pi Imager, pre-setting WiFi and SSH, EEPROM boot order, PARTUUID root filesystem, PCIe Gen 3, systemd services, and Raspberry Pi Connect.]]></summary></entry><entry><title type="html">My Raspberry Pi AI Agent: The Full Setup</title><link href="https://erikzocher.dev/technology/2026/08/02/raspberry-pi-ai-agent.html" rel="alternate" type="text/html" title="My Raspberry Pi AI Agent: The Full Setup" /><published>2026-08-02T09:00:00+02:00</published><updated>2026-08-02T09:00:00+02:00</updated><id>https://erikzocher.dev/technology/2026/08/02/raspberry-pi-ai-agent</id><content type="html" xml:base="https://erikzocher.dev/technology/2026/08/02/raspberry-pi-ai-agent.html"><![CDATA[<h1 id="my-raspberry-pi-ai-agent-the-full-setup">My Raspberry Pi AI Agent: The Full Setup</h1>

<p><em>2026-08-02 · 6 min read · [raspberry-pi] [hermes] [ai-agents] [homelab] [nvme] [telegram]</em></p>

<p>I run a personal AI agent from a Raspberry Pi 5 that sits in my living room in Berlin. It watches the tram schedule, updates an e-ink dashboard on my wall, scans job listings, drafts blog posts, and talks to me through Telegram. This post is the overview. If you want the step-by-step technical details, I wrote those up in a separate post: <a href="https://erikzocher.github.io/technology/2026/08/02/raspberry-pi-technical-deep-dive.html">Raspberry Pi 5 From Scratch</a>.</p>

<p><strong>The 30-second version:</strong> a Pi 5 with 16 GB RAM, booting from a 1 TB NVMe SSD, runs an open source agent framework (Hermes) as an always-on daemon. The model itself lives in the cloud (DeepSeek), because local models on the Pi could not do tool calls reliably, and tool calls are what separate a chat from an agent. You talk to it over Telegram, fix it over SSH, and it runs a small fleet of scheduled jobs: trams, dashboards, job scans, blog drafts. Total cost: hardware once, single-digit euros per month for the model API, a few watts of power.</p>

<h2 id="the-hardware">The Hardware</h2>

<p><strong>Raspberry Pi 5 Model B (Rev 1.1)</strong> with <strong>16GB RAM</strong>. That matters: the Pi 5 is the first Pi where 16GB is an option, and for an always-on agent, the extra headroom is worth it.</p>

<p>The star of the show is the storage: a <strong>1TB NVMe SSD</strong> (BIWIN CE930) connected through the official <strong>Raspberry Pi M.2 HAT+</strong>. Booting from NVMe instead of an SD card changes everything. No more worrying about SD card corruption, which is the classic way a headless Pi dies. Reads and writes are fast enough that the agent never waits on disk.</p>

<p>A 238GB microSD is still in the slot from the old setup, but the system boots from the NVMe drive now. How I made that work is in the <a href="https://erikzocher.github.io/technology/2026/08/02/raspberry-pi-technical-deep-dive.html">technical post</a>.</p>

<h2 id="the-os">The OS</h2>

<p><strong>Raspberry Pi OS, based on Debian 13 (Trixie)</strong>, kernel 6.18. Running headless, no monitor, no keyboard. The only cables are power and network. Everything else happens over SSH and Telegram.</p>

<h2 id="the-agent-hermes">The Agent: Hermes</h2>

<p>The agent itself is <a href="https://hermes-agent.nousresearch.com">Hermes Agent</a> by Nous Research. It is an open source AI agent framework that runs as a daemon on the Pi, and it is the whole point of this box. Instead of a chatbot that answers questions, Hermes is an agent that takes actions: it runs shell commands, reads and writes files, searches the web, manages scheduled jobs, and remembers things across sessions.</p>

<h2 id="why-a-cloud-model-not-a-local-one">Why a Cloud Model, Not a Local One</h2>

<p>I used to run local models on the Pi, served by Ollama. I tried four of them over a few days:</p>

<table>
  <thead>
    <tr>
      <th>Model</th>
      <th style="text-align: center">Parameters</th>
      <th style="text-align: center">File size</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Gemma 4 (e2b)</td>
      <td style="text-align: center">5.1B</td>
      <td style="text-align: center">7.2 GB</td>
    </tr>
    <tr>
      <td>Gemma 4 (64k context)</td>
      <td style="text-align: center">5.1B</td>
      <td style="text-align: center">7.2 GB</td>
    </tr>
    <tr>
      <td>Gemma 4 (e4b)</td>
      <td style="text-align: center">8.0B</td>
      <td style="text-align: center">9.6 GB</td>
    </tr>
    <tr>
      <td>Qwen 3.5 (9b)</td>
      <td style="text-align: center">9.7B</td>
      <td style="text-align: center">6.6 GB</td>
    </tr>
  </tbody>
</table>

<p>The size mattered, and the reason is the Pi’s <strong>16GB of RAM</strong>. A model has to fit entirely in memory while the operating system, the agent, and everything else still need their share. The 5.1B models fit comfortably. The 8B and 9.7B models worked too, but they left much less headroom, and inference on CPU was slow: with no GPU, every token comes from the CPU, and the bigger the model, the longer you wait between words.</p>

<p>But the real reason I moved to the cloud was not speed or RAM. It was this: <strong>none of them could do tool calls reliably.</strong></p>

<p>Tool calls are the difference between a chat and an agent. The model has to decide “I need the tram times” and produce a structured request to fetch them. Small models running on a Pi simply do not have the capacity for that reliably. They answer questions, but they cannot operate a harness. A local model on this hardware was a fun experiment and a dead end for real agent work.</p>

<p>So I switched to <strong>DeepSeek v4 Flash</strong> through the DeepSeek API. It is cheap, fast, and capable enough for tool use. The Pi sends requests to the API, the model decides what to do, and the Pi executes it. For this workload, cloud is the right call: the model lives in the cloud, the agent lives on the Pi, and the Pi stays quiet and cool.</p>

<pre><code class="language-mermaid">flowchart LR
    YOU["You (phone / laptop)"] --&gt;|"Telegram, everyday"| PI["Raspberry Pi 5&lt;br/&gt;Hermes agent&lt;br/&gt;16 GB, NVMe boot"]
    YOU --&gt;|"SSH, serious work"| PI
    PI --&gt;|"tool calls"| API["DeepSeek v4 Flash&lt;br/&gt;(cloud model)"]
    API --&gt;|"decisions"| PI
    PI --&gt;|"fetch data"| EXT["Trams, weather,&lt;br/&gt;job boards, events"]
    PI --&gt;|"render"| KIND["e-ink dashboard&lt;br/&gt;on the wall"]
</code></pre>

<h2 id="how-i-interact-with-it">How I Interact With It</h2>

<p>Two front doors, both headless:</p>

<p><strong>SSH</strong> for everything serious. <code class="language-plaintext highlighter-rouge">ssh ezocher@192.168.178.54</code> gets me a shell, and from there I can reach the agent’s files, logs, and configuration. When something breaks, this is where I look.</p>

<p><strong>Telegram</strong> for everything daily. The agent is connected to Telegram as a bot, so I can message it like a friend. Ask for the tram, tell it to draft a blog post, ask it to check a job listing. It replies in the same chat. This is the interface that makes the whole thing feel alive. I check in from my phone, anywhere.</p>

<p>For the rare moments when I need a graphical interface, I use <strong>Raspberry Pi Connect</strong>, the Foundation’s own remote desktop service. It works from any browser, no port forwarding needed. The setup steps are in the <a href="https://erikzocher.github.io/technology/2026/08/02/raspberry-pi-technical-deep-dive.html">technical post</a>.</p>

<h2 id="the-skills-and-plugins-that-make-it-useful">The Skills and Plugins That Make It Useful</h2>

<p>Hermes organizes capabilities into skills, which are like plugins for the agent. These are the ones I actually use:</p>

<ul>
  <li><strong>BVG tram departures</strong>: checks the real tram times for my station (Buschallee). The most-used skill, honestly.</li>
  <li><strong>Kindle dashboard</strong>: generates an e-ink dashboard image for an old Kindle Touch I mounted on the wall. Shows trams, weather, todos, and a Spanish word of the day.</li>
  <li><strong>Blog pipeline</strong>: drafts posts and opens GitHub pull requests for me to review before anything goes live. This very post was drafted this way.</li>
  <li><strong>Berlin events</strong>: scrapes event listings (rausgegangen.de and friends) for things to do.</li>
  <li><strong>Design job scan</strong>: a weekly cron job that scans German design job boards. (For a friend, not me. I am a software engineer, not a designer.)</li>
  <li><strong>Humanize writing</strong>: a checklist that keeps AI text from sounding like AI text.</li>
  <li><strong>Weather</strong>: Open-Meteo, no API key needed.</li>
  <li><strong>Spanish tutor</strong>: vocabulary and grammar drills, because I am learning Spanish.</li>
</ul>

<p>The skills run on a schedule (cron) or on demand when I ask.</p>

<h2 id="what-runs-automatically">What Runs Automatically</h2>

<p>The Pi is a small fleet of always-on jobs:</p>

<ul>
  <li>The Kindle wall dashboard refreshes every two minutes with trams, weather, and todos.</li>
  <li>A daily todo review at 20:00.</li>
  <li>A design job scan every Monday morning.</li>
  <li>Blog drafts that turn into pull requests for review.</li>
  <li>A real-time tram departure watcher.</li>
</ul>

<p>Uptime at the time of writing: one week, two days, and counting. It just sits there and does its job.</p>

<h2 id="the-costs">The Costs</h2>

<ul>
  <li><strong>Hardware</strong>: Pi 5 16GB, NVMe HAT + 1TB SSD. One-time, roughly the price of a mid-range phone.</li>
  <li><strong>Power</strong>: a few watts. Negligible on the electricity bill.</li>
  <li><strong>Model API</strong>: DeepSeek v4 Flash is cheap. My usage costs single-digit euros per month.</li>
  <li><strong>Everything else</strong>: free. Open source agent, free weather API, free job boards, Telegram bot is free.</li>
</ul>

<h2 id="the-pattern-reusable">The Pattern, Reusable</h2>

<p>The key decision was separating the model from the machine. The model is in the cloud where the compute is, the machine is at home where the actions are. If you want to build the same thing, the shape is simple:</p>

<ol>
  <li><strong>Put the agent where the actions are, the model where the compute is.</strong> A Pi 5 is a terrible GPU server but a fantastic always-on host: silent, cheap, and strong enough to run the harness, the schedules, and the memory.</li>
  <li><strong>Pick the front door for the audience.</strong> Telegram for daily use, SSH for serious work, remote desktop for the rare graphical moment. Three interfaces, one agent.</li>
  <li><strong>Test local models before committing to them.</strong> I ran four Ollama models before concluding the bottleneck was tool calls, not speed. Small models answer questions; they struggle to operate a harness.</li>
  <li><strong>Let the scheduled jobs do the proving.</strong> A fleet of small cron jobs (trams, dashboard, job scans) turns the agent from a toy into something you rely on daily, and it costs nothing to run.</li>
</ol>

<p>If you want a personal AI agent that actually does things instead of just talking, a Pi 5 with an NVMe drive and a cloud model is a really good place to start.</p>

<p><em>Want the technical how-to? Read <a href="https://erikzocher.github.io/technology/2026/08/02/raspberry-pi-technical-deep-dive.html">Raspberry Pi 5 From Scratch</a>.</em></p>]]></content><author><name>Erik Zocher</name></author><category term="technology" /><category term="raspberry-pi" /><category term="hermes" /><category term="ai-agents" /><category term="homelab" /><category term="nvme" /><category term="telegram" /><summary type="html"><![CDATA[How I built a headless AI agent on a Raspberry Pi 5 with an NVMe SSD, DeepSeek in the cloud, and Telegram as the front door.]]></summary></entry><entry><title type="html">Working with AI: Five Ways, From Prompts to Agent Graphs</title><link href="https://erikzocher.dev/technology/2026/07/31/working-with-ai.html" rel="alternate" type="text/html" title="Working with AI: Five Ways, From Prompts to Agent Graphs" /><published>2026-07-31T00:00:00+02:00</published><updated>2026-07-31T00:00:00+02:00</updated><id>https://erikzocher.dev/technology/2026/07/31/working-with-ai</id><content type="html" xml:base="https://erikzocher.dev/technology/2026/07/31/working-with-ai.html"><![CDATA[<h1 id="working-with-ai-five-ways-from-prompts-to-agent-graphs">Working with AI: Five Ways, From Prompts to Agent Graphs</h1>

<p><em>2026-07-31 · 6 min read · [ai] [prompt-engineering] [agents] [workflows] [beginner]</em></p>

<p>Most people think working with AI means typing a good prompt. That’s level one. There are at least five levels of working with AI, and each one gives you more control than the last. Here they are, from the simplest to the most powerful.</p>

<p>Before we start: this is a moving target. The way we work with AI keeps evolving, and new approaches appear faster than anyone can write about them. Treat this list as a snapshot of where things stand now, not a complete map. The next level might already be on its way.</p>

<h2 id="level-1-prompt-engineering">Level 1: Prompt Engineering</h2>

<p><strong>What it is:</strong> the craft of asking well. The words you type into the box.</p>

<p>A prompt is more than a question. It carries context, constraints, examples, and a desired format. Compare these two:</p>

<blockquote>
  <p>“Write a polite refusal.”</p>
</blockquote>

<blockquote>
  <p>“Write a polite refusal to a client who just cancelled their contract. Three sentences. No apology for the delay, thank them for the years of collaboration.”</p>
</blockquote>

<p>Same intent, very different results. The second prompt tells the model who the audience is, how long the answer should be, and what to leave out.</p>

<p><strong>What it gets you:</strong> better answers from a single request. That’s it. No memory, no tools, no follow-up. One shot, one answer.</p>

<pre><code class="language-mermaid">flowchart LR
    A[You] --&gt;|prompt| B[LLM]
    B --&gt;|answer| C[You]
</code></pre>

<h2 id="level-2-context-engineering">Level 2: Context Engineering</h2>

<p><strong>What it is:</strong> controlling everything the model sees before it answers. The system prompt, the instructions, the documents, the conversation history, the search results. All of these are context.</p>

<p>Prompt engineering is asking the right question. Context engineering is deciding what’s in the room when you ask it.</p>

<p>There are many ways to get context in. RAG is the most famous, but it’s only one:</p>

<ul>
  <li><strong>RAG (retrieval-augmented generation):</strong> the model looks up your notes, documentation, or database and answers from what it finds. Great for grounding answers in your own knowledge base.</li>
  <li><strong>IDE integration:</strong> the model sees the actual code you’re working on. Tools like Cursor, GitHub Copilot, or Claude Code run alongside your editor, so the context is the project itself, not a summary of it. You don’t have to explain what your codebase looks like, the model already sees it.</li>
  <li><strong>Links and URLs:</strong> give the model a documentation page, an API reference, or an article URL, and it reads the content before answering. Handy when the answer lives on the web and the model’s training data is outdated.</li>
  <li><strong>System prompts and memory:</strong> standing instructions plus what you’ve discussed before. The model remembers the rules and the history you’ve set up.</li>
</ul>

<p><strong>What it gets you:</strong> grounded, consistent answers that use the right source, whether that’s your database, your codebase, or a webpage. This is how customer support bots stop hallucinating product details: they read the manual first.</p>

<pre><code class="language-mermaid">flowchart LR
    A[You] --&gt;|prompt| B[LLM]
    D[System prompt + memory] --&gt; B
    E[Your docs via RAG] --&gt; B
    F[IDE: your open code] --&gt; B
    G[Linked documentation] --&gt; B
    B --&gt;|grounded answer| C[You]
</code></pre>

<h2 id="level-3-harness-engineering">Level 3: Harness Engineering</h2>

<p><strong>What it is:</strong> giving the AI hands. Tools, APIs, and actions it can call. Instead of just answering, it can do things.</p>

<p>Think of the model as a brain. The harness is the body you build around it: which tools exist, what they can touch, and what the guardrails are. This is where reliability comes from. You decide the boundaries, not the model.</p>

<p>Examples of harness pieces:</p>

<ul>
  <li>Web search</li>
  <li>Database queries</li>
  <li>Sending emails</li>
  <li>Running code</li>
  <li>Reading and writing files</li>
  <li>Calling other APIs</li>
</ul>

<p>My own setup lives at this level. My agent (<a href="https://erikzocher.github.io/technology/2026/07/28/meet-puck.html">Puck</a>, yes I named it) has a terminal, a browser, file access, and messaging. It can check the BVG schedule, generate a dashboard image, push a blog post to GitHub, and send me the result on Telegram. That’s a harness around a model.</p>

<p><strong>What it gets you:</strong> AI that acts, not just talks.</p>

<pre><code class="language-mermaid">flowchart LR
    A[You] --&gt;|task| B[Agent]
    B --&gt;|tool call| C[Web search]
    B --&gt;|tool call| D[Database]
    B --&gt;|tool call| E[Run code]
    B --&gt;|tool call| F[Send email]
    C --&gt;|result| B
    D --&gt;|result| B
    E --&gt;|result| B
    F --&gt;|result| B
    B --&gt;|result| G[You]
</code></pre>

<h2 id="level-4-loop-engineering">Level 4: Loop Engineering</h2>

<p><strong>What it is:</strong> building feedback cycles. The AI acts, observes the result, adjusts, and repeats. Instead of a single shot, you get a loop.</p>

<p>A chef tastes the sauce and adjusts it. Single-shot prompting is cooking without tasting. Loop engineering is the tasting step.</p>

<p>The most common pattern is plan-act-observe: the AI makes a plan, takes a step, looks at what happened, and decides what to do next. It’s also called ReAct (reason + act). Related patterns:</p>

<ul>
  <li>Self-correction: the AI reviews its own output and fixes mistakes</li>
  <li>Retry with more info: when a tool call fails, feed the error back and try again</li>
  <li>Evaluation loops: a second AI checks the first one’s work before it ships</li>
</ul>

<p><strong>What it gets you:</strong> much harder problems solved, because the system can course-correct instead of hoping the first try was right.</p>

<pre><code class="language-mermaid">flowchart LR
    A[Plan] --&gt; B[Act]
    B --&gt; C[Observe result]
    C --&gt;|needs adjustment| A
    C --&gt;|goal reached| D[Done]
</code></pre>

<h2 id="level-5-graph-engineering">Level 5: Graph Engineering</h2>

<p><strong>What it is:</strong> wiring many AI steps, or many agents, together like a flowchart. Different parts do different jobs and hand their results to each other.</p>

<p>An assembly line instead of a single worker. One step researches, the next drafts, another checks quality, and a final step publishes. Each node in the graph is a focused piece, and the connections define how work flows between them.</p>

<p>Examples in the wild:</p>

<ul>
  <li>A research agent that finds sources and passes them to a writing agent</li>
  <li>A review agent that checks the writing agent’s work against a style guide</li>
  <li><a href="https://n8n.io/">n8n</a>-style workflows with AI nodes in the middle</li>
  <li>Multi-agent systems built with tools like <a href="https://langchain-ai.github.io/langgraph/">LangGraph</a></li>
</ul>

<p><strong>What it gets you:</strong> complex, robust pipelines that split work across specialized pieces. If one step fails, you know exactly which one, and you can fix or replace it without touching the rest.</p>

<pre><code class="language-mermaid">flowchart LR
    A[Research agent] --&gt; B[Drafting agent]
    B --&gt; C[Review agent]
    C --&gt;|needs work| B
    C --&gt;|passes| D[Publish agent]
</code></pre>

<h2 id="beyond-the-five-levels">Beyond the Five Levels</h2>

<p>Two more approaches exist outside this ladder, and they work on a different axis.</p>

<p><strong>Fine-tuning</strong> changes the model itself instead of the interaction around it. You take a base model and train it further on your own data: your writing style, your company’s support tickets, your code conventions. The result is a model that knows your stuff before you even prompt it. This is heavier than any level above. It needs data, compute, and maintenance, and it is usually the last resort when context and harness engineering are not enough.</p>

<p><strong>Evaluation engineering</strong> is the quality control layer. Instead of asking “how do I make the AI better?”, it asks “how do I know whether the AI is good?”. You build a test set of known cases, run the model against it, and measure how often it gets things right. This matters more than people think. A system that is right 90% of the time and fails silently is worse than one that is right 70% of the time and flags its own uncertainty.</p>

<p>Both of these are their own crafts, and both keep evolving. They are worth knowing about, but they are not where a beginner should start.</p>

<h2 id="which-level-should-you-use">Which Level Should You Use?</h2>

<table>
  <thead>
    <tr>
      <th>Your problem</th>
      <th>Level</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Quick answer, one-off question</td>
      <td>1</td>
    </tr>
    <tr>
      <td>Answers must use your own data</td>
      <td>2</td>
    </tr>
    <tr>
      <td>The AI should take actions</td>
      <td>3</td>
    </tr>
    <tr>
      <td>Tasks where the first try often fails</td>
      <td>4</td>
    </tr>
    <tr>
      <td>A whole process with many moving parts</td>
      <td>5</td>
    </tr>
  </tbody>
</table>

<p>Start at level 1 and add levels as the problem demands. There’s no prize for using the most advanced level. The prize is solving the problem with the least complexity that works.</p>

<h2 id="a-personal-note">A Personal Note</h2>

<p>Everything on this blog now runs through levels 3 to 5. Puck has tools (harness), loops through steps until tasks are done (loop), and I run separate jobs for separate purposes: a weekly design job scan, a dashboard updater, a blog draft pipeline that opens pull requests for me to review (graph). I rarely type a raw prompt anymore. I engineer the context, the tools, and the loops around it.</p>

<h2 id="the-takeaway">The Takeaway</h2>

<p>Don’t be impressed by people who “prompt well”. The real craft is deciding what the AI sees, what it can touch, and how it learns from its own results. Prompts are the entrance. The levels above are where the actual work happens.</p>

<p>Start with prompts. Then work your way up one level at a time. Each level you add makes the AI more useful, and makes you think more like an engineer and less like a typist.</p>]]></content><author><name>Erik Zocher</name></author><category term="technology" /><category term="ai" /><category term="prompt-engineering" /><category term="agents" /><category term="workflows" /><category term="beginner" /><summary type="html"><![CDATA[A beginner's guide to the five levels of working with AI: prompt engineering, context engineering, harness engineering, loop engineering, and graph engineering.]]></summary></entry></feed>