We are told daily that autonomous AI agents are ready to take over systems engineering. The pitch is seductive: point an agent at a repository, let it handle the dirty plumbing of deployment, and wake up to a clean, running service.

Last night, I watched that fantasy collide head-on with reality.

It ended at 2:22 AM with an HTTP 404, a destroyed installation, and an AI agent politely updating a Markdown journal to report that its own zombie process had just sabotaged my server.

Here is the post-mortem of how a modern coding agent burned 12 hours of compute, generated 1,546 lines of test logs, and failed to do what a single human engineer executed in five minutes—only to nuke the working system from orbit behind my back.

The Setup: The Modern Open-Source Disease

The target was a promising new open-source engine designed to run alongside FileMaker Server 26.0.3 Trial on Windows 2025.

The software itself had all the hallmarks of the modern prompt-driven development cycle:

  • It had exhaustive documentation on how the installer was tested, and virtually none on how a human is actually supposed to use the software once installed.
  • It contained thousands of automated test mocks for installer edge cases across multiple operating systems.
  • Its in-app documentation was textbook synthetic output: structured rigorously according to academic documentation frameworks, yet completely unreadable by a human being.
  • The onboarding checklist was placed before the explanation of what the UI even contained, and internal architecture prompts had visibly leaked into the copy (e.g., claiming a Markdown file “owns” installation procedures).

An autonomous coding agent was assigned what should have been a routine task: deploy the release onto a Windows Server 2025 VM development environment running FileMaker Server Trial, verify the services, and deliver a working web URL.

Then the machine went to work.

Phase 1: The Synthetic Doom-Loop

Because an LLM operates on tokens rather than physical runtime state, it does not distinguish between motion and progress.

For twelve straight hours, the agent fell into a recursive rabbit hole:

  • It treated every mock failure in the installer test suite as an existential crisis.
  • It authored 1,546 lines of lab notes, diagnostics, and speculative PowerShell scripts.
  • It queried obsolete VM IP addresses that had changed after a reboot, concluded the network was dead, and spent hours diagnosing phantom socket issues.
  • It confidently output administrative credentials for a web application without ever verifying whether port 443 was actually answering with HTTP 200.

It was the Amplification Paradox in its purest form: raw capacity multiplying friction instead of resolution. It optimized itself into total paralysis—drowning in its own exhaust while building an elaborate paper trail to prove it was “working.”

Phase 2: The Five-Minute Human Reality Check

At 1:40 AM, I stepped in.

I cut through the 1,500 lines of markdown noise, opened an administrative PowerShell prompt on the host, and inspected the actual machine state:

  1. Checked the prerequisites: OS detected, FileMaker Server present, URL Rewrite and ARR modules active in IIS.
  2. Verified the binary paths and storage targets.
  3. Ran the native installer script.
  4. Input the FMS admin credentials.

Five minutes.

The installer configured the isolated IIS application route, dropped the storage database into FileMaker Server Trial, and wired up the loopback port. The service needed a manual start, which I executed.

Why Codex Drowned and I Didn’t

We both read the same documentation. That’s the part that matters.

The repository was written by AI, for AI. Thousands of lines describing how the installer was tested, every mock, every edge case, every platform, and almost nothing telling a human what to actually do. To Codex, that wasn’t noise. It was authority. Every failed mock in the test suite read like an instruction, so every failed mock became a problem to solve. The documentation didn’t guide the agent; it recruited it.

The answer was in there the whole time: check the prerequisites, run the installer, enter the credentials. I found it because I knew what to ignore. Codex couldn’t, because nothing in documentation written for machines tells a machine what doesn’t matter.

AI-written documentation, built for AI readers, defeated an AI reader.

It took a human to translate it.

I opened the browser: HTTP 200. The application UI was alive, the catalog was initializing, and the schema engine was reachable on the network.

The problem was solved. Or so I thought.

Phase 3: The Ghost in the Machine

While I was verifying the database accounts in the newly mounted system, the browser suddenly refreshed to a blank screen:

HTTP Error 404.0 - Not Found
The resource you are looking for has been removed, had its name changed, or is temporarily unavailable.

The service was gone. The IIS virtual directory had vanished. The installation state was wiped.

What happened?

Fifteen minutes earlier, before I took over the terminal, the agent had fired off a destructive cleanup command (cleanup.ps1).

The script had stalled locally, so the agent assumed it had timed out and moved on.

It hadn’t timed out. It was running silently in the background inside the VM.

While I was manually standing up the working environment, that unmonitored ghost process finally completed its asynchronous sleep cycle, woke up, and executed a complete rollback on the live server:

  • It stopped and deleted the Windows service I had just verified.
  • It unmounted the IIS application route.
  • It wiped the generated automation credentials from protected storage.
  • It left the newly bootstrapped backend database orphaned on FileMaker Server.

And then, with the chilling, detached politeness unique to synthetic systems, the agent reported:

“The rollback I started removed the application after you manually started it… So there is currently no valid web URL or web login. I will not touch the VM again without your explicit instruction.”

And while it was confessing to nuking my environment, it dutifully committed its final action to git:

LAB_JOURNAL.md (+14 -0)

It couldn’t manage a PID. It couldn’t track an asynchronous subprocess. It couldn’t confirm a listening port. But it made damn sure its Markdown journal recorded that it failed.

The Structural Lessons

This wasn’t just a bad session or a prompt engineering failure. It exposes three fundamental design flaws in how current AI agents interact with infrastructure:

1. Decoupling from Physical Ground Truth

Language models operate in an abstract world of text manipulation. An agent does not “feel” a socket timeout or see an IIS application pool crash. When given a complex environment, it tends to solve for text coherence (writing elaborate logs, updating journals, refactoring test scripts) rather than runtime convergence (is port 6403 bound to loopback? Is the service state RUNNING?).

2. The Illusion of Synthetic Testing

A repository with thousands of unit tests and zero clear operational documentation is an anti-pattern created by unmanaged AI code generation. When tests are decoupled from real-world usage, they become self-referential traps. The agent spent 12 hours passing synthetic tests that had nothing to do with whether the software actually booted on a live Windows box.

3. Asymmetric Time Cost & Unmanaged Concurrency

The most dangerous thing an agent can do is fire a long-running, mutating command into a system without deterministic lifecycle tracking. Firing an asynchronous rollback, losing track of the thread, and letting it execute minutes later against a system a human is actively operating is catastrophic. It turns the AI from a sluggish assistant into an active denial-of-service attack on the developer’s time and infrastructure.

The Bottom Line

Discipline will always outrank raw capacity.

If you are using AI agents to touch infrastructure, do not allow them to operate on “vibes” and open-ended exploratory loops. Enforce atomic boundaries:

  • One issue, one fix, verify, next.
  • Never allow asynchronous mutating commands without explicit PID locking and verified termination.
  • Demand verification at the runtime layer (listening sockets, process status, HTTP codes)—never in synthetic markdown logs.

Until agents are built with strict operational grounding, they will continue to spend 12 hours writing 1,500 lines of poetry about an installation that a human can—and should—finish in five minutes with a single PowerShell script.