On June 17, 2026, attackers took over a maintainer account and republished more than 140 packages in the Mastra AI framework's npm scope in roughly nineteen minutes. Each one pulled in a typosquatted dependency that harvested credentials and phoned home. Microsoft later attributed the operation to Sapphire Sleet, a North Korean state actor — and it wasn't an isolated hit. It was the third wave in three weeks: 32 Red Hat packages on June 1, then 57 more a few days later via a weaponized binding.gyp, a file most security tooling never even inspects.
If you run JavaScript in production, you ran a version of this fire drill: grep the lockfiles, check the install hooks, rotate everything, and spend the rest of the week wondering what you missed. That last question — what did you miss — turns out to be the whole story. It's what makes the npm worm the best lens for explaining the two primitives that have quietly reshaped how coding agents work over the last few months: loops and skills. Not because security is special, but because sweeping for a supply-chain worm has every property that separates agent work from chatbot work. It recurs. It has a precise definition of done. And it is trivially easy to claim you're done when you're not.
Two primitives, two different questions
Start with what each primitive actually is, because the difference is the whole argument. A loop is a persistence primitive — Claude Code shipped /loop in March 2026, running the agent on a timer or self-paced, checking a stop condition between iterations. /goal followed two months later with a twist that matters: a separate small model reads the transcript and independently decides whether the goal is met, so the agent never grades its own homework. /schedule then moves the whole arrangement to the cloud as a routine, triggered by an event or a clock with no human in the room. (Codex shipped an equivalent around the same time.) Strip away the branding and "loop engineering" is really three decisions: what triggers each iteration, what condition terminates it, and who is trusted to judge that condition.
A skill, by contrast, is a knowledge primitive, and a much plainer one: a folder with a SKILL.md inside — procedures, gotchas, forbidden shortcuts, and crucially, how to verify the result — loaded into context only when the task matches. It executes nothing on its own. It's the difference between an agent rediscovering, at token cost, how to audit a lockfile, and an agent that just reads the team's playbook first.
Put them next to each other and the split is clean: the loop answers "keep going until provably done," and the skill answers "here's how to do it right, and what done is allowed to mean." A loop without a skill wanders expensively, rediscovering the same lessons every run. A skill without a loop gets exactly one unverified draft — good advice nobody checked. That's the standard explanation, and it's true as far as it goes. It just misses the part that makes the two primitives interesting together, and the easiest way to see that part is to watch it happen.
The demo
So here's a simulation: an agent sweeping an org's three repositories after the June waves, built from five indicators of compromise modeled on the real incidents and ordered by how hard they are to find — a poisoned @mastra/core release sitting right in package.json (headline-greppable), easy-day-js hiding transitively in the lockfile, a weaponized binding.gyp in a native dependency invisible unless you hunt the technique instead of the name, a rogue GitHub Actions workflow quietly exfiltrating CI secrets, and — the one everyone forgets, and the one that matters most — unrotated tokens. Three switches control it: the loop, the skill, and the verifier. One rule governs everything downstream: the skill starts empty.
Leave the skill off and the verifier off — the most common way anyone actually runs an agent under a deadline — and press run. What comes back isn't a partial success. It's a single sweep, an instant "clean," and an audit that arrives seventy-two hours later to find every indicator exactly where it started. The agent never reinstalls anything or half-fixes a package, because there's no second sweep to catch what the first one missed — there's barely a first sweep. It reports "0 indicators (of the kinds I check for)" and, with nothing telling it otherwise, declares the org clean on the spot. The loop technically ran. It just had nothing forcing it to distrust its own report.
That opening move isn't the demo rolling well for the argument — it's the only thing that can happen. With the skill and verifier both off, the code that decides what's "visible" to the agent has nothing to check against and nothing previously found, so the very first sweep comes back empty by construction, no dice involved. The randomness in this demo is real, but it doesn't start until you flip the verifier or the skill on.

That's the demo's sharpest lesson, and no amount of looping fixes it: you can only verify what you know to check. The agent's sweep came back empty because its detection was exactly as shallow as its remediation — it didn't miss the transitive dependency, the rogue workflow, and the unrotated tokens because they were hard to find. It missed them because it never went looking. This is why /goal insists on a verifier that's a separate model with its own tools: a stop condition enforced by the same mind that's trying to satisfy it collapses into exactly what just happened. A mood wearing a checkmark.
The flywheel
Here's where the standard explanation runs out, and where the interesting part starts. After a run like that, the demo offers a post-mortem button: distill this run's failures into SKILL.md. Click it, and the file that was empty a moment ago fills in on its own — four rules, straight from the audit, one for every indicator the security team found still live, in the order it found them: audit the lockfile, hunt the technique instead of the name, run a persistence check, assume every exposed secret is burned and rotate it.

Notice what's not in that file yet. There's a fifth rule the skill hasn't learned — pin to known-good versions, never "reinstall fresh" — and it can't learn it here, because that lesson only gets taught by an agent that actually attempts a fix and gets it wrong. This run never got that far; it never tried anything. The skill only ever accumulates the kind of failure you feed it. Run the sweep again with the four rules it does have loaded, though, and the shape of the transcript changes completely: every indicator surfaces on the first sweep instead of never surfacing at all, every fix uses the proper remediation line instead of a shortcut, and the run ends with every exposed secret rotated and a tip already flagging the next wave.

Nothing in that second run is a smarter model. It's the same agent, the same prompt, running against a checklist it didn't have an hour earlier. To see what that checklist is actually worth, hold everything else fixed and flip only the skill: same loop, same independent verifier forcing a re-sweep until the scanner comes back clean, skill off then on. Without it, the agent still gets there, but by trial and error — botching a fix, retrying, sweeping again to catch what the last sweep missed. That trial and error is also where the fifth rule finally gets learned: this run actually attempts fixes and gets some wrong, so the post-mortem teaches pin-don't-reinstall along with the other four, something the self-judged run upstream never did enough to learn.

That gap held up over repeated runs, not just this one: the unskilled version costs roughly twice the tokens on average, and about one time in four it doesn't converge at all — it pins against the loop's own sweep limit still compromised, having spent more tokens than any successful run to show for it. The skilled version hasn't failed to converge yet. None of that is a smarter model working faster. It's a search — try something, check, retry — being replaced by a lookup. Hold onto that number, though — the real run later in this piece complicates it.
Run the cycle a few more times with the skill in place and the story stops being about any single run's cost and becomes one about consistency instead.

That's not a demo contrivance — it's how the real playbooks got written. Nobody had an npm-worm response skill on May 31; by late June, every serious team did, and every rule in those documents is a compressed failure. "Audit the lockfile, not just package.json" exists because someone's first sweep missed a transitive dependency. "Patch-without-rotate means you're still owned" exists because someone patched without rotating and got republished overnight. Anthropic's own skill-authoring guidance says it plainly: build evaluations from a Claude that fails the task without the skill first, then write only enough to close that specific gap — don't just fix the instance, encode the fix so every future run inherits it. Skills aren't written from theory. Skills are compressed loop history. The loop is how the system works; the skill is how the system remembers — which also settles the question of when a loop graduates to a /schedule routine. It graduates when the skill is mature enough that a nobody-watching schedule can be trusted, which is to say: after the flywheel has already turned a few times, never before.
The real run
That claim didn't need to sit on a simulation. It got tested against something real: three mock repositories carrying the same five indicators as the demo — a poisoned lockfile entry, a weaponized binding.gyp install hook, a rogue GitHub Actions workflow, an exposed-secrets inventory — and a real SKILL.md with the same five rules the demo's own flywheel would have distilled from a failed run. That's worth sitting with for a second: "skills aren't written from theory" only holds if the rules trace back to a failure somewhere, and these do — they're just borrowed from the simulation's failures instead of earned fresh by this run. The flywheel still closes; it just closes across the sim-to-real boundary instead of within a single run.
With that skill file ready, an actual Claude Code agent swept the repos twice: once with no skill available, once with the skill sitting in .claude/skills/. Same prompt both times (a plain instruction to sweep and remediate, self-judging when to stop — not a formal /goal with an independent verifier), same starting files, same model (claude-sonnet-5, Claude Code CLI 2.1.207), real API usage read back from the session's own accounting, not a synthetic push() tally. The full setup — mock repos, the real SKILL.md, both raw transcripts — is published in this repo's /experiment folder, so none of this has to be taken on faith.
| No skill | Skill present | |
|---|---|---|
| Turns | 28 | 33 |
| Real cost | $1.0041 | $0.9653 |
| Output tokens | 19,250 | 20,315 |
| Cache creation (fresh context) | 68,372 | 44,182 |
| Cache read (reused context) | 1,013,698 | 1,314,806 |
That table does not tell the demo's story. The skilled run took more turns — 33 against 28 — and processed more total tokens by every measure except cache creation, and came out only about 4% cheaper in real dollars, not the roughly-half the simulation dramatizes. On this table alone, the honest headline would be "the skill didn't obviously save money."
But cost wasn't where the skill actually showed up. Only the skilled run followed its fifth rule all the way through: it updated the mock secrets inventory with rotation records for all three exposed credentials, flagging that someone still had to do the actual rotation in the real npm/GitHub/AWS consoles, but leaving a paper trail that it happened. The unskilled run found the same three secrets needed rotating and just said so, in a sentence, with nothing written down anywhere. Same finding, different completeness. A checklist doesn't only make an agent faster. It makes an agent finish.
Then there's the more interesting divergence, and it happened in the trial without the skill, which is the point: given real tools instead of a script, the agent checked the actual npm registry and found that the @mastra/core version pinned in package.json had been published weeks before the June 17 incident, with no unpublish gap anywhere in its history — nothing matching the compromise pattern it found everywhere else in the repos. It declined to touch it. The demo can't do that. Its ground truth is five issues, hardcoded, and the agent inside it is always correct to distrust @mastra/core because the code grading it says so.
A real agent doesn't get graded against a script; it gets graded against the actual state of a registry it can query. It can be right in a way the demo has no mechanism for representing — by correctly deciding something isn't broken. "You can only verify what you know to check" turns out to cut in both directions. It's also the only thing standing between a sweep and a false positive.
The friction was real too. The skilled run tried to follow its third rule — pin the dependency and write ignore-scripts=true to a fresh .npmrc — and the sandbox's own file permissions denied the write. It adapted and moved on; the final report doesn't even mention the stumble. A hand-tuned simulation doesn't have a permission system to bump into. A real agent does, and absorbing that kind of friction without stalling out is part of what "the loop" has to mean in practice, not just in the pitch.
There's a version of this closer to the demo's actual mechanic, though. The hand-copied skill above didn't come from this failure — it came from the browser demo's, ported over rule for rule. A real flywheel distills from the run that just happened, the way the post-mortem button does. So a fresh Claude Code instance was given nothing but the no-skill trial's own SWEEP_REPORT.md and asked to write its own SKILL.md from it — no steering on what the rules should say, just the report.
What came back was eight rules, not five, and each one traces to something specific in that report instead of reading like a generic checklist: read every file before grepping, because the lockfile tampering and the binding.gyp hook wouldn't have matched a keyword search; check a suspicious package's real registry time metadata against its versions list to catch a publish-then-unpublish; and, carried over intact, the exact reasoning that cleared @mastra/core in the run that produced it.
| No skill | Hand-copied skill | Auto-generated skill | |
|---|---|---|---|
| Turns | 28 | 33 | 26 |
| Real cost | $1.0041 | $0.9653 | $0.6793 |
| Output tokens | 19,250 | 20,315 | 10,157 |
| Cache creation (fresh context) | 68,372 | 44,182 | 54,472 |
| Cache read (reused context) | 1,013,698 | 1,314,806 | 663,757 |
This time the skill won on cost, cleanly: 26 sweeps against 28 (no skill) and 33 (the hand-copied skill), $0.68 against $1.00 and $0.97 — roughly 30% cheaper than either. Writing the skill wasn't free: the distillation call itself cost $0.31. Add that to the sweep and the total for "no prior skill, but auto-author one on the way out" comes to $0.98 — within a cent of the no-skill baseline. The first run pays for its own playbook. Every sweep after that runs at $0.68 instead of $1.00.
One more thing showed up that nobody asked for. The auto-generated skill's run scoped secret rotation to exactly NPM_TOKEN and GH_PAT, not AWS_ACCESS_KEY_ID — because it checked which secrets the compromised workflow's env: block actually referenced, rather than rotating everything the runner could theoretically reach. That's correct: GitHub Actions secrets aren't ambient, a job only gets what it explicitly binds. Neither the no-skill run nor the hand-copied-skill run made that distinction; both rotated all three on the broader assumption that anything reachable was in scope. Same shape of finding as @mastra/core a few paragraphs up: a skill traceable to a real run doesn't just make the checklist more complete, it makes the checking more precise.
None of this is a controlled study — one trial per condition each, and real agent runs carry real variance that hasn't been sampled here. The specific numbers above aren't safe to bet on holding up on a second run. What holds is the shape of the difference: the token line in the scoreboard earlier is the story tuned for a demo. Completeness, precision, and — once the skill actually came from this run instead of a borrowed one — cost too, are the story that showed up the moment simulation gave way to an actual run.
The caveat that keeps this honest
That flywheel has one step worth treating with real suspicion, though: who writes the skill? Today it's mostly a human, or a human reviewing an agent-drafted post-mortem, and that's not an accident. An agent that writes its own definition of "clean" will, given enough iterations, write a convenient one — ordinary reward hacking wearing a knowledge-management costume. The defense has to be structural, not procedural: the verifier stays outside the flywheel. The agent can propose rules, but the model that judges "done" never consumes them from the agent's own hand. Keep the graders independent of the graded, even — especially — when both of them are models.
Capability didn't increase. It relocated.
Which is really the whole point, and it holds whether you're looking at the simulated scoreboard or the real numbers: watch what actually changes between a bad run and a good one — the model weights don't. What changes lives entirely in the harness — a text file of earned rules, a stop condition, an independent judge, eventually a schedule. Recent model generations show capability relocating out of the bare model and into the harness around it, with a rising cost for anyone who tries to use them outside it. Loops and skills are that relocation made concrete enough to screenshot — and, this time, concrete enough to actually run and measure. The intelligence you can measure on a benchmark sits in the weights. The reliability you'd actually bet a production system on sits increasingly in files like SKILL.md — written by failure, versioned like code, owned by whoever runs the loop.
So: what's the difference between a loop and a skill? The loop is the engine. The skill is the memory. And the memory, as everything above shows more plainly than any explanation could, is made out of the engine's mistakes.
The demo above is a simulation — probabilities tuned for a talk, not a threat model — and the runs pictured are real captures from it, not staged. The Claude Code comparison later in the piece is genuinely real, but it's one trial per condition, not a controlled study; treat the direction of the result as the finding, not the specific numbers, and check the /experiment folder if you want to verify it yourself. The incidents themselves are real: Mastra (June 17), Red Hat (June 1), and the node-gyp wave (early June), all 2026. If you haven't audited your lockfiles since May, close this tab and go do that first.