Agent skill · Documentation

harness-evolve

Run `@metaharness/darwin evolve <repo>` to mutate a harness's seven policy surfaces (planner/contextBuilder/reviewer/retryPolicy/toolPolicy/memoryPolicy/scorePolicy), sandbox-score each variant, and promote only measured wins. The model is frozen; the harness evolves. Closes the loop ADR-150 opens (score+genome describe; evolve changes). Degrades gracefully when @metaharness/darwin is absent (ADR-150 + ADR-153 architectural constraints).

rUv71,307★ · +1,002/wk · 3 repos on radarProfile →
claude-codecodexcan modify filesMIT
Install
npx skills add ruvnet/ruflo --skill harness-evolve --agent claude-code

Same command for any agent — swap --agent for codex, cursor, copilot.

Facts
Files in the skill folder: 1
SKILL.md size: 6 KB
Bundled scripts: none
Allowed tools: Bash
Path: plugins/ruflo-metaharness/skills/harness-evolve/SKILL.md
Open the folder on GitHub →
Where it comes from
Source: ruvnet/ruflo
Stars: 67,015 · +629 this week
Language: TypeScript
Read our review of the source →

Weekly change comes from our own snapshots, not the repository page — it measures attention, not adoption.

From the SKILL.md

Surfaces the upstream `metaharness-darwin evolve` CLI as a ruflo skill. The **write** layer that pairs with ADR-150's read layer (score / genome / mcp-scan / threat-model / oia-audit). Use when you have a harness whose readiness scores are flat and you want to discover *which* surface mutation moves them — without retraining the foundation model. ## When to use - A `harness-score` result is below target and you don't know which policy surface is responsible. - You're seeding a harness for a new vertical and want to find a good starting configuration empirically rather than hand-tuning. - You're comparing your hand-tuned harness against an evolved baseline (treat darwin's champion as the strawman). ## When NOT to use - For continuous background optimization. Darwin Mode is human-initiated. Wire it into CI for one-shot exploration, not for autonomous self-modification. - For ruflo itself in CI. ADR-153 §5 explicitly rejects auto-evolving ruflo — the CI gate verifies graceful degradation, not convergence. ## Algorithm Implementation: [`scripts/evolve.mjs`](../../scripts/evolve.mjs). 1. Validate args (`--repo` exists, caps on `--generations` ≤ 50, `--children` ≤ 20, `--concurrency` ≤ 8

What's inside
Steps it walks through
  1. When to use
  2. When NOT to use
  3. Algorithm
  4. The seven mutation surfaces
  5. Output
  6. Failure diagnosis (--diagnose)
  7. Exit codes
  8. Graceful degradation (ADR-150 constraint 3 + ADR-153)
More from ruflo
All skills →
About this skill
What does the harness-evolve skill do?

Run `@metaharness/darwin evolve <repo>` to mutate a harness's seven policy surfaces (planner/contextBuilder/reviewer/retryPolicy/toolPolicy/memoryPolicy/scorePolicy), sandbox-score each variant, and promote only measured wins. The model is frozen; the harness evolves. Closes the loop ADR-150 opens (score+genome describe; evolve changes). Degrades gracefully when @metaharness/darwin is absent (ADR-150 + ADR-153 architectural constraints).

How do I install it?

Run `npx skills add ruvnet/ruflo --skill harness-evolve --agent claude-code` — it drops the skill into your project so the agent can pick it up. Swap the --agent value for codex, cursor or copilot if you use one of those.

Where does this skill come from?

From ruvnet/ruflo, a repository with 67,015 stars. We read it straight from the repository tree rather than a submitted listing, so what you see here is what is actually published.

Is a popular skill a good skill?

Not necessarily. Stars measure attention, not adoption — a repository can trend for a week and be abandoned. That is why we show the weekly change from our own snapshots next to the total, instead of a single flattering number.

Keep going