Yielded Is Not Finished: Guarding Partial-Clone Git Scans Without Breaking Git
An agent tool returned control after launching a broad Git scan. The easy mistake was to read that as completion.
The process was still alive. In a partial clone, a broad history or object walk can fetch missing objects on demand, and later maintenance can amplify temporary allocation. The caller had yielded, but the resource owner had not finished.
Yielded is a scheduling state. Finished is a terminal state.
The right response was not “ban Git.” It was to identify the reproduced command families, give every long-running process an owner and deadline, and prove that the narrow guard stopped those canaries without blocking ordinary work.
The Failure Chain
broad historical scan in a partial clone
→ missing objects may be fetched on demand
→ maintenance may amplify temporary allocation
→ the caller yields while the process tree remains alive
→ storage pressure continues without an owner or reap path
Each arrow is plausible and operationally useful, but the chain still needs careful wording. A private incident can show the process tree and storage movement without proving the exact split among fetching, repacking, open-file retention, and filesystem recovery.
That limitation changes the claim. I can justify a guard for the reproduced incident class. I cannot claim a complete causal model for every partial-clone disk spike.
The State Model the Tool Needed
| State | Meaning | Owner obligation |
|---|---|---|
| Running | The process is active and the caller is still attached. | Track identity, deadline, output, and resource signals. |
| Yielded | Control returned before the process terminalized. | Persist an owner, poll path, cancellation path, and deadline. |
| Cancellation requested | The owner asked the process tree to stop. | Wait for and verify the actual terminal state. |
| Reaped | The process tree is gone and terminal output/state was read back. | Reconcile the claimed effect and remaining resource state. |
| Completed | The command reached a verified terminal outcome. | Only now may the caller make a completion claim. |
A timeout alone is not ownership. Neither is returning a session handle. The process must have a named owner that can poll, cancel, reap, and reconcile it after the original call yields.
A Narrow Classifier Beats a Global Git Ban
The sanitized reproduction distinguishes five shapes:
| Command shape | Guard result | Reason |
|---|---|---|
| Ordinary status inspection | Allow | No broad object or history walk. |
| Bounded single-reference history | Allow | The scope is explicit and narrow. |
| Broad content search across history | Block in protected partial clones | Matches a reproduced materialization family. |
| Broad object enumeration | Block in protected partial clones | Matches the second reproduced family. |
| Same broad scan with a no-lazy-fetch override | Allow to fail rather than materialize | The operator explicitly chooses a bounded failure mode. |
The classifier should consider both the command shape and repository context. A pattern that is risky in a protected partial clone may be acceptable in a disposable full clone. That is why a global alias, blanket shell restriction, or universal Git disablement would be disproportionate.
if repository_is_protected_partial_clone:
if command_matches_reproduced_broad_scan:
if explicit_no_lazy_fetch_override:
allow_bounded_failure()
else:
deny_with_diagnostic()
else:
allow_ordinary_git()
else:
leave_policy_unchanged()
A Denial Log Is Not Non-Execution Proof
One ingress can say “denied” while another adapter still executes the nested shell command. That makes policy logs necessary but insufficient.
For each tested ingress, my public-safe canary requires four pieces:
- the same classifier identity and semantics;
- an independently observed attempt;
- an explicit denial event;
- an absent execution marker after the denial.
A separate allowed positive control must create the marker successfully. Otherwise marker absence could mean the observer was broken rather than the risky command was stopped. In the canary, the fake executable writes the marker immediately on process entry, before any Git history or object work; an absent marker therefore proves that the tested command never entered that fake execution path. It does not prove that no wrapper or unrelated side effect ran.
| Ingress | Attempt seen | Denial seen | Marker absent | Positive control |
|---|---|---|---|---|
| Default agent-runtime shell | Required | Required | Required | Required |
| Nested coding-agent shell | Required | Required | Required | Required |
Fail-Open Is an Availability Tradeoff
The guard described here fails open if its adapter itself errors or times out. That avoids turning a classifier fault into a broad shell outage. It also means the guard is not a security boundary.
Fail-open can still be useful for accidental resource control when the bypass is observable. A sanitized fail-open control models adapter error as command allowed plus a visible diagnostic, and direct fixtures exercised malformed input and missing-classifier cases. End-to-end timeout behavior remains a specified availability tradeoff, not a published timeout canary. Silent classifier failure would make the operational claim false.
If the threat model includes malicious bypass, untrusted operators, or adversarial commands, this pattern is insufficient. Use a real sandbox, resource limits, repository isolation, and least-privilege execution boundaries.
When Not to Use This Pattern Blindly
- Do not infer that every broad Git command has the same materialization behavior.
- Do not use a private incident as proof of exact byte attribution.
- Do not treat a yield handle as a terminal receipt.
- Do not count denial text without independent marker evidence.
- Do not let a narrow accidental-resource guard masquerade as a security sandbox.
- Do not block ordinary Git merely because a small set of reproduced shapes was expensive.
Known uncovered accidental paths include Git auto-maintenance reached through otherwise allowed commands, other broad history/object walks, IDE-initiated scans, manual terminals, and unlisted harness ingresses. The classifier covers the two reproduced families; it is not a complete resource policy for Git.
The Checklist I Would Reuse
- Separate incident observation from causal attribution.
- Identify the exact reproduced command families and repository context.
- Keep ordinary and explicitly bounded Git shapes available.
- Give every yielded process an owner, deadline, poll, cancel, and reap path.
- Validate every ingress with attempted/denied/marker-absent evidence.
- Run an allowed positive marker control.
- Make fail-open adapter faults observable.
- State clearly that the guard is neither a sandbox nor a complete Git cost model.
Conclusion
The expensive lesson was not that Git is dangerous. It was that an asynchronous caller can lose ownership of a still-running command while the system quietly continues doing exactly what the command asked.
The proportionate repair is small: narrow classification for the reproduced incident class, explicit lifecycle ownership, and proof that each intended ingress stopped the canary without breaking normal Git.
That is a better operational boundary than either extreme: trusting a yield as completion, or disabling the tool because one command shape ran away.