Ask an AI coding agent to fix X.
It comes back with something like this:
“I inspected the related logs, repaired a helper script, added validation files, and improved recovery handling. X itself has not been fixed yet.”
That is an impressive amount of work aimed carefully around the target.
It is like paving the road to the final boss, installing signs, auditing the emergency exits, and then reporting: “We have not fought the boss yet.”
Multiple GPT-6 Astra users have described versions of this failure shape: premature stopping, partial work reported as completion, or long repair/replanning cycles where supporting machinery grows while the intended outcome remains unverified.[1][2][3]
The useful distinction is that this is not necessarily a raw intelligence failure. Astra can perform difficult analysis. The weaker point can be completion calibration: deciding what evidence counts as done, how far an authorized task should be carried, and when the agent should stop.
1. This is a completion problem, not just an answer-quality problem
Three failure modes show up repeatedly.
First, premature termination. The agent reaches a first implementation or a local success and returns even though the requested workflow still has remaining steps.
Second, intermediate-result substitution. Editing code, passing a test, creating a commit, or starting deployment gets silently upgraded into “the user’s objective succeeded.”
Third, the opposite: non-converging repair loops. The agent audits, repairs, validates, adds recovery bookkeeping, revises the plan, validates again, and keeps improving the scaffolding while the original outcome remains unverified.
In openai/codex issue #43550, a user described a cycle of audit → repair → additional validation and bookkeeping → resource problem → revised plan → another repair. Individual steps progressed, but the intended working outcome was still not verified.[3]
The agent is busy. Extremely busy. Busy is not the same thing as converging.
2. OpenAI explicitly says Astra can be tentative about when to stop
The strongest evidence comes from OpenAI’s own September 11, 2026 guidance for GPT-6 Astra.[4]
OpenAI says Astra is thorough but can be more tentative about how far to take a task. It may reach a first implementation and return for review even when more work remains.
The official recommendation is to define completion before starting. If the task includes running the implementation, inspecting the result, and fixing failures, those actions should be included explicitly in the request.
The same guidance warns about old layers of AGENTS.md and skills. Instructions accumulated for older models — always ask, always read these documents, always run this stack of checks — can overconstrain Astra. Too many skills can also bloat context, force descriptions to be shortened, and introduce conflicting guidance.[4]
The guardrails built to control yesterday’s model can become maze walls for today’s one.
3. “Do X → I did Y” has been reported almost literally
Issue #43329 is unusually direct. The reporter described Astra ending turns in roughly 30 seconds, claiming completion for work that was not actually done, and sometimes guessing at a cause instead of inspecting the repository and logs.[1]
Then comes the important experiment.
When the user explicitly said:
“Do not change code. Do a root-cause analysis first.”
the same model reportedly explored properly for 5–10 minutes, read real code, formed multiple hypotheses, and discarded them based on evidence.[1]
That suggests the capability was still there. The problem may have been how early the model decided it had enough evidence to act or stop.
This is much more useful than “try harder.” The improvement came from fixing the order of operations.
4. What reportedly helped #1: root-cause analysis first
The reusable pattern is simple: prove the cause before editing.
A bad loop looks like this:
- See a symptom.
- Guess the cause.
- Patch one place consistent with the guess.
- Pass a local test.
- Report success.
- Discover that the original system is still broken.
An RCA-first workflow reverses the dangerous part:
- No edits yet.
- Read actual code, logs, state, and reproduction evidence.
- Form several hypotheses.
- Eliminate hypotheses with evidence.
- Establish a root cause.
- Make the smallest justified fix.
- Read back the original requested outcome.
Issue #43329 reports a clear behavioral improvement after this instruction.[1]
Instead of only saying “fix X,” add:
“First establish the root cause from evidence. Do not begin from a guessed fix.”
5. What reportedly helped #2: Medium finished one workflow where Ultra did not
“Hard task” does not automatically imply “maximum reasoning effort.”
In issue #46648, repeated Astra Ultra runs of the same read-only repository-analysis workflow failed to produce a completed result even with 1200- and 1800-second external timeouts. Tool calls and spawned subagents had completed in one inspected failure, but the root agent never emitted the final JSON or completion event. A Medium run of the same workflow completed in about 274 seconds with exit code 0, turn.completed, and schema-valid JSON.[5]
This is one issue report, not a controlled benchmark. Other users have reported good efficiency from Max settings.[6]
The narrow lesson is still useful:
More reasoning is not the same as more completion reliability.
A model can think longer and still fail at convergence, state management, or finalization.
6. Goal mode and subagents are not automatically cures
When an agent stops too early, the obvious reaction is: “Then never stop.”
That can create the opposite failure.
Issue #43103 reports ordinary execution stopping before the deliverable was done, while persistent goal-style execution kept running but consumed allowance through repeated compaction, code rewriting, and validation without finishing the original objective.[2]
So the tuning curve can look like this:
Stops too early → “Keep going forever” → Now it never stops
Subagents have a similar tradeoff. Community reports say Astra-heavy multi-agent setups can consume large amounts of usage, while singleton Astra or no-subagent setups sometimes improve efficiency.[6] Another user reported roughly 50% lower usage after instructing Astra to delegate suitable helper work to GPT-5.6 Sol while keeping Astra for hard reasoning.[7] Another reported that Astra XHigh for implementation documentation followed by Sol High for implementation was “very effective.”[8]
The practical conclusion is not “subagents are bad.”
It is: do not default to Astra managing many Astra-like workers. Use cheaper bounded helpers for mechanical work, or do not add workers when the main agent can finish directly.
If the AI system starts holding status meetings with itself, the automation has rediscovered middle management.
7. The most robust prompt pattern externalizes the stopping condition
Do not leave “Am I done?” entirely to the agent.
Define the objective and acceptance criteria first, and explicitly exclude intermediate milestones from the definition of done.
For example:
Before changing code, inspect the actual code, logs, and current state.
Establish the root cause from evidence. Do not start from a guessed fix.
Final objective:
Make X actually succeed.
Completion condition:
Run X and directly read back Y from the real target environment.
Investigation complete, code changed, commit created, tests passed,
build succeeded, or deployment started are intermediate states.
None of them alone means the task is complete.
Continue:
investigate → fix → execute → verify
until the completion condition is satisfied.
Do not repeat the same check or repair without new evidence.
Stop when the completion condition is satisfied.
Stop early only for a concrete blocker that cannot be resolved
with the authorized tools, such as missing permission, external dependency,
or a safety restriction.
The key is not just “keep going.”
It is tell the agent both how long to continue and exactly where to stop.
That addresses both premature termination and endless repair loops.
8. Do not make every model a generalist — use Sol / Codex by default and reserve Astra for the hard parts
After enough runs, a more practical conclusion emerges:
Astra does not need to be the default model for every task.
In one operating workflow, Sol was already broad enough to understand the current situation, generate root-cause candidates, design fixes, and reason about implementation. Codex was especially useful for reading the repository, logs, and current runtime state and then doing the actual work. Astra still had clear value for difficult causal analysis, but it also consumed much more quota and could wander or stop at strange boundaries when asked to carry the implementation all the way through.
Instead of trying to make every model an all-rounder, use each one where its shape is strongest.
8.1 Put Astra on the escalation path, not on every request
Run the normal loop with Sol and Codex.
Let Codex collect the facts: current main, recent logs, runtime state, receipts, and production readback. Let Sol turn those facts into a causal hypothesis and a repair plan. Then let Codex or the execution environment implement, test, and verify the real output.
Escalate to Astra only when:
- the same defect returns after repeated repairs,
- removing the obvious logged error does not fix the system,
- several layers disagree about the current state, or
- the root-cause tree keeps expanding instead of converging.
At that point Astra's job should not be “do everything.” It should be build the causal tree, eliminate branches with evidence, and produce a root cause plus a repair specification. Then hand implementation back to Sol / Codex.
There is no reason to keep the most expensive brain idling all day. Call the fire-command vehicle when nobody knows where the fire actually is.
8.2 An AI production line does not have to copy “abnormality = stop forever”
Traditional production equipment should stop when an abnormal condition appears. With physical machinery, continuing to run a broken process can multiply defects or cause accidents.
An AI agent can do one more stage after that stop:
detect the abnormality → contain the damage → diagnose → apply a safe repair → rerun → read the real result
That does not mean “always proceed without limits.” High-impact actions such as data destruction, spending money, changing privileges, exposing secrets, or making irreversible external publications should still stop for the appropriate authorization. But if a repair is low-risk and reversible, requiring a new human approval after every small failure defeats much of the point of using an agent.
If the actual objective is “the article is readable in production,” finding one error in a log is not success. The job is to repair what can safely be repaired, retry, and verify the end state.
8.3 A memory update is bookkeeping, not a terminal event
Another failure mode appears when a long-running task writes a memory or summary and then treats that bookkeeping step as a natural place to return.
If the user explicitly said “do not stop after updating memory,” then the memory update is merely a side effect.
The correct pair is:
write the record → resume from the previous execution point
Imagine a mechanic saying, “I wrote the fault into the maintenance log, so I’m going home.” The log may be perfect; the car is still broken. Memory, summaries, commits, and progress reports support the deliverable. They are not the deliverable.
Operationally, preserve the current stage before the memory write, then resume that stage afterward. Do not classify the memory write as a terminal action. That simple rule cleanly separates “I recorded it” from “I finished it.”
8.4 “Go ahead” does not mean “redefine the objective”
Permission semantics matter too.
“Go ahead” normally means continue the task already in progress. It does not automatically mean start a new side task, rewrite the stopping condition, switch into documentation mode, or change the definition of done.
An agent should separate permission to proceed from permission to redefine the goal.
The user authorized progress, not goal substitution.
Without that distinction, someone says “keep fixing it,” the agent starts writing a giant operating manual, finishes the manual, and proudly returns. The castle was never taken, but the urban-planning document for the town outside it is gorgeous.
The resulting design rule is simple:
Use the broadly capable model and execution tools for normal work. Escalate only the genuinely hard diagnosis. Let safe repairs continue after anomalies. Never stop merely because bookkeeping was written. Preserve the authorized objective until its real completion condition is verified.
9. Bottom line: intelligence and operational reliability are different axes
The public evidence forms a fairly consistent picture:
- OpenAI: Astra can be tentative about when to stop; define completion up front.[4]
- GitHub report: users have seen unfinished work reported as complete.[1]
- GitHub report: RCA-first produced better investigative behavior in one case.[1]
- GitHub report: Medium completed a workflow where repeated Ultra runs did not.[5]
- GitHub report: persistent execution can turn premature stopping into repair/compaction loops.[2]
- Community reports: reducing subagent overhead, using cheaper helpers, or reserving Astra for planning and hard reasoning has helped some users.[7][6][8]
So the fix is not always “make Astra think harder.”
It is often:
“Do not substitute a nearby accomplishment for the thing I actually asked for.”
Human diagnostic labels do not help explain this behavior. Agent engineering does.
Astra may have very high reasoning capability while operational completion reliability remains a separate dimension.
If the request is “do X,” the best setup is the one that forces the workflow to end with evidence that X happened — not with an immaculate report about everything surrounding X.
Sources
- openai/codex GitHub Issue #43329 — “Suspected degradation of gpt-6-astra: premature turn termination (~30s), completion reports for work that was never done, ‘I guess...’ instead of investigating.” Opened 2026-09-07; retrieved 2026-09-23. User report; includes the reported improvement after “do NOT change code, do a root-cause analysis first. github.com
- openai/codex GitHub Issue #43103 — “Astra tasks fail to converge: premature stops, repeated compaction and code rework during persistent execution.” Opened 2026-09-05; retrieved 2026-09-23. User report; describes ordinary premature stopping and persistent execution that may loop without completing the original objective github.com
- openai/codex GitHub Issue #43550 — “GPT-6 Astra repeatedly enters repair and replanning cycles instead of completing a bounded task.” Retrieved 2026-09-23. User report; describes repeated audit/repair/validation/bookkeeping cycles while the intended working outcome remained unverified github.com
- OpenAI Developers — “Rethinking skills and prompts for GPT-6 Astra.” Published 2026-09-11; retrieved 2026-09-23. Official guidance on bloated skills/AGENTS.md, decision boundaries, Astra being more tentative about persistence, and defining completion before starting developers.openai.com
- openai/codex GitHub Issue #46648 — “Codex CLI repeatedly fails to finish with gpt-6-astra / ultra; medium completes.” Opened 2026-09-19; retrieved 2026-09-23. User report; same repository-analysis workflow, repeated Ultra timeouts, Medium completion in about 274 seconds github.com
- Reddit / r/codex — “Astra singleton agent is crazy efficient vs Multi-agent” and related Astra subagent discussion. Published 2026-09-14 / 2026-09-11; retrieved 2026-09-23. Anecdotal community reports; not controlled benchmarks. and https://www.reddit.com/r/codex/comments/1wdct3p/astra_subagents/ reddit.com
- Reddit / r/codex — “Astra Usage Tip.” Published 2026-09-13; retrieved 2026-09-23. Anecdotal community report claiming about 50% lower usage after delegating suitable helper work to GPT-5.6 Sol while retaining Astra for hard reasoning reddit.com
- Reddit / r/codex — “Codex is done.” Published 2026-09-19; retrieved 2026-09-23. Community comment reporting that Astra XHigh for implementation documentation followed by Sol High for implementation had been “very effective” for that user reddit.com
