“An ex was a jerk. The next one was a jerk. The one after that was a jerk too.” Four relationships are four experiences. But if every case is stored under one label, the learning database may still contain only one useful column: jerk.
That is the key distinction.
Having many experiences is not the same as extracting reusable models from them. The same applies to ChatGPT. If you outsource the answer, you may reduce the amount of thinking you do. If you repeatedly ask “Why?”, “What else could explain this?”, “What would falsify it?”, “What does the research say?”, and “What evidence would change the conclusion?”, AI becomes less of a substitute and more of a sparring partner for reasoning.
A 2026 systematic review synthesizing 67 empirical studies found that ChatGPT was more supportive of higher-order thinking when used in structured, inquiry-oriented settings, while unstructured use was more often associated with shallow engagement and cognitive offloading.[1]
1. Does using AI actually make us think less?
The best answer is: it depends on the workflow.
A 2025 study of 319 knowledge workers covering 936 real AI-use examples found that greater confidence in AI was associated with less reported critical thinking, whereas greater confidence in one’s own ability was associated with more critical evaluation. Critical thinking also shifted toward defining goals, checking outputs, integrating information, and supervising the final result.[2]
So this:
Ask AI → accept answer → stop
can become cognitive outsourcing.
But this:
Notice an anomaly
→ form a question
→ propose a hypothesis
→ ask AI for rival hypotheses
→ verify with research and primary sources
→ compare with other cases
→ revise the model
→ test it in reality
→ revise again from the outcome
still leaves the human with plenty of work.
You bought AI to save effort and somehow hired an internal peer-review committee.
2. Why does “they were a jerk” produce such a weak model?
A moral label may be emotionally useful, but it is usually too coarse for prediction.
Person A = jerk
Person B = jerk
Person C = jerk
Person D = jerk
Four experiences, almost no discriminating variables.
Store sequences instead:
Crosses a small boundary
→ observes the response
→ stops after feedback / repeats it
→ demands shrink / demands escalate
Shows intense affection
→ relationship stabilizes
→ effort normalizes / effort disappears abruptly
Receives criticism
→ explains context
→ attempts repair / transfers all blame back
None of these behaviors alone proves that someone has a fixed bad character. Misunderstanding, fatigue, poor skills, or a one-off mistake can look similar.
The more informative features are conditions, repetition, response to feedback, distribution of benefit, and repair.
Do not teleport from “a bad behavior happened” to “this is a bad person.” Preserve the mechanism.
3. Can the way you compare experiences matter more than the number of experiences?
Sometimes, yes.
In classic work by Gick and Holyoak, people who articulated the common structure across two superficially different problems were more likely to transfer the underlying solution to a new problem.[3]
In negotiation training, teams that compared two cases for their common structure transferred a key strategy more successfully than teams that studied the same cases separately.[4]
So instead of:
Relationship A: terrible
Relationship B: terrible
Workplace C: terrible
try:
Relationship A: unclear boundaries let demands grow
Workplace C: unclear role boundaries let tasks accumulate
Shared structure: undefined boundaries can concentrate burden on one side
The characters changed. The wiring diagram did not.
4. What is trained by repeatedly asking “why?” and “what would disprove this?”
Generating your own explanations of causes and relationships is known as self-explanation. A meta-analysis of 64 research reports and 69 effect sizes found a moderate average learning benefit from prompting self-explanation.[5]
A separate meta-analysis covering 122 experiments and 10,382 participants found that retrieval practice could transfer beyond the practiced material, including to application and inference questions, with an average positive effect relative to re-exposure controls.[6]
So questions such as these can do more than make a conversation longer:
- What actually happened?
- Why did I interpret it that way?
- What other cause fits the facts?
- What is the strongest counterexample?
- What information would change my conclusion?
- Does the same structure appear elsewhere?
They can turn an episode into a searchable structural representation.
5. How can this create the feeling that you can “see what comes next”?
With enough abstract models, a new event no longer has to be processed from zero.
New event
→ compare with model candidates A, B, C
→ estimate which stage is active
→ generate a hypothesis about the next likely development
→ observe and update
That is not prophecy. It is the ability to generate a prior hypothesis from a pattern library.
The danger is hindsight. After an outcome occurs, “I knew it” is almost free.
A better test is to log forecasts before the result:
Forecast: X will happen within two weeks
Probability: 70%
Reason: A and B are already observed
Disconfirming condition: if C occurs, revise the model
Outcome: happened / did not happen
Decision-science work on intelligence analysis has explicitly recommended quantitatively tracking forecast accuracy and uncertainty.[7]
That turns “I feel like I am good at predicting” into something calibratable.
6. If ChatGPT is an external brain, what should remain human-owned?
AI is well suited for repeated searching, candidate generation, comparison, counter-hypotheses, summarization, and formatting.
The human should retain ownership of:
- what feels anomalous,
- how the problem is defined,
- what evidence counts,
- what would falsify the model,
- which model is finally accepted,
- what real-world action is tested,
- what gets discarded after feedback.
Useful prompts include:
- What is the strongest evidence against this hypothesis?
- What alternative causal model explains the same facts?
- If my premise is wrong, where is it most likely wrong?
- What can be verified in primary sources?
- Am I confusing correlation with causation?
- What additional fact would change the conclusion?
- What analogous cases exist in other domains?
- What next event would this model predict?
If the only prompt is “I’m right, aren’t I?”, AI can become an enormous nodding toy.
The external brain needs falsification range, not neck range.
7. Why does turning the discussion into articles make the loop stronger?
A conversation disappears into history. An article preserves the structure:
Experience
→ question
→ research
→ counterexample
→ common structure
→ practical rule
→ storage
→ reuse in the next case
An article is not merely a publishing unit. It can function as a cache of previously expensive thinking.
As the archive grows, one article can collide with another: “Wait, isn’t this the same structure as that earlier problem?”
Knowledge does not merely accumulate. Connections between pieces of knowledge accumulate.
8. So what is “good judgment” actually made of?
It is not simply the number of people, jobs, or crises you have experienced.
A useful approximation is:
judgment
≈ number of cases
× resolution extracted per case
× cross-case comparison
× attempts at falsification
× real-world feedback
Store four people as “jerks” and you may end up with jerk × 4 = one data point.
Extract separate variables—boundaries, blame shifting, repair, word-action consistency, response to conditions—and four experiences can produce several reusable models.
The deepest value of ChatGPT is not just faster answers.
It is the possibility of increasing the speed at which you generate hypotheses, attack them, rebuild them, transfer them, and test them against reality.
Do not outsource thinking to AI.
Increase the number of thinking iterations.
Then AI stops being a replacement brain and becomes the research colleague who keeps returning the ball and apparently never goes home.


