Why Test Scores Alone Cannot Identify Who Will Perform Well at Work

Employers commonly use aptitude tests to assess verbal reasoning, numerical reasoning, logic, and personality.

Advertisement
Advertisement

Introduction

Employers commonly use aptitude tests to assess verbal reasoning, numerical reasoning, logic, and personality.

These tests are convenient. They allow organizations to compare many applicants quickly, and research does show that general cognitive ability is related to learning speed, training performance, and some measures of job performance.

The problem begins when that relationship is treated as an identity.

Scoring well on a cognitive test is not the same as performing well in a real workplace.

Real work does not arrive as a clean question with four answer choices. Information is incomplete. Stakeholders disagree. Several solutions may be defensible. Someone must decide what matters, what can wait, and who owns the result.

Generative AI has widened the gap. AI can now handle many tasks that resemble traditional aptitude-test questions, including summarization, text ordering, arithmetic, and condition matching.

That means employers need to assess more than unaided speed under artificial constraints.

They need to examine whether a person can define the real problem, make provisional judgments with incomplete information, set priorities, coordinate stakeholders, clarify accountability, finish work consistently, acquire local operational knowledge, verify AI outputs, and protect other people’s decision rights.

Job performance is a portfolio of capabilities, not one score.

Part 1: What Cognitive Tests Measure—and What They Miss

1. General cognitive ability is not useless

General mental ability, or GMA, includes broad capacities such as reading comprehension, reasoning, numerical processing, and learning.

A substantial research literature shows that GMA predicts job-specific performance and training outcomes to a meaningful degree. It can therefore serve as one source of information, particularly in complex roles that require rapid learning.

The wrong conclusion is not that aptitude testing has no value. A minimum level of comprehension and quantitative reasoning may genuinely matter for many jobs.

2. GMA is not identical to job performance

Actual work requires capabilities that conventional tests only partly capture.

Problem definition

Tests provide the problem. Work often requires people to discover it.

Employees must distinguish causes from symptoms, decide whose outcome matters, and determine whether an issue deserves action at all.

Prioritization

A person may be capable of solving every problem presented to them and still fail because they work on the wrong problem first.

Judgment under uncertainty

In practice, waiting for complete information is often impossible.

Effective work requires a cycle of separating facts from unknowns, forming a provisional hypothesis, taking reversible action, collecting new evidence, and revising the decision.

Stakeholder coordination

Knowing the correct answer does not move an organization.

There is also a difference between upward coordination—persuading senior leaders or peer departments—and downward coordination—giving employees the information, authority, and role clarity they need.

Follow-through

Starting an initiative, assigning it, and discussing it are not the same as completing it.

Domain knowledge

General intelligence does not substitute for knowledge of equipment, contracts, regulations, history, systems, and operational constraints.

AI verification

The relevant capability is increasingly not “Can you solve everything alone?” but whether the inputs are correct, whether the tool misread a condition, whether the answer follows from the assumptions, what must be verified before action, and whether the person can take responsibility for the final decision.

3. The “GMA-first” assumption is being revised

For decades, GMA tests were often described as the strongest single predictor of job performance.

More recent work has questioned some of the statistical corrections used in earlier validity estimates. A 2024 updated meta-analytic matrix reported that, when structured interviews, biodata, conscientiousness, integrity, and situational judgment methods are combined, removing GMA tests often produces little or no loss in overall predictive validity while reducing adverse impact.

This does not prove that cognitive ability is irrelevant. It suggests that GMA should be treated as one component rather than the center of every selection system.

4. AI changes the meaning of unsupervised testing

When an applicant can use AI during an unsupervised test, the result reflects a mixture of the applicant’s own ability, whether AI was used, prompt quality, image or text quality, verification skill, and interface speed.

Employers may respond by increasing surveillance, but a more future-oriented approach is to allow AI and assess the full workflow: how the problem was framed, what was delegated, which output was questioned, how it was checked, and how the final decision was justified.

Part 2: A Strong Individual Contributor Is Not Automatically a Strong Manager

1. The high-inference, high-coordination manager

Some managers are excellent at gathering incomplete information, building hypotheses, learning operational details, and persuading senior stakeholders.

They often appear highly capable because they can personally move difficult work forward.

These are real strengths. They do not automatically create a scalable team.

2. Management is not simply solving more problems personally

An individual contributor creates value by solving problems.

A manager must also create conditions in which other people can solve problems.

That requires defining the objective, stating decision criteria, assigning authority, setting priorities, removing obstacles, monitoring progress, confirming completion, and increasing the employee’s future independence.

A manager who investigates, decides, and personally rescues every issue may look efficient while keeping the organization dependent.

3. Upward coordination and downward enablement are different

A manager may be exceptional at influencing executives and peer functions while remaining poor at clarifying responsibilities, communicating priorities, allowing employees to challenge, matching authority with accountability, and creating reproducible processes.

This creates a split reputation. Senior leaders see a reliable operator. Direct reports experience ambiguity, late changes, and constant intervention.

4. Development can become self-serving

Employee development is healthy when it increases the employee’s independent judgment.

It becomes self-serving when the real goals are making employees reproduce the manager’s style, reducing the manager’s explanation burden, training employees to anticipate unspoken expectations, treating disagreement as lack of understanding, and retaining final control.

The activity looks like coaching. The outcome is dependence.

Part 3: When Benevolence Becomes Control

1. Good intentions do not guarantee healthy boundaries

Advice becomes healthy when the adviser can say: “This is my view. You may decide differently.”

Control begins when the adviser assumes: “I understand what is best for you better than you do.”

Good intentions can make this more dangerous because they reduce self-doubt. Intrusion can be renamed as care. Pressure can be renamed as commitment. Resistance can be interpreted as immaturity.

2. Thinking about someone is not the same as deciding for them

Employers and managers have legitimate reasons to discuss job requirements, working conditions, development, and support needs.

They do not need to construct an employee’s private future for them.

Japan’s Ministry of Health, Labour and Welfare states that fair selection should be based on job-related aptitude and ability, that organizations should avoid collecting unrelated private information, and that interview questions and evaluation standards should be defined in advance.

The design principle is simple: Do not collect information that the job does not require merely because it might help someone build a more detailed personal story.

3. Strong inference can produce sophisticated false stories

A coherent story is not necessarily a true story.

A skilled reasoner can connect limited information into a detailed narrative about another person’s motives, future, and best interests. Without verification, that narrative remains a hypothesis.

Healthy reasoning requires explicit labels: fact, interpretation, hypothesis, employee decision, organizational decision, reversible action, and irreversible action.

4. The ability to end the conversation is a boundary skill

A mature adviser can offer a view and return the decision.

A controlling adviser keeps talking until the other person adopts the preferred conclusion.

The problem is not a lack of explanation by the recipient. It is the adviser’s inability to accept a different decision.

Part 4: Returning Execution Responsibility to the Person Who Demands the Ideal

1. Correct principles do not allocate resources

“Improve quality,” “prevent every error,” and “communicate with everyone” are reasonable principles.

They do not specify available staffing, deadlines, trade-offs, decision authority, or the accountable owner.

2. The unhealthy split

In an unhealthy system, one person supplies ideals and criticism while another person absorbs all implementation costs.

The first person remains correct. The second person carries prioritization, coordination, execution, and failure risk.

3. A boundary-based response

A practical response is:

I understand the requested standard. Under the current staffing and deadline, it conflicts with existing priorities. If it is mandatory, please identify the work to defer, the owner, the deadline, and the decision authority.

Where the requirement reflects a personal preference:

If that method and quality level are mandatory, the person holding that requirement should also hold the decision and execution responsibility.

This is not refusal. It reconnects preference, authority, and accountability.

4. “The world does not run on logic alone” applies to everyone

Managers sometimes use realism to dismiss employee objections while treating their own principles as exempt from constraints.

That is not realism. It is asymmetric responsibility.

Part 5: Why Training Alone Does Not Change Behavior

1. Knowledge is not self-application

People can understand a harassment or human-rights training module while excluding themselves from the category it describes.

They may define harassment as hostility while defining their own conduct as guidance, care, or development.

2. Everyday rewards overpower one-time training

If intrusive behavior is rewarded as commitment, expertise, or leadership—and produces no consequences—the individual has little reason to change.

The U.S. Equal Employment Opportunity Commission has emphasized that training cannot stand alone. Prevention requires leadership, accountability, reporting systems, investigation, proportional corrective action, and evaluation of managers.

Training communicates knowledge. Organizational consequences communicate reality.

3. A small number of powerful people can shape the culture

A behavior does not need to be common to become culturally dominant.

If a small number of influential people can intervene in hiring, evaluation, assignments, meetings, and priorities, their habits become system-level conditions.

Part 6: Walking Dialogue as a Thinking and Health Practice

1. Walking changes the mode of thought

Sitting at home can keep a person inside the same emotional and cognitive loop.

Walking while speaking thoughts aloud creates a sequence: externalize the mental load, impose verbal order, receive a response, reduce emotional heat through movement, separate evidence from interpretation, compress events into a reusable model, and choose one next action.

This combines movement, reflection, dialogue, and decision-making.

2. Walking appears especially helpful for divergent thinking

A 2026 systematic review and meta-analysis covering 23 studies and 1,036 participants found a large positive effect of walking on divergent thinking, while evidence for convergent thinking was uncertain and close to null.

That supports a practical division.

During the walk

  • generate explanations
  • identify patterns
  • develop alternatives
  • build an article outline
  • express emotion

After the walk

  • verify numbers
  • inspect evidence
  • finalize decisions
  • send messages
  • confirm commitments

3. Dialogue turns rumination into processing

Rumination repeats. Structured dialogue transforms.

A set of disconnected experiences can be compressed into one model—for example, “crossing into another person’s decision rights and execution responsibility.”

Once the model becomes clear, the response changes from persuasion to boundary-setting.

4. The health return is real even without immediate weight loss

Regular walking increases physical activity and supports cardiovascular health, mood, sleep, and functional fitness.

When reflection itself becomes the reason to walk, exercise is easier to sustain because one habit serves several purposes.

5. A five-part closing template

End a walking reflection by identifying facts, interpretation, hypothesis, controllable factors, and one next action.

Part 7: A Better Way to Assess Work Capability

A modern selection system should combine several methods:

  • basic ability testing for minimum comprehension and numeracy
  • structured interviews with predefined scoring criteria
  • behavioral examples from past work
  • job-relevant work samples
  • AI-enabled tasks that assess prompting, checking, and judgment
  • evaluation of how candidates define authority, responsibility, and priorities

Managers should also be evaluated on whether they make decision criteria visible, match authority with accountability, reduce unnecessary ambiguity, respect disagreement, return personal decisions to the employee, finish work through the team, increase employee independence, and create systems that function without them.

The goal is not to find a flawless person.

It is to determine whether strengths and weaknesses can be recognized, balanced, and corrected without transferring the cost of one person’s weaknesses to everyone below them.

Conclusion

Cognitive testing can provide useful information, but it cannot represent the whole of work capability.

Real performance depends on problem definition, judgment under uncertainty, prioritization, stakeholder coordination, follow-through, domain learning, AI verification, and respect for decision boundaries.

A person may be a powerful problem solver and still be a weak manager.

When inference is not separated from fact, advice from decision, and preference from responsibility, competence can become control.

The practical response is to restore boundaries: separate facts from hypotheses, separate advice from authority, connect preferences with execution responsibility, and return personal decisions to the person who owns them.

And when thought becomes overheated, walk.

Movement and dialogue can turn repeated mental noise into structure, action, and health.


参考文献・References

  1. 厚生労働省「公正な採用選考の基本」
    https://www.mhlw.go.jp/stf/seisakunitsuite/bunya/koyou_roudou/koyou/newpage_56780.html

  2. 厚生労働省「公正な採用選考 よくある質問」
    https://kouseisaiyou.mhlw.go.jp/question.html

  3. Berry, C. M., Lievens, F., Zhang, C., & Sackett, P. R. (2024). Insights from an updated personnel selection meta-analytic matrix: Revisiting general mental ability tests’ role in the validity-diversity trade-off. Journal of Applied Psychology, 109(10), 1611–1634.
    https://pubmed.ncbi.nlm.nih.gov/38695805/

  4. Berry et al. (2025). Correction to the 2024 article. The correction did not change the study’s conclusions.
    https://pubmed.ncbi.nlm.nih.gov/40658550/

  5. Hambrick, D. Z., Burgoyne, A. P., & Oswald, F. L. (2024). The validity of general cognitive ability predicting job-specific performance is stable across different levels of job experience. Journal of Applied Psychology, 109(3), 437–455.
    https://pubmed.ncbi.nlm.nih.gov/37843546/

  6. Woods, S. A. et al. (2024). A critical review of the use of cognitive ability testing for selection into graduate and higher professional occupations. Journal of Occupational and Organizational Psychology.
    https://bpspsychub.onlinelibrary.wiley.com/doi/full/10.1111/joop.12470

  7. U.S. Equal Employment Opportunity Commission. Select Task Force on the Study of Harassment in the Workplace.
    https://www.eeoc.gov/select-task-force-study-harassment-workplace

  8. Thabane, A. et al. (2026). The impact of walking on creative thinking: A systematic review and meta-analysis. PLOS One.
    https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0347878

  9. PubMed record for the walking and creative thinking meta-analysis.
    https://pubmed.ncbi.nlm.nih.gov/42127076/


関連記事・内部リンク候補

  • 適性検査で測れる能力と、実務で必要な能力の違い
  • 優秀なプレイヤーが管理職で失敗する理由
  • 善意の押し付けと心理的境界線
  • 責任と権限が一致しない職場で起きること
  • 「正論」を実行責任とセットで返す方法
  • ハラスメント研修だけでは職場が変わらない理由
  • 歩きながら考えると頭が整理される理由
  • AI時代の採用試験は何を測るべきか

  • ☑ 個人名を記載していない
  • ☑ 企業名を記載していない
  • ☑ 地域名を記載していない
  • ☑ 部署名・役職の組み合わせで特定できる記述を避けた
  • ☑ 具体的な採用面接の質問内容を記載していない
  • ☑ 個別事例を一般化し、複数の事象を再構成した
  • ☑ 日本語版とUS英語版を1ファイルに収録した
  • ☑ 見出し階層を統一した
  • ☑ 研究上の知見と構造分析を区別した
  • ☑ 参考文献を末尾にまとめた
  • ☑ published: false / status: private に設定した
Advertisement
Mendoi-chan

Written by

Mendoi-chan

She turns friction at work and in everyday life into clear structure and practical next steps.

About
Advertisement

Latest articles

  1. 1Do AI Agents Make Humans Unnecessary? How Environment Design and Trend Signals Can Build a Media System That “Kicks the Boss Out of the Factory”
  2. 2Should Long-Running AI Agents Keep Progress Logs? A Heartbeat Design That Prevents “Did It Stop?”
  3. 3Is ¥15,000 a month for AI expensive? It looks different when you are buying back your evenings and weekends
  4. 4The Third Eye Is for Gacha: Where Intuition Helps and Where Logic Must Take Over
  5. 5How to Stop Wasting ChatGPT Pro’s Weekly Message Limit: What Counts as One Use, Retries, and Accidental Sends

You may also like

Advertisement