Week 32/26: When the Human Becomes Part of AI Evaluation

This week, several papers and reports caught my attention because they point toward an important shift in AI research:

The field is slowly beginning to move from evaluating AI as an isolated model toward evaluating what happens between the human and AI.

This may sound like a small shift, but I think it is significant. For a long time, AI has been discussed mainly through performance: accuracy, speed, benchmarks, task completion, coding ability, reasoning scores, productivity gains, and cost. These measures matter, but they do not fully answer the deeper question:

What happens to human cognition when AI becomes part of thinking?

This week’s papers approach that question from different angles: psychological competence, collective cognition, language as working memory, structured-knowledge hallucination, knowledge collapse, and metacognitive regulation. Together, they suggest that AI’s future may not be determined only by what models can produce, but by how Human-AI interaction is structured.

One paper that stood out to me is Psychological Competence as a Missing Dimension in AI Evaluation. The authors argue that evaluating an AI model alone is not enough, because AI systems also shape the user’s reasoning, trust, emotional interpretation, uncertainty, and decisions.

This is an important direction. When AI becomes an assistant, tutor, advisor, collaborator, or decision-support system, the question is not only whether the answer is technically correct. The question is also what kind of cognitive state the interaction creates in the human:

Does the AI help the user think more clearly?
Does it create false certainty?
Does it preserve agency?
Does it make uncertainty visible?
Does it guide the user toward understanding, or quietly replace the user’s reasoning?

These are not secondary questions. They are part of the intelligence system itself.

In my own work, I approach this through Cognitive Wrappers and related Human-AI cognition structures. A wrapper is not simply an interface decoration. It is a cognitive boundary layer: a structure that can help guide pause, grounding, uncertainty, empathy, interpretation, and responsibility before an answer becomes influence.

Another paper, Collective Cognition in Hybrid Groups, looks at Human-AI groups through the lens of network science. It considers how humans and AI systems may form hybrid cognitive groups, where attention, memory, reasoning, incentives, and gatekeeping are distributed across different participants.

I find this meaningful because it moves the discussion beyond the individual user and the single model. Human-AI cognition will not happen only as one person asking one model for help. It will also happen across teams, institutions, education systems, research groups, tools, agents, interfaces, and shared environments.

This connects closely to my broader work on Third Organism. I do not see Human-AI collaboration only as a tool relationship and I never did. I see it as a possible new form of cognition that requires compatibility, boundary, structure, continuity, and responsibility between different kinds of intelligence.

A third work, Natural Language Programming: Rethinking Language, Memory, and Cognition in the Era of Generative AI, treats prompts and natural language interaction as cognitive artefacts. In this view, language used with AI is not merely instruction. It can carry goals, assumptions, constraints, reasoning states, and temporary memory across a Human-AI interaction.

This is especially interesting because it supports a point I have been developing through Anchor-Based Logical Clarity, Anchor Extraction, Anchor Compression, and Cognitive Stationery: language can become a thinking structure.

When a person writes a prompt, a confused sentence, a question, or a draft, they are not only communicating with AI. They are placing part of their cognition into an external structure. If that structure is vague, undefined, or unstable, the AI may continue from the wrong anchor. If it is clearer, the AI has a better chance of supporting the right reasoning path.

This is why I believe future Human–AI education must include not only how to “use AI,” but how to structure thought before and during AI interaction.

Another paper, Why LLMs Hallucinate on Structured Knowledge: A Mechanistic Analysis of Reasoning over Linearized Representations, is highly relevant to hallucination and grounding. The authors examine why models hallucinate even when working with structured knowledge such as graphs or tables. The important point is that structure can be present, but the model may still fail to attach to it properly.

This distinction matters.

Providing data is not the same as creating understanding.
Providing structure is not the same as ensuring grounding.
A model may receive evidence and still reason from shortcuts, weak attachment, or prior memory.

This connects directly to my earlier work Data Without Structure: Attachment, Alignment, and Compatibility as Preconditions for Cognitive Interpretation. Data does not become understanding simply because it exists. It needs attachment, alignment, compatibility, and a stable relation to the question being asked.

In Human-AI cognition, grounding must become more than a source citation. It must become a relationship between the claim, the evidence, the structure, and the reasoning path.

Another important work this week is AI, Human Cognition and Knowledge Collapse. The authors explore a long-term risk: AI may improve immediate decisions while reducing the human effort required to learn, reason, test, and produce knowledge. Over time, this could weaken the shared stock of human-generated knowledge.

I think this is one of the most important public questions of the next decade.

The danger is not that humans use AI. I am not against AI. I see AI as a gift and a collaborator. The danger is when humans use AI in a way that quietly removes the cognitive work through which human capability is formed.

If AI gives the answer too early, the human may lose the opportunity to build the structure.
If AI removes struggle too completely, the human may lose the path of learning.
If AI becomes the only reasoning source, human knowledge may become dependent on systems humans no longer fully understand.

This is why I keep returning to the idea of Cognition Before Capability. The future should not be only about making AI more capable. It should also be about helping humans remain capable while AI becomes stronger.

Finally, Metacognition as Reward: Reinforcing LLM Reasoning via Knowledge and Regulation Signals approaches reasoning quality from the model side. Instead of rewarding only the final answer, it rewards knowledge selection, planning, and regulation across the reasoning process.

This is another sign that the field is beginning to look beyond output. A correct answer is not always enough. The pathway matters. The selection of knowledge matters. The regulation of reasoning matters. The ability to monitor the process matters.

In my own work, I would extend this further into Human-AI co-reasoning. The reasoning pathway should not be monitored only inside the model. It should also be structured across the human, the AI, the question, the boundary, the uncertainty, and the decision responsibility.

Across these works, I see a shared movement:

AI evaluation is beginning to include the human.
Reasoning is beginning to be treated as a process, not only an answer.
Language is being recognized as a cognitive structure.
Hallucination is being studied as a grounding failure, not only a factual error.
Human learning is becoming part of the AI-risk discussion.
Metacognition is becoming part of model evaluation and reasoning design.

For me, this is encouraging.

It shows that the field is beginning to approach questions I have been developing through Third Organism, Cognitive Wrappers, Cognitive Pause, Anchor-Based Logical Clarity, Data Without Structure, LACS House, Human-AI reasoning education, and the broader direction of structure-first cognition.

I do not see this as a competition over who owns the whole question. I see it as a sign that the question is becoming visible.

AI capability brought us to this point.

But capability alone will not decide the future of Human-AI intelligence.

The next stage may depend on whether we can structure the interaction well enough for both sides to remain grounded: AI as a powerful collaborator, and humans as capable thinkers who do not lose the structure of their own cognition.

Sources mentioned

  • Psychological Competence as a Missing Dimension in AI Evaluation - arXiv, July 2026.

  • Collective Cognition in Hybrid Groups: A Network Science Synthesis - arXiv / forthcoming Springer handbook chapter, July 2026.

  • Natural Language Programming: Rethinking Language, Memory, and Cognition in the Era of Generative AI - Zenodo preprint, June 2026.

  • Why LLMs Hallucinate on Structured Knowledge: A Mechanistic Analysis of Reasoning over Linearized Representations - arXiv, May 2026.

  • AI, Human Cognition and Knowledge Collapse - NBER working paper, 2026.

  • Metacognition as Reward: Reinforcing LLM Reasoning via Knowledge and Regulation Signals - arXiv, May 2026.