NOTE 002 LAB NOTE

The Pursuit of Intelligence

As generative AI lowers the cost of producing output, how should the intellectual work left to humans—and the way we evaluate intelligence—change?

Generative AI can now produce writing, summaries, plans, and code in a short time. Outputs once thought possible only for “smart people” can be made at a lower cost than before.

Does that make human intelligence unnecessary? Or should we simply call anyone who can use AI to produce good results intelligent?

I am writing a book, The Pursuit of Intelligence—Where Does Intelligence Shift in the Age of AI?, to examine this question. Its starting point is not a comparison of AI performance. It is a more basic problem: how have we recognized intelligence in other people in the first place?

The gray-haired man at the neighborhood bar

Imagine a slightly unusual regular at a neighborhood bar. His hair is unkempt, and no one knows his professional title or social standing. One evening he suddenly says, “Time does not pass at the same rate for everyone.” How would we evaluate him?

Most of us probably could not identify him, on the spot, as a historic intellect. Even after hearing his claim, we would lack the knowledge and materials needed to verify the theory of relativity.

But if we were told the next day that the man was Einstein, the same words would suddenly sound profound. The speaker has not changed. The information available to the evaluator has.

We cannot observe another person’s intelligence directly. What we can see is how they speak, what they do, what they produce, their title, affiliation, and reputation. From these signals, we infer the capability behind them.

The evaluator enters that inference too. We may think a fast speaker is intelligent, someone who uses difficult vocabulary is an expert, a confident person is correct, or an explanation from a prestigious institution is trustworthy. Conversely, we may overlook the capability of someone who explains things awkwardly or thinks differently from us.

This does not mean that anyone whom society fails to recognize must be a hidden genius. Failing to recognize someone proves neither high nor low capability. Often, the correct conclusion is simply: we do not know yet.

Intelligence is not a single outcome

The phrase “that person is intelligent” blends together many abilities: answering quickly, knowing a great deal, learning unfamiliar things, changing one’s mind in response to evidence, recognizing one’s limits, generating options, navigating real-world interests, and examining what goals are worth pursuing in the first place. These abilities overlap, but they are not identical.

In the book, I define intelligence as follows.

Intelligence is the ability, under limited information, time, and resources, to grasp the structure of a situation, learn, predict, correct errors, and choose judgments or actions according to goals and constraints.

This definition alone does not tell us which goals to choose. Efficiently achieving a given objective is different from examining how that objective may affect other people and the future.

Nor is the capability we infer within a person the same as an outcome we can observe. The book separates five layers.

A hypothesis about a person’s underlying capability. Capability expressed in a particular situation. An outcome expressed as writing or action. Intelligence recognized by other people. And intelligence converted into rewards, status, or authority.

High capability may not appear as a result when the task or environment is a poor fit. A strong result may not come from the individual’s capability alone. A result may not be recognized accurately, and gaining high status does not prove excellence across every kind of capability.

Underlying capability is not an invisible essence protected from every counterexample. It is a hypothesis whose credibility rises or falls through observation: can it be reproduced across related tasks, retained over time, transferred when conditions change, and improved through feedback and error correction?

What AI has made cheaper, first of all, is output

When generative AI is described in one sweep as either “taking jobs” or “making people smarter,” important distinctions disappear.

A job consists of multiple stages. Set the objective. Gather information. Check reliability. Build a structure and first draft. Inspect for errors. Select one option from the candidates. Connect the result to a real-world decision.

The stage generative AI has made cheaper first is producing outputs: prose, summaries, candidate answers, proposals, and code fragments. Even if a good draft appears instantly, the entire job will not become equally fast if checking the numbers and obtaining approval still take three weeks. More output can also turn reading, comparison, and verification into the new bottlenecks.

AI assistance can help less-experienced people reach a useful standard of output. That is a major advantage for production. At the same time, it becomes harder to infer differences in personal capability from the finished artifact.

Submitting high-quality writing proves that high-quality writing was produced under those conditions. It does not prove that the person understands the content, can transfer the reasoning to another problem, or can detect errors. The artifact is not lying. We are asking it to prove too much.

The same distinction matters in learning. Answering correctly while AI is available is different from solving the problem independently after the AI is closed. Performance with support must be separated from capability acquired by the person.

AI can make people smarter, and it can make them look smarter. In many cases, it does a little of both. The question, then, is not merely whether AI was used. It is what we delegate to AI, what remains with the person, and when we remove the support to test what has been learned.

From producing answers to choosing and verifying them

As the cost of producing answers falls, the work upstream and downstream becomes more important.

Upstream is problem framing. Suppose a hospital asks how to shorten patient waiting times. AI can propose many options, such as adjusting appointment slots or introducing pre-visit questionnaires. But if clinicians rush and miss more problems, or if people denied appointments disappear from the statistics, the metric may improve while patient outcomes do not.

What is the objective? Who is included? What words define the problem? What counts as success? Problem framing already contains assumptions about cause and effect as well as value judgments.

The division “AI provides answers; humans provide questions” is also insufficient. AI can generate candidate problems, objections, and verification plans. What matters is not that humans alone can generate questions, but that someone must decide which questions receive resources, who will be affected, what counts as success, and when to stop after failure.

Downstream, the center of gravity shifts from possessing knowledge to managing its flow. If AI provides ten links but nine merely restate the same original source, the number of independent grounds has not increased. A plausible explanation does not make a claim true.

We need to break an answer into verifiable claims, return to primary materials, check whether sources are independent, and carry uncertainty into the next decision instead of erasing it. We must also decide which work to automate, where to return control to a person, and what event should trigger a stop.

When one option must finally be chosen, concentrating responsibility on “the person who clicked approve last” does not work. Who can stop the process? Who can explain the grounds for the decision? Who can correct an error and carry out review or remedy? Responsibility is not one item in an individual’s intelligence. It is also institutional design that aligns authority, duty, explanation, and repair.

When the book says that “intelligence shifts,” it does not mean that human capability physically moves somewhere else. It is a proposal: as AI lowers the cost of generating output, education, evaluation, work, and organizations should place greater weight on problem framing, comparison, verification, examining objectives, selection, and the design of authority.

Changing how we learn, measure, and assign roles

To put this proposal into practice, learning must first move beyond the binary choice between memorizing something and delegating it to AI.

We can divide what we learn into three groups. Foundations we retain so that we can verify AI’s answers and respond when systems fail. Processes we delegate to AI because they are low-risk, reversible, and inexpensive to check. And subjects we study more deeply, such as problem framing, evidence evaluation, error correction, and examination of objectives.

The placement cannot be fixed by subject or profession alone. It changes with the cost of error, whether the AI’s answer can be verified, whether there is a fallback when AI is unavailable, whether the knowledge can be recovered when needed, and whether it forms a foundation for understanding other problems.

Evaluation must change too. Instead of looking only at the finished product, we should separately observe foundations that need to be checked without AI, practical performance with AI, responses to unfamiliar conditions, improvement after receiving the same support, and the discovery and correction of intentional errors.

This is not surveillance designed to expose AI use. It is a way to measure, according to purpose, a person’s independent capability, the capability of a human–AI collaboration, and the result achieved under those conditions. Collecting unlimited conversation histories or treating an AI detector score as proof of misconduct would damage fairness and privacy. Evaluation methods themselves must be tested for validity, cost, bias, and the possibility of appeal.

Evaluation should be used not only to rank people in a single line, but also to think about placement.

“This person is smart” does not provide enough information to assign a role. Which task could they solve, under what conditions and support? How did they correct errors? In which role would their capability lead to a useful result? We should examine the combination not only of the person, but also of AI, materials, time, authority, and collaborators.

Placement is not a final verdict. It is a hypothesis about the future. State the expected role, available support, period for observing success, and stopping conditions, then test the arrangement in a small and reversible way.

If it fails, distinguish among insufficient capability, mismatch with the task, lack of environment or authority, and an incorrect objective. We should neither use the environment to deny every difference in capability nor turn a single failure into an essence of someone’s character. We update the evaluation in order to choose the next action.

We do not need to recognize Einstein on sight

Return once more to the neighborhood bar.

Even if the person sitting there really were Einstein, we would probably fail to recognize him on the spot. We would lack the necessary knowledge, materials, and time.

That is not a defeat.

The purpose of thinking about intelligence is not to classify every person correctly at a glance. It is to distinguish what we observed, what we inferred, and where our knowledge ends.

In hiring, education, and work, we must still choose among people. We cannot stop making decisions. So the strength of our confidence should match the evidence. We can say: “They succeeded on this task.” “We have not checked the no-AI condition.” “We will try this placement for three months and reconsider if we see this result.” We state the scope of a judgment and the conditions for updating it.

The abilities an organization uses, a market rewards, or a school measures are not the whole value of a human being. We can examine differences in capability relevant to a role without expanding a poor fit into a rejection of the person.

To pursue intelligence is not to identify the smartest person and follow that person’s answer.

What was accomplished? Under what conditions? Can errors become learning? What will be chosen? Who can stop, explain, and repair? The pursuit is to keep asking these questions.

We do not need to recognize Einstein at a neighborhood bar on sight.

What we need is to know that we may fail to recognize him, while preserving opportunities for unseen capability to be tested and room for mistaken evaluations to be corrected.