Analysis and opinion

LLMs and System 2: Where Does the Checking Happen?

An opinion on LLM reasoning and the System 1/System 2 analogy, using one queue decision to ask where assumptions meet evidence.

François Guéguen 8 min read

TL;DR: An AI can recommend an action without establishing what will happen if it is taken. The general problem is how a system checks consequences, exposes assumptions and revises its choice when evidence changes. More reasoning can help; the question is what it actually checks. The queue-design example below makes that problem concrete.


Imagine an engineering team designing report generation for a customer app. They ask an AI assistant, a software application using a large language model (LLM), how to handle work that might take too long for a web request. The assistant recommends a queue: put requests on a waiting list, let a background process called a worker generate each report, and notify the user when it is ready. The engineering team owns the design decision. Before accepting the recommendation, what would they want the AI assistant to have checked?

Perhaps the customer needs the result immediately. Perhaps the worker creates a report but fails before telling the queue the job is done, so the queue may retry it. Perhaps most reports are so quick that a queue adds more operating burden than it removes. The recommendation could be right, but its reasons depend on conditions the team has not established.

The System 1/System 2 analogy is most useful as a question about where an AI application investigates the consequences of a proposed choice. That checking may happen within the AI assistant or through checks the engineering team runs. The queue is ordinary software; the AI question is how the assistant's recommendation meets the requirements and failure behavior of the real app. More reasoning may help, but the team needs to see what was checked and what could change the choice.

What did the longer answer check?

Suppose the LLM spends longer reasoning about the queue question, and the AI assistant returns a detailed analysis of waiting time, retries and cost. That may be a better answer. A reasoning sequence can compare alternatives, notice a mistake and revise a plan.

OpenAI's o1 research account reports improved performance with more inference-time computation. It also describes sampling and selecting candidate answers in particular evaluation settings. Those are concrete procedures, but that public account does not establish a maintained tree of future states as the default mechanism behind every answer.

The GPT-4 technical report describes that model's next-token pretraining objective. It does not settle what computations a model can support. Training objectives, inference procedures and the AI assistant's actions are different levels of description.

For the queue decision, the useful distinction is between describing what might happen and checking a consequence through an identifiable procedure. What evidence did the assistant use: measurements from an actual report run, results from a worker-failure test, or a comparison against the stated deadline? Those results would require suitable tools connected to the assistant or checks run by the engineering team. Or did the LLM explain what usually makes queues useful without that evidence?

Tree of Thoughts is one explicit search method, branching over model-generated intermediate reasoning. Its states and evaluations still need scrutiny. A single reasoning sequence can also deliberate meaningfully; neither format alone tells the team whether the queue will meet this app's needs.

System 1 is a role, not a verdict on LLMs

Human dual-process theories distinguish relatively automatic judgments from more effortful hypothetical reasoning. Evans and Stanovich's account concerns human cognition. System 1/System 2 is a limited analogy here, not a classification of LLMs.

In that analogy, the LLM supplies a learned judgment through the assistant: a queue looks promising. Further investigation could use the same model, another model, a test or a search procedure. Two model calls do not automatically create two cognitive systems; one model does not automatically lack deliberation.

The roles become useful when they make the work visible. A component can choose an action from a short list without investigating what follows from it. Another component can check a narrow consequence without answering the whole design question.

Two current examples of typed decision interfaces

TypeSafe AI's Jev launch announces early access for its "System One Model"; its docs describe typed questions over supplied state. OpenAI's September 29, 2026 recap announces Decisions API in limited preview: finite answers over text or image context, using Luna.

These statements were checked on October 3, 2026; they do not establish current general availability. Choosing "queue" from a short list makes the answer format explicit, but still does not show whether downstream failures or waiting time were evaluated.

AlphaGo used learned policy and value networks to suggest promising moves and estimate how favorable a position was. Monte Carlo tree search combined those estimates with rollouts, or simulated play, to guide further investigation. The 2016 Nature paper describes how those parts worked together to inform the choice.

In Game 2 of its 2016 Go match against Lee Sedol, AlphaGo's 37th move was so unconventional that professional commentators initially thought it was a mistake. As DeepMind later recounts, roughly a hundred moves later the stone's position helped AlphaGo win the game.

That retrospective does not reveal the exact computation that caused the move, and the Nature paper predates the match. Searching candidate explanations for a queue design would still leave the team needing evidence about the app. Go's search evaluates positions under defined rules; software decisions often begin with disputed assumptions.

The queue has no Go board

Go supplies precise rules and executable transitions. The report-export decision begins with unresolved requirements: how long a user can wait, when the data must be captured, what happens after a failure, and what the team can operate. The assistant may need to help define the problem before searching it.

Suppose a simulation assumes every job runs once and the worker always acknowledges it. A queue could look excellent under those assumptions while the app's real failure cases remain unexamined. If the LLM generates both the simulation assumptions and the evaluation method used to judge the results, agreement between them can reflect one shared mistake.

Some parts of this decision are testable now. Others concern future traffic, customer behavior or requirements the team cannot yet predict. Keep that distinction visible. The following checks are proposals for the engineering team to run or review on the imagined customer app, not experiments performed for this article.

Evidence to seek before accepting the queue recommendation
QuestionCheck to planHow it could affect the choice
Must the user wait for the result?Agree on acceptable waiting time; measure report runs under realistic inputs and load.If the direct request meets that requirement, it may be simpler. If it cannot and background completion is acceptable, investigate the queue.
What if a report is produced but the job is not acknowledged?Exercise that failure and a retried job; observe reports and notifications produced.A queue still needs explicit duplicate-handling behavior.

These two checks could challenge the recommendation before the team commits to it. They leave questions about future traffic and operating burden open. A simulation can help explore those questions, but its assumptions need to remain visible.

It may also turn out that the problem was framed badly. Optimizing completion time is of little use if the customer actually needs a reproducible report from a fixed snapshot. More search through that formulation would not correct the missing requirement on its own.

What would count as progress?

To test whether an added search or checking layer helps, the engineering team should define success and what should change the recommendation before comparing approaches. Use an AI agent built around a strong reasoning model as the comparison baseline, with the same information and tools, then compare at matched resource budgets or show the cost-quality tradeoff. Track time, model and tool costs, and attempts; use observed outcomes where available and expose assumptions where they remain.

An evaluation based on the same imagined traffic can favor the same mistake. Bound the investigation too: the agent harness article explains how to limit a run and hand off evidence, but those controls do not establish that the agent's recommendation is correct.

Ask what could change the choice

The queue recommendation is still open. A compact decision note can make the next step explicit:

  • Choice and owner: the engineering team chooses direct report generation or a queue.
  • Required outcome: record the agreed waiting time and failure behavior.
  • Evidence and change condition: name the next check and which result would change the recommendation.
  • Unknowns: record unresolved assumptions, including operating cost.

For a founder handing this to an engineer, the question can be plain: "What requirement makes this need to run in the background? What is the simplest alternative? Show me what happens if the report is created but the job retries, and what result would change your recommendation."

The process behind an AI recommendation, including checks run by the engineering team, should be judged by whether it exposes assumptions, checks relevant consequences and supports revising the recommendation when the evidence changes. System 1/System 2 can name those roles, but the label reveals little about how well they are performed. The open question is which forms of investigation improve the decision enough to earn their cost, especially when the situation itself is uncertain.

All articles