Agentic LLM Reasoning (I): From Chain of Thought to Search, Reflection, and RLVR
Contents
-
- 1. First define the boundary: what is an Agentic LLM?
- 2. Why a model may “know” but still fail to “do”
- 3. Show examples before asking: in-context learning
- 4. Make the process visible: Chain of Thought
- 5. Do not trust the first path: verification and voting
- 6. Let programs execute precisely: PAL, interpreters, and debuggers
- 7. From one chain to a backtrackable search tree
- 8. Failure is not the end: self-reflection and reusable experience
- 9. Reasoning also needs research: retrieval augmentation
- 10. Write successful reasoning back into the model: RLVR and GRPO
- 11. A unified view: the model proposes; the system structures and verifies
- 12. After Reasoning: Action and Interaction
- References
A model that can answer a question is not necessarily able to complete a task.
A conventional chatbot receives a prompt and returns a response. An agent must decide what to do next in a changing environment, act, inspect the result, and choose whether to continue, backtrack, or try another path. That transition begins with reasoning.
A complete agent system needs at least three capabilities:
- Reasoning: analyze state, decompose problems, compare paths, and form decisions.
- Action: call tools and translate decisions into external operations.
- Interaction: observe outcomes, exchange information with the environment or other agents, and revise the strategy.
This is the first article in the Agentic Large Language Models series. Rather than listing reasoning terms in isolation, it follows one question: when one generation or one reasoning path is unreliable, what can the system add? The next two articles will cover Action and Interaction.
Series: (I) Reasoning · (II) Action (coming next) · (III) Interaction (coming next)
1. First define the boundary: what is an Agentic LLM?
An Agentic LLM can be defined as:
An agent that receives natural-language or multimodal input from its environment, reasons to make decisions, and takes autonomous actions that affect the environment in pursuit of a goal.
A conventional chatbot often follows:
|
|
An Agentic LLM operates in a loop:
|
|
Agent capability therefore depends on more than the base model. It also depends on reasoning strategies, tools, external memory, state management, environmental feedback, and the harness surrounding the model.
1.1 From training a model to using it
A modern LLM typically passes through the following stages:
| Stage | Purpose |
|---|---|
| General corpus acquisition | Build a large unlabeled dataset |
| Pretraining | Learn language, knowledge, and patterns through self-supervision |
| Supervised fine-tuning | Adapt to tasks with labeled instruction–answer data |
| Instruction tuning | Improve natural-language instruction following |
| Preference alignment | Align behavior through RLHF, DPO, RLVR, or related methods |
| Training and deployment optimization | Reduce cost with LoRA, mixed precision, or distillation |
| Inference | Solve tasks through prompts, context, tools, and decoding strategies |
The first six stages primarily alter model parameters. Much of agentic reasoning happens in the final stage: without necessarily changing the weights, a system gives the model more intermediate computation, candidate paths, feedback, and external machinery.
Reasoning is therefore not only about what happens “inside” one model call. It also includes how an external algorithm organizes calls, preserves state, and uses verification results.

The taxonomy spans step-by-step prompting, ensembles, search, self-reflection, retrieval, and reinforcement learning. The rest of the article follows this progression.
2. Why a model may “know” but still fail to “do”
Consider a typical GSM8K problem:
Romeo bakes four trays containing two dozen cookies each. If the cookies are shared equally among sixteen people, how many does each person receive?
A person writes:
|
|
Yet GPT-3 175B achieved only about 15% accuracy on GSM8K when the benchmark was introduced. The model may know arithmetic and understand every word while still failing to organize a reliable multi-step calculation.
A direct autoregressive answer implicitly requires the model to retrieve facts, identify variables, order operations, execute each step, and prevent early errors from propagating. Reasoning methods turn this hidden process into a longer, more controllable computational path.
From a probabilistic perspective, direct generation still follows the autoregressive factorization:
Reasoning methods do not replace this basic generation rule. They add intermediate steps, alternative paths, verification results, or observations to the condition on which later tokens are generated.
3. Show examples before asking: in-context learning
In-context learning (ICL) places examples and a query together in the context window. The model does not update its parameters; it infers the task pattern from the current prompt.
Zero-shot
|
|
Few-shot
|
|
Few-shot examples do more than specify formatting. They steer the model toward relevant knowledge and a suitable solution pattern. Ordinary ICL, however, can still compress the process into one forward generation. The natural next step is to include intermediate steps in the output itself.

The same problem enters a different computational path depending on whether the prompt contains examples, reasoning traces, or executable code.
4. Make the process visible: Chain of Thought
Chain-of-Thought Prompting asks the model to generate a sequence of intermediate reasoning steps rather than only the final answer:
|
|
Few-shot CoT demonstrates complete reasoning traces. Zero-shot CoT can be elicited with a phrase such as “Let’s think step by step.” In the original work, an example GSM8K result improved from roughly 16% to 47%.
One intuition is that every intermediate token re-enters the context for subsequent prediction. The model obtains more serial computation and gradually narrows the probability space around the final answer.
CoT creates a new problem as well: longer chains provide more opportunities for error accumulation. A wrong early assumption can produce a coherent but incorrect conclusion.

Direct prompting encourages an answer guess; a CoT demonstration teaches the model to calculate before concluding.
5. Do not trust the first path: verification and voting
5.1 Self-verification
Self-verification treats the conclusion as a condition and checks it against the original problem:
|
|
The same LLM may act as generator and verifier, or a separate evaluator can be used. The important move is not merely “think again,” but changing the form of the problem so that mistakes encounter a new constraint.
5.2 Self-consistency
Self-Consistency samples diverse reasoning paths and aggregates their final answers:
|
|
Complex problems often admit multiple valid paths to the same correct answer, while mistakes tend to diverge. Majority voting reduces the impact of one unlucky sample.
The paper reported a 17.9 percentage-point improvement over CoT on GSM8K and gains on several arithmetic and commonsense benchmarks. The tradeoff is inference cost: twenty sampled paths consume far more tokens and latency than a single generation.
If $r_i$ is the $i$-th reasoning trajectory and $a(r_i)$ is its extracted final answer, majority voting can be written as:

Correct paths tend to converge on one answer, while errors are more likely to scatter. Self-Consistency exploits that difference.
Reasoning quality is therefore not solely a property of model weights; it can also be purchased with inference-time compute.
6. Let programs execute precisely: PAL, interpreters, and debuggers
Natural language is flexible but ambiguous and unreliable for exact execution. A practical division of labor lets the LLM understand and decompose a problem while an interpreter executes it.
6.1 Program-Aided Language Models
PAL translates a natural-language problem into an executable program and delegates evaluation to a runtime:
|
|
The model is good at mapping a problem into program structure; Python is good at exact variables, loops, and state. PAL with Codex reported about 72% accuracy on GSM8K and outperformed larger models using natural-language CoT on multiple symbolic tasks.
The division of labor can be summarized as a model generating a program $g_{\theta}(x)$ and an executor producing the final answer:

The LLM translates intent into a program; the interpreter performs the unambiguous execution.
6.2 Interpreters, debuggers, and self-debugging
Execution errors, failed tests, and runtime output can be returned to the model:
|
|
Self-Debugging shows that an LLM can identify mistakes by inspecting execution results and explaining its own code. Ground truth now comes from compilers, interpreters, and tests rather than from the model’s confidence.
This also addresses the knowing–doing gap: a model may describe the correct algorithm while failing to maintain every state transition in tokens. Letting a conventional program execute the algorithm is often more accurate and cheaper than simulating execution through language.

Natural language is easier to read, while PDDL and other formal languages represent states, actions, and constraints more precisely.
7. From one chain to a backtrackable search tree
CoT follows a single path. Search-based methods generate several candidates at each step and create a backtrackable state space.
7.1 Tree of Thoughts
Tree of Thoughts (ToT) treats intermediate thoughts as search nodes:
|
|
The system performs three operations:
- Generate candidate thoughts from a state.
- Evaluate how promising each candidate is.
- Search with BFS, DFS, or another policy.
BFS can be summarized as:
|
|
DFS follows one branch and backtracks when it fails or reaches a boundary.
If $s_t$ denotes the search state at step $t$ and $z_t$ is a candidate thought, ToT can be summarized as generating new states, scoring them, and retaining the best $b$:

CoT expands one chain, Self-Consistency samples independent chains, and ToT branches, scores, and prunes intermediate states.

BFS keeps promising states at each depth; DFS follows one branch and backtracks after failure.

In Game of 24, the model both proposes candidate steps and evaluates which intermediate results deserve more search.
ToT exposes paths, supports rollback, and makes constraints easy to add. Its expansion policy is usually fixed by humans: branching factor, retained candidates, and stopping rules are configured in advance. Learned policies can dynamically choose where to search next, at the price of training and debugging complexity.
8. Failure is not the end: self-reflection and reusable experience
When an LLM is called separately as actor, critic, or evaluator—and feedback is used to construct the next prompt—the overall system performs engineering-style self-reflection.
This is usually not mysterious introspection inside one call. It is an external control algorithm repeatedly invoking the model, saving trajectories, and restructuring context.

An action changes the environment; the resulting state and reward become input to the next decision.
8.1 Self-Refine
Self-Refine uses a simple loop:
|
|
The same model can generate, critique, and revise without changing its weights. The method works best when quality criteria can be expressed in language.

Self-Refine improves an output through a generate-feedback-revise loop without updating model weights.
8.2 ReAct: the bridge from reasoning to action
ReAct interleaves thoughts, actions, and observations:
|
|
CoT can only continue from the model’s current context. ReAct can retrieve evidence, call a tool, or query an environment and incorporate the observation into the next reasoning step. This grounding helps reduce hallucinations from purely linguistic reasoning.
ReAct crosses the boundary between Reasoning and Action, making it the starting point for the next article in this series.

ReAct writes external observations back into the reasoning trace instead of relying only on parametric knowledge.
8.3 Reflexion: extracting reusable experience from a trajectory
Reflexion adds three roles around a ReAct-like agent:
- Actor generates reasoning, actions, and a trajectory.
- Evaluator determines success and provides feedback.
Reflector compresses failure into a reusable natural-language lesson.
1 2 3 4 5 6 7 8 9Actor trajectory ↓ Evaluator score / feedback ↓ Reflector creates a lesson ↓ Store in memory ↓ Actor tries again
Short-term memory stores the current reasoning and action trace. Long-term memory stores reusable reflections across attempts.
|
|
The relationship is easy to remember:
ReAct = Think → Act → Observe
Reflexion = ReAct → Evaluate → Reflect → Remember → Try Again

The actor attempts the task, the evaluator judges it, and the reflector compresses failure into a reusable lesson.

Reflexion separates the current trajectory from experience retained across attempts.
8.4 Buffer of Thoughts
Buffer of Thoughts distills high-level solution structures from past tasks into thought templates stored in a meta-buffer:
|
|
Instead of solving every task from scratch, the agent accumulates reusable cognitive patterns such as “identify constraints, enumerate candidates, then verify.”

Buffer of Thoughts stores transferable solution templates rather than every detail of every previous trace.
8.5 Other prompt-improvement loops
| Method | Core idea |
|---|---|
| Progressive Hint Prompting | Feed the previous answer into the next prompt until the result stabilizes |
| Self-Discover | Select and compose reasoning modules suited to the current problem |
| Prompt Improvement | Rewrite the next prompt from evaluation feedback rather than updating weights |
Self-reflective systems share two challenges:
- Multiple roles and prompts interact unpredictably and can amplify errors.
- States, actions, observations, rewards, and reflections eventually overflow the context window.
Reflection therefore needs summarization, selective memory, context resets, and external state storage.
9. Reasoning also needs research: retrieval augmentation
Reasoning should not remain closed inside parametric memory. RAG connects unstructured documents, databases, and knowledge graphs:
|
|
Basic RAG retrieves once before answering. Adaptive retrieval lets the model decide when to search, what to search for, and whether the available evidence is sufficient.
The connection to self-reflection is direct: an evaluator can detect missing or stale evidence and trigger retrieval; the retrieved result becomes the next observation. Reasoning evolves into a think–retrieve–verify process.
10. Write successful reasoning back into the model: RLVR and GRPO
Most methods above operate at inference time without changing parameters. High-quality reasoning traces can also feed back into training.
10.1 RLVR
Reinforcement Learning with Verifiable Rewards uses fast, deterministic verifiers instead of a subjective reward model:
- Is the mathematical answer correct?
- Does the code pass its tests?
- Are format and constraints satisfied?
Does a proof checker accept the proof?
1Prompt → Sample Reasoning + Answer → Verifier → Reward → Policy Update
Rewards are objective and scalable, but only some real-world tasks have cheap, reliable verification functions.
If a verifier returns reward $R(\tau)$ for trajectory $\tau$, the training objective can be summarized as:
10.2 GRPO
Group Relative Policy Optimization samples a group of answers to the same prompt and estimates advantages from relative group scores, reducing dependence on a separate critic model. It extends the diverse sampling intuition of self-consistency into training: the system does not merely select a better answer at inference time; it increases the probability of high-reward reasoning policies.
For answer $i$ in a group, an intuitive relative-advantage estimate is:
|
|
11. A unified view: the model proposes; the system structures and verifies
LLMs belong to connectionist AI: knowledge and capability are distributed across neural parameters and learned from data. Search, logic, planning, interpreters, and state machines resemble symbolic AI: rules and operations are represented explicitly.
Modern agentic systems combine both:
|
|
This is often compared with System 1 and System 2:
- Fast reasoning directly uses learned associations.
Slow reasoning introduces intermediate steps, search, planning, and tools.
1 2 3 4 5 6Fast: 17 × 6 → 102 Slow: 17 × 6 → 10 × 6 + 7 × 6 → 60 + 42 → 102
The analogy should not be taken literally. LLMs still operate through next-token prediction. A more precise engineering description is that “slow thinking” supplies additional serial tokens, parallel candidates, external state, and verifiable feedback.
12. After Reasoning: Action and Interaction
The path from ordinary generation to agentic reasoning is now visible:
|
|
Reasoning alone cannot change the world. A correct plan remains text if it cannot call tools; an action cannot form a loop if the system cannot observe its outcome.
The next two articles will examine:
- Action: planning, tool use, world models, vision-language-action models, and the translation from tokens to operations.
- Interaction: observations, environmental feedback, multi-agent coordination, state management, and continual adaptation.
Reasoning decides what should happen next. Action determines how to do it. Interaction tells the system what happened afterward. Together, they form an Agentic Large Language Model.
References
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- PAL: Program-Aided Language Models
- Teaching Large Language Models to Self-Debug
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- Self-Refine: Iterative Refinement with Self-Feedback
- ReAct: Synergizing Reasoning and Acting in Language Models
- Reflexion: Language Agents with Verbal Reinforcement Learning
- Buffer of Thoughts: Thought-Augmented Reasoning with Large Language Models
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Author zhengxz
LastMod 2026-09-29