Observed Signal · Jul 3, 2026 · Technical Publication · Source: DEV Community · Impact: 2/5 · Sentiment: Positive

ReWOO: Plan Tool Calls Upfront, Only Two Model Calls

Executive Signal Summary

The article introduces ReWOO (Reasoning WithOut Observation), an agent design that asks an LLM to plan every tool call up front, executes the planned tool calls in plain code without further model calls, then calls the LLM once more to produce the final answer. By splitting the agent into a Planner (one LLM call), a Worker (zero LLM calls) and a Solver (one LLM call), ReWOO keeps model-call count constant at two regardless of chain length, reducing token retransmission, inference cost and latency for multi-hop queries. The piece explains the trade-off: ReWOO is efficient when the step sequence is predictable but less adaptable to surprising or errorful tool results, and it suggests hybrid approaches (re-planning on surprises). Published on dev.to on 2026-07-03.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

Presents a practical agent architecture that materially reduces LLM calls, tokens and latency for multi-hop workflows — relevant for teams deploying cost-sensitive RAG/agent systems — but is not a platform-level or industry-shifting announcement.

SIGNAL RADAR

Track DEV Community Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • ReWOO (Reasoning WithOut Observation) asks an LLM to produce a complete plan of tool calls up front and then performs the tool calls in code before a final LLM call.
  • ReWOO splits the agent into three roles: Planner (1 LLM call), Worker (0 LLM calls), and Solver (1 LLM call).
  • Model-call count under ReWOO is constant at 2 regardless of the number of hops; ReAct typically requires k+1 model calls for k hops.
  • ReWOO reduces repeated transcript token usage and can lower latency and inference cost, but it sacrifices mid-run adaptability if tool results are unexpected.
  • Article published on dev.to with a metadata publication date of 2026-07-03.
Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: DEV Community•Published: Jul 3, 2026
Original Coverage Title: “ReWOO: plan every tool call up front, then call the model only twice”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

Large Language Models (LLM) & AIMay 24, 2026

Plan-and-Solve Agent Architecture: Plan First, Then Execute

This technical article explains the Plan-and-Solve agent paradigm: use an LLM to generate a complete, ordered plan of 3–7 concrete steps (Plan phase), then execute each step sequentially with tool-enabled sub-agents (Solve phase). The author implements the architecture using LangGraph's StateGraph abstraction, showing a Plan node, an Execute node that embeds a ReAct sub-agent (for tool calls like web_search and calculator), a Replan node, and a Finalize node. Demos highlight practical failure modes—information loss when step results are transmitted as short natural-language summaries, planner over-splitting, and tool-specific errors—and propose engineering fixes (structured step outputs, dedicated collected_data state, planner-annotated data flow). The article concludes with five production findings and guidance on when to choose ReAct versus Plan-and-Solve.

Read assessment
Large Language Models (LLM) & AIMay 4, 2026

Open‑Source RightModel Tackles Token Consumption Anxiety

Developer Regnard Raquedan published an article on May 4, 2026 describing RightModel, an open-source tool that recommends the best LLM for a task without making live LLM calls in the default request path. RightModel uses a human-owned, versioned ruleset to classify task types and map them to model tiers; pricing data is refreshed asynchronously (via OpenRouter and a scheduled workflow using Google Cloud Scheduler) to avoid runtime API calls. For ambiguous cases the app exposes a user-triggered "Deep Analysis" escalation that calls an LLM (currently Gemini 2.5 Flash). Raquedan frames the architecture as an instance of a broader pattern he calls "Precomputed AI," which shifts reasoning out of real-time request paths into asynchronous build pipelines with explicit staleness controls and escalation paths.

Read assessment
Large Language Models (LLM) & AIMay 27, 2026

ARTIST: RL-Powered Tool Use for LLM Agents

ARTIST (Agentic Reasoning and Tool Integration in Self-improving Transformers) is a Microsoft Research training framework that teaches LLMs when and how to call external tools by using outcome-only reinforcement learning. Published (paper/arXiv 2505.01441) and described in the article, ARTIST interleaves tool calls inside the model’s chain-of-thought tokens rather than appending results as separate turns, and trains with GRPO (Group Relative Policy Optimization) using final-answer rewards, format checks, and efficiency signals. At 7B scale, ARTIST reportedly outperforms GPT-4o on multiple math and multi-turn function-calling benchmarks. An independent Effloow Lab proof-of-concept reproduced the interleaving execution loop in a minimal Python sandbox and observed improved accuracy and fault-tolerant recovery. The paper focuses on a training recipe rather than a production SDK; some public implementations (TRL, verl) can approximate parts of the GRPO loop.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.