What happens between a message and a finished answer?
A source-backed comparison of execution, stopping rules, tool discovery and channel routing—with actual recorded turns kept separate from architecture descriptions.
Source refresh, September 28: 25 relevant files were reviewed at the newly fetched OpenClaw and Hermes commits. Eliza develop still matches the reviewed base. The update below explains meaningful changes; recorded runs keep their actual older version labels.
Prepared September 28, 2026Recorded model/tool tracesPinned source reviewNot a reliability leaderboard
The direct answers
These sections answer Shaw’s questions separately for each implementation. The flow labels are explanations; exact code identifiers appear in monospace and the linked source.
Without an evaluator, how do they stop?
The model proposes more tool work or an answer. Runtime code accepts, continues, finalizes or aborts that proposal using its own rules. Both OpenClaw and Hermes run code that checks whether a turn may finish.
Stopping, satisfying the request and delivering the reply are different outcomes.
Is it always one model call?
No. A direct reply may need one. Our recorded file-read turns needed two: request the read, then answer from its result. Discovery, retries, finalization and additional work can add calls.
Eliza has three named stages, not an invariant three-call limit.
What happens across channels?
Adapters identify the sender and conversation, apply admission rules, select a session/history scope, then return the result through a delivery route. Sharing depends on configuration.
A session key is not, by itself, proof of complete long-term-memory isolation.
Correction to the initial shorthand: “They have no evaluator and just trust the model more” is too broad. OpenClaw has completion/delivery guards and conditional finalization; Hermes has final-text recovery and configurable stop gates. Eliza explicitly requests structured routing, planning and completion judgments.
Follow the data
First the source-level structure; then select an actual scenario below. A model call chooses or assesses work. The runtime—not the model itself—executes the tool.
OpenClaw
Source verified
Admit and routeResolve agent, session, policy and execution owner.
Prepare contextInstructions, skills/workspace context, history and permitted tools.
Model / tool roundsTool requests execute and return results; continuation can require another response.
Resolve completionApply terminal, pending-work and delivery rules; finalization may add a call.
Deliver and settleRecord the lifecycle result and source-reply delivery separately.
↳ More tool work or a continuation signal returns to the model/tool round.
Hermes
Source verified
Admit and select sessionAdapter/source authorization and configured session key.
Build model inputInstructions, selected session history and enabled tool definitions.
Model proposes next stepTool calls go to a validated execution round; results are persisted and returned.
Final-text and stop gatesNo-tool answers can still require repair or configured verification.
Persist and deliverAccepted final text returns through the source adapter and delivery ledger.
↳ Tool results, recovery or a stop gate can request another completion.
Eliza
Source + live trace
Admit and normalize messageBind actor, room, channel and request context.
Response handlerDecide response/work, intents and context. Simple paths may finish here.
Planner → toolsChoose operations from selected schemas; execute through runtime checks.
EvaluatorAssess outcomes and remaining work; return a structured decision.
Guard and deliver replyRuntime checks completion/evidence and delivery requirements.
↳ Continue, replan or restore context when the request remains incomplete.
Actual recorded scenarios
Eliza-only application checks: these are not matched competitor benchmarks.
Read the count correctly. A model attempt, a successful provider response, a tool execution and a background task are not the same unit. Recorded transport failures are retained; unavailable usage is not zero. These scenarios do not establish an apples-to-apples performance ranking.
Who decides the turn is over?
“The model wrote text” is not a universal terminal condition. The framework combines model output with execution, interruption, recovery and delivery state.
Boundary
OpenClaw
Hermes
Eliza
More tool work
Execute admitted tools and continue through the selected agent backend.
The conversation loop branches on tool calls, persists them, executes, appends results and continues.
The planner supplies a queue; runtime evaluation can continue it, replan or restore context.
Proposed final answer
The core loop can end when no executable tool calls, continuation signal, pending steering or follow-up require another round. The outer runtime separately checks required replies and delivery; an already delivered reply avoids duplicate generation.
When the model returns no tool calls, code checks the proposed answer for recoverable problems and runs configured stop gates. If accepted, it persists the final answer and exits the loop.
On the evaluated path, structured FINISH/success/coverage must satisfy runtime guards; pending work or failures can prevent acceptance.
No usable answer
A settled tool batch can receive a tool-free finalization pass when a reply is required.
Empty/stalled/truncated or dropped-tool-call outputs can trigger bounded recovery.
Missing or invalid completion/reply information can require repair or another planning/evaluation step.
Explicit verification
Completion and delivery guards exist; this is not a claim of an Eliza-style evaluator after every action.
Conditional stop gates include verification hooks. verify_on_stop defaults off at the studied pin.
An explicit evaluator model call is part of the ordinary planned path, but simple replies/navigation can finish earlier.
Limits / cancellation
Abort and execution-owner rules apply. A wait timeout alone need not cancel the underlying run.
Iteration/shared budgets and interruption handling can end or recover the run. The constructor’s iteration-limit parameter defaults to effectively unbounded; callers/configuration can supply finite limits. There is no universal ten-call cap.
Cancellation, failures and loop/budget bounds can terminate an incomplete request.
OpenClaw’s backend matters
In embedded Responses mode, a completed response with end_turn: false can continue even when it contains only text. In the Codex app-server path, native turn completion owns the terminal outcome; a quiet stream or an assistant message alone is insufficient.
A model can incorrectly propose that work is done. A separate evaluator can also make a wrong judgment. Our acceptance checks inspect actual tool results, navigation/delivery and stored records—not merely a success-shaped answer.
An exhausted budget, denied tool or interrupted turn must not be presented as a completed task.
Messages from different channels
Before calling a model, channel adapters identify the sender and conversation. Runtime code selects the history to load and the destination for the reply.
1 · IngressApp, Discord, DM or group event
2 · AdmissionSender, account, mention and permission policy
3 · Context keyChoose agent and session/room history
4 · ExecutionRun the model/tool workflow
5 · DeliveryReturn through the bound channel/thread route
OpenClaw
Configuration matters
Channel/account/peer bindings choose an agent and session. DM scope can share the agent’s main session or isolate peers, channels and accounts.
At the reviewed pin: the resolver’s DM fallback is main and group fallback is per-group; installation settings can override both. Thread and account handling depends on the route/configuration.
Do not assume every DM has independent history. Session queues and delivery ownership are distinct from long-term memory policy.
Hermes
Pinned gateway rules
Adapters normalize platform, chat, thread and sender identity. Authorization runs before an ordinary turn, and the session key includes profile/platform/chat scope.
At the reviewed pin: DMs use chat identity; groups default to per-user sessions, while threads default to shared sessions unless configured otherwise.
The reply uses the current adapter and chat/thread anchor. Profile-scoped memory providers mean session separation is not a blanket privacy guarantee.
Eliza
App + Discord source
Ingress becomes a message with actor/entity, room, source and channel type. In Discord, the channel ID becomes a runtime-scoped room; guild/server context becomes a world.
Channel/DM access and mention policies govern admission. The message service receives the normalized message plus a transport callback for the response.
App conversations and connector rooms are distinct. Providers can include broader authorized facts; room identity alone is not proof of total memory isolation.
Evidence boundary: the OpenClaw/Hermes runs reconstructed here used CLI sessions. Channel behavior above is source-verified, not a matched Discord/Telegram/Slack delivery benchmark. We have not proved universal cross-channel privacy or equal reliability across these systems.
What actually enters the model request?
“Context” means the instructions, messages, tool definitions and results supplied to a particular model call. A stored conversation, an installed plugin and a tool visible in that request are different things.
OpenClaw
Runtime code prepares instructions and workspace context, loads the selected session’s history, and exposes tools allowed by configuration and policy. After a tool runs, its result becomes available to the next model round.
The recorded Cerebras and Responses setups used different request protocols. The Responses setup can refer to earlier server-retained context; request size alone does not measure total model context.
Hermes
Runtime code builds instructions, loads the selected conversation and enabled tool definitions, then calls the model. Tool calls and results are persisted and included as the conversation continues. Memory and skills depend on the enabled configuration. Current Hermes source can pin and reuse a session’s workspace context and prompt inputs; it does not necessarily rebuild every input from scratch on each call.
The original run exposed 24 definitions. This is an observed configuration, not a permanent limit or a claim that every installation has the same tools.
Eliza
The response handler receives its own instructions and selected request context. Its decision helps select context and domain operations for the planner. The evaluator receives accumulated results and completion instructions rather than the same executable tool catalog.
Each stage is a separate model request with a different purpose and input. Supplied history and provider output can be repeated; exact provider token totals are shown below, not inferred from text length.
Search, describe, execute: search finds candidate tools; describe returns an operation’s argument schema and instructions; execution actually invokes it. Catalog search and describe do not themselves execute the Calendar read operation. Context providers may separately supply facts. Some tools are visible from the start and need no discovery. Eliza’s load mode can combine search with making selected schemas available on the next planner call.
Terms: a tool schema describes accepted arguments; a session or room identifies conversation history; a channel adapter connects a transport such as Discord; a stop check is runtime logic that accepts or rejects ending a turn. BM25 is a lexical ranking method based on term occurrence and document length, not evidence of vector search.
Which tools can the model see and use?
Separate what is installed, what is visible to the model, what was actually called, and what input was cached. Collapsing these categories produces misleading comparisons.
Agent
Tools shown before discovery
How it finds additional tools
What happens next
OpenClaw
In the recorded setup, the model initially saw 12 tool definitions, including file/process operations and tools for finding, describing and calling additional tools. This count belongs to that setup, not every OpenClaw installation.
Ranks catalog text using word matches and BM25, prioritizing exact tool names and literal matches before broader query expansion.
Search returns candidate tools. The model can request a candidate’s full argument definition, then request its execution. A tool already shown directly can be called without this discovery sequence.
Hermes
In the recorded setup, the model initially saw 24 tool definitions, including file/terminal, browser/web, memory, skills, delegation, and tools for finding and inspecting more tools. This count belongs to that setup, not every Hermes installation.
Ranks catalog text using BM25, prioritizing exact names and checking distinctive query words and how much of the query a candidate matches.
Search, describe and call are separate operations for additional tools. Tools already shown directly—such as the file reader in our recording—can be called immediately.
Eliza
In the recorded Notes + Calendar request: 2 definitions in the response handler, 9 in the planner, and 0 executable tools in each evaluator call. The evaluator instead returns a structured decision. The first model call decides how to handle the message and whether more context is needed. If planning is required, the planner receives selected action definitions—for example, open a view, read Notes, or read the next Calendar event—plus controls for returning an answer or finding more actions.
Code uses the request and the handler’s decision to select relevant actions. Search combines exact names, patterns, keywords and BM25 word ranking, with context-based weighting and a limit on loaded definitions. Optional embedding scores were not supplied by the inspected callers.
The planner can search for more actions. Load mode makes the selected actions’ full argument definitions available on a later planner call. Describe mode only explains an action. Neither mode executes the requested Notes or Calendar action.
Concrete cache evidence: Eliza combined read
58,698foreground input tokens
29,696cached subset
4.78swhole foreground turn
Cached: 50.6%Other input: 49.4%
One successful September 28 run. Cached input varied by stage; this is not a cache-hit guarantee or proof that every input was necessary.
Three accounting rules
Cached tokens still occupy context. They are a subset of input, not an additional quantity.
Small wire payload does not mean small context. Server-retained response chains can carry earlier instructions, schemas and history.
Characters are not token counts. Section sizes and provider usage are reported separately where exact attribution is unavailable.
The original OpenClaw Responses lane retained context through response references; it must not be merged with the Cerebras lane as one benchmark.
Latest Eliza stage measurements
Both evaluator rows are calls to the same evaluator component with different accumulated results, not different agents or evaluator types.
Stage
Input
Cached subset
Model latency
Response handler
11,362
9,216
649 ms
Planner
16,050
2,048
1,097 ms
Evaluator call 1 — after Notes
14,957
4,096
1,014 ms
Evaluator call 2 — after Calendar
16,329
14,336
702 ms
Separate background task: one call, 9,149 input tokens, zero reported cached input, 125 output tokens, 972 ms task duration. Recorded per-call model latencies sum to 3,462 ms; total turn latency is 4,780 ms. The under-three-second target was not met.
What this report proves—and does not
Architecture descriptions, recorded outcomes and unresolved limits are deliberately separate. No private transcripts, credentials or raw account data are embedded in this file.
Established
Recorded simple file reads in both OpenClaw and Hermes used two model calls and an actual tool result.
Both frameworks have runtime rules around proposed endings; “no separate evaluator” is not “no checks.”
Eliza’s recorded compound request used handler, planner and two evaluator calls, with actual navigation and read results.
Tool availability, tool execution, final text and confirmed delivery are different evidence.
Not established
That Eliza is faster or more reliable overall.
That the competitors lack guardrails or that every turn follows one fixed call count.
Equivalent Notes/Calendar/reminder integrations across all three.
Latest-upstream behavior beyond the stated source pins, or universal memory isolation for every configuration.
New shared-fixture tests — partial results
These runs use identical local JSON files with note and event records. They test reading, conditions and conversation recall; they do not test equivalent native Notes/Calendar integrations. OpenClaw runs the pinned 2026.9.6 build; the new Hermes run uses revision 9460cc11, different from the older recordings above. The model is Qwen in both. A buffering test proxy enforces the budget and allowed operations, so these are not normal streaming-latency measurements.
Scenario
OpenClaw
Hermes
Eliza
Read latest note and next event
Pass · 2 model calls. Actual file results returned; exact note body and local event times verified.
Not completed. The model requested filename search, which the test proxy did not support; execution was blocked. This is a test-harness incompatibility, not an incorrect answer.
Not run yet.
Do not read the second resource when the condition is false
Pass · 2 model calls. Only the control file was read; no second-resource read was proposed or executed.
Not run after the harness stop.
Not run yet.
Recall the first answer without rereading
Pass · 1 model call. Retained conversation/tool history verified; no new read proposed or executed.
Not run after the harness stop.
Not run yet.
Earlier transport attempts failed before a model result; the HTTP client was corrected and a separate diagnostic succeeded before these runs. All failed attempts remain in private evidence. The resumed batch made 7 provider requests: 5 OpenClaw foreground calls, one Hermes title call and one Hermes foreground call. These partial results do not establish a winner. Test restrictions and version differences must accompany any comparison.
What changed in the newer source?
This is a bounded source review of context, stopping, tool discovery and channel handling—not a rerun on the newest binaries or an audit of every feature. OpenClaw: 0b3221fbfbf3. Hermes: 6d42313deee6. Eliza develop: d0cd0399e710. Checked September 28, 2026.
OpenClaw
The core continuation conditions and finalization eligibility remain consistent with the explanation above. Finalization dispatch was reorganized by backend. Prompt composition changed, including stable-prefix handling and conditional memory sections, so old token counts cannot predict current requests.
Search ranking is unchanged in the compared file. Discord routing now carries a bound agent into fallback route resolution. Tool-result handling also preserves certain error explanations for code-mode callers.
Current code adds a check for runaway repeated final text before accepting it. Optional verification remains off by default. Workspace context and prompt inputs can be pinned to a session and reused; tool changes can refresh those inputs.
Queued-message handling preserves pending reply obligations. SimpleX access rules now use stable contact IDs rather than display names; this is a platform-specific change, not a new Discord default.
The fetched develop revision has not changed since the base used for this source review. Separate local fixes address the launcher’s unnecessary app-catalog request, rejecting a trigger with no schedule, and duplicate rendered conversation context.
Those local fixes do not retroactively change older trace counts. The reminder and matched-fixture acceptance results must identify the code actually running, and no local fix is presented here as a universal reliability guarantee.
The refresh inspected 25 current files and completed 23 old/new file comparisons. One old Hermes file fetch was rate-limited; its current verification default was read directly, without claiming a byte-for-byte comparison. Source changes are not measured speed improvements.