September 2026
Text-to-World Bench: Can agents turn a text into a real-world result?
What this benchmark measures
We treat each task as an end-to-end request. Alongside the final outcome, we examine how assistants use context, ask for clarification, handle obstacles, communicate progress, and support their completion claims. We also record response and completion latency and assess the conversational experience, including tone and cadence. The central question is not whether an assistant can understand the request, but whether it can carry that request through to a verifiable result.
Across 7 agents and 11 tasks, we tested whether a short message could become a reservation, delivered meal, purchased ticket, or flight booking ready for checkout. Existing efforts such as Assistant Benchmark cover a broad range of capabilities; we deliberately focus on consequential real-world actions that people genuinely want to delegate and that remain difficult for today's agents to complete reliably.
To make comparisons consistent, we reused the same opening prompts, presented tasks in the same order, and reviewed each outcome as a pass or failure.
Task design and evaluation protocol
- Real workflows. Tasks came from the places where we would actually delegate work: reservations, food delivery, movie tickets, and flights. We excluded low-consequence tasks that agents already handle reliably.
- Shared conversations. Each agent remained in one persistent chat. Later requests could therefore test whether earlier constraints and preferences changed the next decision.
- Outcome-first review.The headline score is binary against each task's stated checkpoint. Partial progress and timeouts do not receive fractional credit, even when the intermediate work is useful.
What the runs revealed
Useful assistance can stop short of completion
Agents frequently found the right option, selected a plausible seat or itinerary, and reached checkout without completing the final action. Muse's live browser handoff and Pally's dedicated flight interface made those moments easier for the user, but both could still shift consequential work back to the person who delegated it. Helpful preparation and autonomous completion should therefore be reported separately.
Context matters when it changes a decision
The strongest memory behavior was not repeating a stored fact. It was using prior context at the right decision point. Muse reused a party size and seating preference across reservations, then warned that a San Francisco movie conflicted with an earlier San Jose dinner. Persistent memory needs relevance control, not just retention.
Execution barriers interrupt otherwise promising workflows
Several unsuccessful attempts reached relevant options or checkout before encountering authentication, payment, browser-state, or website obstacles. These observations identify visible points of failure, but they do not isolate the contribution of the underlying model, agent implementation, external service, or account configuration.
Instinct
finished the most tasks, passing seven of eleven. It responded most like a human assistant would: concise, action-oriented, and clear about the next decision.
Muse
(through WhatsApp) produced the strongest qualitative experience in our assessment, with preference reuse, conflict detection, and clear handoffs. Its successful runs also had the shortest median completion time. It passed five tasks; many of its other runs ended in handoffs before independent completion.
Poke Ultra
also passed five tasks. We tested its Ultra plan at $199 per month. Its human operators sometimes completed the request, as with reservations; other handoffs ended with progress updates and no result.
Pally
offered a dedicated flight interface that made its proposed itinerary easier to review than a long thread of messages.
Catch
completed three restaurant reservations, but said it could not place food orders or buy movie tickets.
Folk
built the requested food cart and reached checkout, but could not complete payment.
Arlo
quoted a flight for two instead of one and shifted reservation attempts to phone calls without completing the bookings.
Annotated cases
A reservation becomes context for the next task
“you've got dinner at must be thai in san jose at 5pm that same wednesday, so a ~7:15pm showing in sf would clash pretty hard.”
- Request
- Book dinner at Must Be Thai and Bar in San Jose, then a Spider-Man ticket near 7:15 p.m. in San Francisco that Wednesday.
- What it knew
- Earlier bookings established two diners and standard seating; the confirmed 5:00 p.m. dinner remained in chat.
- Result
- Muse confirmed a 5:00 p.m. table for two with standard indoor seating and a confirmation number. It flagged the movie conflict; we chose to proceed. Muse said AMC had declined its virtual card on earlier attempts, and no ticket was bought.
- Why it matters
- Muse reused dining preferences, then caught a conflict with the next request.
A strong handoff still leaves the task unfinished
“The payment page is ready ... take over through the Muse app or muse.ai, add your card, and hit purchase.”
- Request
- Buy one good seat for Spider-Man: Brand New Day near 7:15 p.m. at the theater closest to Rincon Hill.
- What it knew
- Muse knew about the 5:00 p.m. San Jose dinner and reported three earlier AMC virtual-card declines.
- Result
- Muse selected a center seat and reached payment. Citing earlier card declines, it asked us to finish checkout in its browser. No ticket was issued.
- Why it matters
- Muse found the seat but left payment to us, so the purchase remained unfinished.
A food order reaches the door
“Order’s in — Cantoo, $27.05 on your card with the $2 tip. Arriving 9:16–9:31.”
- Request
- Order Cantoo for delivery with a rice substitution, extra sauce, and salted pumpkin.
- What it knew
- Instinct flagged unavailable sauce. We supplied the address and sign-in code, then approved the revised cart and tip.
- Result
- We verified delivery and that the food matched the approved cart. Instinct quoted $27.05 with a $2 tip and a 9:16–9:31 p.m. arrival window.
- Why it matters
- Instinct resolved an unavailable item, got approval, and completed delivery.
Poke's human handoff secures a table
“one sec, handing this to our human agent to get your table at grand lake kitchen for monday sept 21”
- Request
- Book Grand Lake Kitchen in Noe Valley for two on September 21 between 5 and 9 p.m.
- What it knew
- We supplied seating and contact details; Poke handed the booking to a human operator.
- Result
- Poke reported the booking; we verified it.
- Why it matters
- Poke Ultra's human operator secured the table, a successful assisted handoff.
The human handoff stalls before a ticket purchase
“a human's on it, grabbing you a good seat at kabuki for the 7pm show”
- Request
- Buy one good seat for Spider-Man: Brand New Day near 7:15 p.m. on September 23.
- What it knew
- Poke found a 7:00 p.m. AMC Kabuki showing and routed the purchase to a human operator.
- Result
- We approved 7:00 p.m. Poke said an operator was buying the ticket, then sent another progress update. No ticket arrived within 20 minutes, and it requested no further action from us.
- Why it matters
- Even with a human operator involved, no ticket arrived. The record does not show whether the breakdown was in routing the task or completing the purchase.
The path toward better assistants
In many runs, the agent found the right restaurant, showing, or flight but could not finish the task. The bottleneck appeared to be its execution environment: the computer and browser available to it, its integrations, and the software that keeps track of progress across steps. Authentication challenges, expiring checkout sessions, and payment failures tested that system. We could see where work stalled, though not whether a given failure came from the model, its tools, or the outside service.
The ideal assistant would remember relevant preferences, work quickly in a persistent computer, use direct integrations where available, and handle sensitive steps through specific user approvals. Muse was the only agent in our runs to visibly offer an OpenTable connection, suggesting how direct integrations could simplify reservations. But no integration can cover every site. To work out of the box, an assistant must navigate unfamiliar services, ask for specific help when required, resume from the same state, and verify the outcome.
About Delphi
Delphi builds evaluations and human-data workflows for non-verifiable domains. Text-to-World Bench starts with observable outcomes while examining the judgment around them: understanding intent, applying context, choosing well, and knowing when evidence is strong enough to say the work is done.
Back to the results