Skip to article
Delphi← Benchmark results

September 2026

Text-to-World Bench: Can agents turn a text into a real-world result?

What this benchmark measures

We treat each task as an end-to-end request. Alongside the final outcome, we examine how assistants use context, ask for clarification, handle obstacles, communicate progress, and support their completion claims. We also record response and completion latency and assess the conversational experience, including tone and cadence. The central question is not whether an assistant can understand the request, but whether it can carry that request through to a verifiable result.

Across 7 agents and 11 tasks, we tested whether a short message could become a reservation, delivered meal, purchased ticket, or flight booking ready for checkout. Existing efforts such as Assistant Benchmark cover a broad range of capabilities; we deliberately focus on consequential real-world actions that people genuinely want to delegate and that remain difficult for today's agents to complete reliably.

To make comparisons consistent, we reused the same opening prompts, presented tasks in the same order, and reviewed each outcome as a pass or failure.

Task design and evaluation protocol

  • Real workflows. Tasks came from the places where we would actually delegate work: reservations, food delivery, movie tickets, and flights. We excluded low-consequence tasks that agents already handle reliably.
  • Shared conversations. Each agent remained in one persistent chat. Later requests could therefore test whether earlier constraints and preferences changed the next decision.
  • Outcome-first review.The headline score is binary against each task's stated checkpoint. Partial progress and timeouts do not receive fractional credit, even when the intermediate work is useful.

What the runs revealed

Useful assistance can stop short of completion

Agents frequently found the right option, selected a plausible seat or itinerary, and reached checkout without completing the final action. Muse's live browser handoff and Pally's dedicated flight interface made those moments easier for the user, but both could still shift consequential work back to the person who delegated it. Helpful preparation and autonomous completion should therefore be reported separately.

Context matters when it changes a decision

The strongest memory behavior was not repeating a stored fact. It was using prior context at the right decision point. Muse reused a party size and seating preference across reservations, then warned that a San Francisco movie conflicted with an earlier San Jose dinner. Persistent memory needs relevance control, not just retention.

Execution barriers interrupt otherwise promising workflows

Several unsuccessful attempts reached relevant options or checkout before encountering authentication, payment, browser-state, or website obstacles. These observations identify visible points of failure, but they do not isolate the contribution of the underlying model, agent implementation, external service, or account configuration.

Instinct finished the most tasks, passing seven of eleven. It responded most like a human assistant would: concise, action-oriented, and clear about the next decision.

Muse (through WhatsApp) produced the strongest qualitative experience in our assessment, with preference reuse, conflict detection, and clear handoffs. Its successful runs also had the shortest median completion time. It passed five tasks; many of its other runs ended in handoffs before independent completion.

Poke Ultra also passed five tasks. We tested its Ultra plan at $199 per month. Its human operators sometimes completed the request, as with reservations; other handoffs ended with progress updates and no result.

Pally offered a dedicated flight interface that made its proposed itinerary easier to review than a long thread of messages.

Catch completed three restaurant reservations, but said it could not place food orders or buy movie tickets.

Folk built the requested food cart and reached checkout, but could not complete payment.

Arlo quoted a flight for two instead of one and shifted reservation attempts to phone calls without completing the bookings.

Annotated cases

Case 01Reservation confirmed

A reservation becomes context for the next task

Muse
“you've got dinner at must be thai in san jose at 5pm that same wednesday, so a ~7:15pm showing in sf would clash pretty hard.”
Request
Book dinner at Must Be Thai and Bar in San Jose, then a Spider-Man ticket near 7:15 p.m. in San Francisco that Wednesday.
What it knew
Earlier bookings established two diners and standard seating; the confirmed 5:00 p.m. dinner remained in chat.
Result
Muse confirmed a 5:00 p.m. table for two with standard indoor seating and a confirmation number. It flagged the movie conflict; we chose to proceed. Muse said AMC had declined its virtual card on earlier attempts, and no ticket was bought.
Why it matters
Muse reused dining preferences, then caught a conflict with the next request.
Case 02Incomplete

A strong handoff still leaves the task unfinished

Muse
“The payment page is ready ... take over through the Muse app or muse.ai, add your card, and hit purchase.”
Request
Buy one good seat for Spider-Man: Brand New Day near 7:15 p.m. at the theater closest to Rincon Hill.
What it knew
Muse knew about the 5:00 p.m. San Jose dinner and reported three earlier AMC virtual-card declines.
Result
Muse selected a center seat and reached payment. Citing earlier card declines, it asked us to finish checkout in its browser. No ticket was issued.
Why it matters
Muse found the seat but left payment to us, so the purchase remained unfinished.
Case 03Delivered

A food order reaches the door

Instinct
“Order’s in — Cantoo, $27.05 on your card with the $2 tip. Arriving 9:16–9:31.”
Request
Order Cantoo for delivery with a rice substitution, extra sauce, and salted pumpkin.
What it knew
Instinct flagged unavailable sauce. We supplied the address and sign-in code, then approved the revised cart and tip.
Result
We verified delivery and that the food matched the approved cart. Instinct quoted $27.05 with a $2 tip and a 9:16–9:31 p.m. arrival window.
Why it matters
Instinct resolved an unavailable item, got approval, and completed delivery.
Case 04Reservation confirmed

Poke's human handoff secures a table

Poke Ultra
“one sec, handing this to our human agent to get your table at grand lake kitchen for monday sept 21”
Request
Book Grand Lake Kitchen in Noe Valley for two on September 21 between 5 and 9 p.m.
What it knew
We supplied seating and contact details; Poke handed the booking to a human operator.
Result
Poke reported the booking; we verified it.
Why it matters
Poke Ultra's human operator secured the table, a successful assisted handoff.
Case 05Timed out

The human handoff stalls before a ticket purchase

Poke Ultra
“a human's on it, grabbing you a good seat at kabuki for the 7pm show”
Request
Buy one good seat for Spider-Man: Brand New Day near 7:15 p.m. on September 23.
What it knew
Poke found a 7:00 p.m. AMC Kabuki showing and routed the purchase to a human operator.
Result
We approved 7:00 p.m. Poke said an operator was buying the ticket, then sent another progress update. No ticket arrived within 20 minutes, and it requested no further action from us.
Why it matters
Even with a human operator involved, no ticket arrived. The record does not show whether the breakdown was in routing the task or completing the purchase.

The path toward better assistants

In many runs, the agent found the right restaurant, showing, or flight but could not finish the task. The bottleneck appeared to be its execution environment: the computer and browser available to it, its integrations, and the software that keeps track of progress across steps. Authentication challenges, expiring checkout sessions, and payment failures tested that system. We could see where work stalled, though not whether a given failure came from the model, its tools, or the outside service.

The ideal assistant would remember relevant preferences, work quickly in a persistent computer, use direct integrations where available, and handle sensitive steps through specific user approvals. Muse was the only agent in our runs to visibly offer an OpenTable connection, suggesting how direct integrations could simplify reservations. But no integration can cover every site. To work out of the box, an assistant must navigate unfamiliar services, ask for specific help when required, resume from the same state, and verify the outcome.

About Delphi

Delphi builds evaluations and human-data workflows for non-verifiable domains. Text-to-World Bench starts with observable outcomes while examining the judgment around them: understanding intent, applying context, choosing well, and knowing when evidence is strong enough to say the work is done.

Back to the results