StepFun unveiled the StepX Neo in Shanghai on 13 July 2026, presenting it as a phone built around an AI agent rather than a conventional smartphone with an assistant added later. The important question is not whether the agent can book a ride, order food or move between apps. It is whether the owner can see, limit, confirm and reverse those actions.
That is the practical test for any AI phone. An agent that can act across communication, payments, files and travel services becomes useful only when authority is explicit. A clever demo proves capability; a trustworthy product must also prove boundaries.
What StepX Neo appears to change
Current launch coverage describes Step AOS as an operating system designed around StepFun's Amoo agent. Instead of opening several apps and completing each stage manually, the user states an outcome and the agent coordinates supported services. Reports also describe a local Step Edge model and a system layer that exposes phone capabilities to the agent.
Some details remain unconfirmed. StepFun had not published a complete global specification, price or international availability schedule at the time of writing. Claims such as “world's first” are therefore best treated as the company's positioning, not an independently settled category record.
The more consequential change is architectural. A chatbot usually answers inside one interface. An agentic phone may need to read screen context, inspect files, select contacts, call services and prepare transactions. Each additional capability expands both usefulness and the cost of a mistake.
The agentic-phone permission matrix
Before trusting any agentic phone, buyers should be able to answer the following questions from the interface itself—not from a marketing presentation.
| Action | Minimum permission | Confirmation that should appear | Recovery control |
|---|---|---|---|
| Draft a message | Selected conversation or contact | Preview recipient and full text | Edit or discard before sending |
| Send a message | One-time send authority | Final recipient, channel and content | Visible sent record; revoke future access |
| Book transport | Location and selected travel service | Route, fare, vehicle and payment method | Cancel within the provider's terms |
| Buy an item | Product, delivery and payment access | Total cost, seller and delivery address | Order record and cancellation route |
| Read a document | Specific file or folder access | Name the files being opened | End session and revoke file access |
| Change a setting | Exact device-control permission | Show old and proposed state | One-tap undo where technically possible |
| Remember a preference | Defined memory category | Show the fact being saved and its purpose | View, edit and delete the memory |
This matrix separates preparation from execution. A phone can safely do more work in the preparation layer—finding options, filling fields, drafting messages—than in the commitment layer. Sending, purchasing, deleting, sharing or changing access should require a clearer confirmation.
Five tests matter more than the launch demo
1. Is authority scoped to the task?
“Use my travel apps” is too broad. “Check train options from Shanghai to Hangzhou tomorrow morning without booking” is a bounded instruction. The operating system should translate that boundary into permissions the user can inspect.
Good design distinguishes one-time access, access for the current workflow and persistent access. If the system only offers a permanent on/off switch, convenience may come at the cost of unnecessary authority.
2. Can the user see the execution plan?
An agent does not need to expose internal reasoning, but it should show the operational plan: which services it will contact, what information it will use and where a confirmation will occur. A buyer should not discover after the fact that a task shared a contact list or selected a different payment account.
3. Are high-impact actions separated?
Booking a car and sending an itinerary are two different commitments. Combining them into one invisible chain makes recovery harder. A trustworthy agent pauses at meaningful boundaries—payment, external communication, deletion, account changes and disclosure of private information.
4. Is there an audit trail?
The owner needs a plain record of what the agent did, when it acted, which service responded and what was confirmed. This is especially important when a task crosses app boundaries. A single chat transcript is not enough if it omits the actual external action.
5. Can memory and sessions expire?
Personalisation becomes risk when context remains indefinitely. The phone should make memories visible, attach them to a purpose and support expiry. The same applies to signed-in sessions and delegated app permissions. Our separate AI memory expiry control matrix explains how to classify what should be retained, reviewed or deleted.
What buyers can verify before purchase
Do not judge an AI phone only by the number of apps mentioned in a launch. Ask for a live demonstration of one complete task and deliberately interrupt it.
Change the destination halfway through a travel request.
Decline the payment confirmation.
Revoke access to one app and repeat the request.
Ask the phone to show what it remembered.
Review the action history and delete the session.
Confirm what still works without a network connection.
The interruption test is revealing. A robust system should stop cleanly, explain which steps were completed and avoid retrying a commitment without fresh approval.
The VERTU perspective: capability needs a human boundary
VERTU's current Hermes Agent positioning uses the same essential principle: actions across supported apps and services should occur with user authorisation, while significant actions require confirmation. Availability depends on device, configuration, software version and supported services. The point is not to claim that one architecture has solved agentic computing. It is to make user control part of the product definition.
The same distinction applies to human concierge service. An AI agent can compare options and prepare a request; a human can handle ambiguity, exceptions and judgement when automation reaches its limit. An intelligent phone should make that hand-off visible rather than pretending every situation is deterministic.
What StepX Neo needs to prove next
StepX Neo is significant because it turns the AI phone from a feature contest into an operating-system question. Yet the launch is the beginning of the evidence, not the conclusion. Independent reviewers still need to test error handling, app compatibility, regional availability, permission revocation, offline behaviour and the reliability of transaction confirmations.
The buyer's rule is simple: the more an AI phone can do, the more precisely it must show what it is allowed to do. Cross-app autonomy without scoped authority is not premium intelligence. It is an uncontrolled dependency.
A 30-minute showroom test
Buyers do not need a laboratory to expose weak control design. Use a harmless test account and ask the phone to plan a restaurant journey involving maps, messaging and transport. Before any external action, note whether the system names the app, recipient, location and payment state. Change one constraint after the plan is prepared, then deny the final confirmation.
Next, inspect the permission page. A strong interface should show which access was temporary, which remains active and how to revoke it. End the session, restart the phone and ask what it remembers. Finally, request an export or readable history of completed actions. If the salesperson cannot explain where these controls live, the autonomy claim is ahead of the ownership experience.
This test does not establish security, but it reveals the product's philosophy. An agent-first operating system should make authority easier to understand than a collection of separate app permissions—not hide more power behind one conversational surface.




