The making of the Yelp Assistant UI
-
Daniel Andrade Groppe, Services Tech Lead
- Sep 29, 2026
As simple as it may look, a chat application is made up of thousands of small decisions that we make along the way, all with the goal to build a delightful experience for users. But each decision comes with a cost, often paid in time. If you ask people to draw a chat application, there is a high chance that every drawing will look very similar to this.

Yelp Assistant is the first large-scale consumer-facing use of large language models at Yelp. We built it to connect users with local professionals with a simple proposition: a user has a problem, opens the Yelp Assistant, and writes “I have a leaky faucet”. Yelp Assistant asks questions until it is certain that it has all the information that a pro might need. When the user is ready, Yelp Assistant sends a project to multiple businesses.
Yelp Assistant has since expanded to answer questions about a business or surface ways to connect with them such as creating reservations, booking appointments, or ordering food.
In building Yelp Assistant, we faced uncertainty from different angles: first, we were learning to use it as part of a product. And second, the technology was evolving fast. We needed to be able to make decisions and feel comfortable reversing them when they no longer made sense. If decisions are paid with time, how can we mitigate the cost of making and changing decisions?
To build native or server driven?
The first version of Yelp Assistant was launched early in 2024 with one goal: guide users to submit a high-quality project using an LLM. This version was built using SwiftUI and was available on iOS. Soon after that, we began planning Yelp Assistant’s expansion to other platforms.
Unlike building for web platforms where users are one refresh away from using the latest client code, mobile applications have a unique constraint: every change requires a new build to be deployed to and reviewed by the platform’s app store, and it still requires that users download the latest version once it is available. This means that any change, however small, can take several days to reach the users in the best of cases. In the worst cases, a user may never update the app, and they will remain in a time-locked experience that is either broken due to a bug or not the experience we intended for them.

We litigated, deliberated, decided, and backtracked multiple times on an alternative: to build Yelp Assistant using a server-driven UI (SDUI). In a server-driven UI application, the backend tells the client what to show to the user and what to do when the user interacts with the UI. If we ever make a change, all we have to do is update the backend and all clients receive the latest configuration within minutes of the change being deployed.
Yelp has a server-driven UI framework called CHAOS, short for Content Hosting Architecture with Optimization Strategies. And it is multi-platform, available for web, iOS, and Android. This allows us to create a tailored experience that looks and behaves in the same way, no matter how the user got to it.
However, as CHAOS was in its infancy, it came with two disadvantages. First, only a handful of people were familiar enough to build something with it. And second, we’d build Yelp Assistant and evolve CHAOS while doing so. But we were making an investment now to make our lives easier in the future. We went all in on CHAOS.
Actually, a hybrid is what we needed
CHAOS shines at rendering dynamic content such as a conversation that might happen in Yelp Assistant, but it doesn’t provide everything that we needed. For example, CHAOS doesn’t provide affordances to send requests to the backend and support for text input fields did not exist at the time. Building support for network requests or input fields required more time than we wanted to invest. Our goal was to build something, ship, and start learning.
We chose a hybrid solution where some parts of the UI are native and others are server-driven. The illustration below shows an early design of Yelp Assistant. There are three parts to it: the top bar, which contains a close button, the chat content area that contains the messages sent and received, and the input bar, where the user types messages. The top and input bars are native driven and the content section is server-driven.

Making native UI and SDUI work together
Whether we use a native UI or SDUI, the UI is self-contained and the UI framework handles all aspects of the user experience, from what the user sees to what they do. However, to make a hybrid approach work, we blur the lines between the native UI and the server-driven UI.
For example, sending a message requires interacting with the native UI, and displaying it requires access to the CHAOS context.
CHAOS is built to be extendable and we leverage this extensibility to make Yelp Assistant possible. We build CHAOS components to be used specifically in Yelp Assistant. A component is a visual element that conveys information to the user such as text, pictures, or buttons. When appropriate, a component uses one or more CHAOS actions to trigger pre-defined behavior when the user interacts with it.
At a high level, the chosen architecture separates the view and the data. The client fetches the view and uses it as a template, and as the conversation progresses, the data hydrates the view. When a message is sent, the bot’s reply includes messages, quick replies, and client actions.
In the following sections we’ll explore actions, components, message events, quick replies, and client actions.

Initializing the chat
We separated the view from the data to allow incrementally adding content to the chat as the conversation progresses without fetching the view on each turn. For this to work, we must be sure that the view is ready before the first message is sent. As soon as Yelp Assistant starts, the mobile app fetches the CHAOS view configuration which provides the blueprint for the chat. The view configuration also includes actions that may be triggered when certain events occur and operations that may be required for the proper functioning of the application.

Once the view is fetched, the application sets up the storage required to keep the state of the conversation. In CHAOS, these are called datasets and each entry in the dataset represents a MessageEvent.
A MessageEvent tells CHAOS the template and contents to use. The CHAOS framework observes the datasets to keep the UI up to date as the user sends or receives messages.
data class MessageEvent(
val id: String,
val type: String,
val text: String? = null,
val url: String? = null
)
The chat loop
Yelp Assistant is reactive and turn-based. It only responds to a message and doesn’t proactively send messages to the user. After a message is sent, the user must wait for the bot to reply or fail before sending a new message.
When the user types a message and presses the send button several things happen:
- the input box is disabled
- quick replies are cleared
- the client sends the message to the backend
- the client optimistically inserts the message and a typing indicator into the dataset
In the happy case, the bot sends a response and we show it to the user. When this happens, the client:
- removes the typing indicator
- inserts the response into the dataset
- if available, adds the quick replies
- enables the input box

We handle two failure modes categorized as recoverable or non-recoverable. In both cases, the client removes the typing indicator and inserts an error message.
In the recoverable cases, the client retries for a number of times before asking the user to send the message once again. In the non-recoverable cases, the client asks the user to try again at a later time and disables the chat for further input, effectively ending the session.
Rendering the conversation
Every message in the conversation is a MessageEvent and to the chat application, it is mostly opaque. As the conversation moves, new MessageEvents are added to the dataset and CHAOS uses component templates to render each one. The specifics of how CHAOS achieves this will be covered in a future blog post.
We started with a few event types to support the experience we were building. Each type has a specific purpose. Some types display information while others perform an action. Some are persistent and are visible from the time they enter the conversation until the end while others are ephemeral and are removed once no longer needed.
| Type | Ephemeral or Persistent? | Interactive? | Notes |
|---|---|---|---|
| user_message | Persistent | No | Shows the text sent by the user. |
| typing_indicator | Ephemeral | No | Shown from the time the user sends a message to the time the response is received. |
| bot_message | Persistent | No | Shows the bot response. |
| link_message | Persistent | Yes | Shows a link to navigate to a different screen. Generally signals that a chat has ended. |
| quick_replies | Ephemeral | Yes | Shows a list of possible replies that the user may use. |
Anatomy of the conversation
Every conversation in Yelp Assistant is a collection of MessageEvents. Some are inserted by the client whenever a message is sent, some are inserted when there is a response from the backend. Only the quick replies carousel and the typing indicator are removed from the dataset.
When the user sends a message it goes to two different places: to the local dataset and to the backend. We insert the user message to the dataset to show it instantly. We call this an optimistic update.
MessageEvent(
id = "user-message-01",
type = "user_message",
text = "I want to fix a leaky faucet"
)
And the user sees a chat bubble with their message:

The client also inserts a typing indicator into the dataset and will remove it once the backend sends a response. The client is fully responsible for inserting and removing the typing indicator at the right time.
MessageEvent(
id = "typing-indicator",
type = "typing_indicator",
)

As soon as the bot message is received, the client inserts it into the dataset.
MessageEvent(
id = "bot_message-01",
type = "bot_message",
text = "Can you tell me more about the issue? Is it a constant drip or does it leak when you turn the faucet on?"
)

On occasion, we also show a carousel with quick replies. These are presented as bubbles underneath the last message received and the text is relevant to the question being asked by the bot. These suggestions save time as the user need only tap a bubble to send it as their next message.
Quick replies have a different schema than other messages and are backed by a different dataset but follow the same principle as messages.
listOf<QuickReply>(
QuickReply(text = "It's a constant leak"),
QuickReply(text = "Only when you open the hot water"),
)

This iteration of Yelp Assistant has one purpose: to guide users through creating a high-quality Request a Quote project. When that is achieved, the chat ends. But simply ending the chat isn’t useful. Instead, we show a link message to direct the user to their newly created project.
MessageEvent(
id = "link_message-01",
type = "link_message",
text = "See your matches",
url = "yelp:///project/..."
)

Acting on the user’s actions
CHAOS lets us configure the desired behavior when a user interacts with the UI by declaring CHAOS actions in event hooks such as onView or onClick. CHAOS actions are signals to the framework to trigger code in the client application and are either built into CHAOS or are created specifically for Yelp Assistant. An action created for Yelp Assistant is still a CHAOS action, however its scope is limited to the chat screen.
There are two interactive components, quick_replies and link_message, and each has specific behavior that is triggered when the user taps.
| Component | Action | Provided by CHAOS or YA? | Notes |
|---|---|---|---|
| quick_replies | quick_reply | YA | Triggered when the user taps a quick reply bubble. This action routes the message through the message-sending logic. |
| link_message | open_url | CHAOS | Triggered when the user taps the link. Uses the web or mobile application context to navigate to the link’s destination. |
The backend as an actor
On occasion it is necessary for the backend to trigger certain actions in the client. For this, we use client actions. Whereas a CHAOS action is generally triggered by a user’s action such as interacting with a visual component, a client action is executed by the client on the backend’s instruction.
The first client action, DisableChat, tells the client to disable all forms of input and prevent further messages from being sent. This can happen for various reasons but the most common is simply that the purpose of the chat, which is to submit a project, is fulfilled. In less common cases, the chat may be disabled due to the user hitting rate limits, or errors.
The mechanism is flexible to support new actions as we require them.
End of the chat
The chat ends when its purpose is fulfilled. When this happens, the user can no longer send messages. However, there are other situations when ending the chat is appropriate such as in the case of errors or outages.
To achieve this, the backend includes a DisableChat client action in its last response to the user to remove the text input bar from the UI.
Response(
messageEvents=listOf(
MessageEvent(
id="bot_message-02",
type="bot_message",
text="Success! 🎉 I’ve matched you with a few local pros.",
),
MessageEvent(
id = "link_message-02",
type = "link_message",
text = "See your matches",
url = "yelp:///project/..."
)
),
clientActions=listOf(
DisableChat
)
)
With the input bar removed, the conversation is now complete.
Putting it all together
Yelp Assistant is the product of the collaboration of several teams across platforms and domains with the goal to build a product that looks and feels uniquely new, yet familiar. The decisions that we’ve made along the way allow us to iterate and improve Yelp Assistant and this is largely facilitated by the decision to build a hybrid native-server-driven UI.
Separating view from data enables us to think about what we’re building in a different way. Whether the client renders a user message, a typing indicator or something completely new does not matter, as CHAOS abstracts it away. This makes it possible to add new functionality such as image recognition or support for logged-out users as each part of the entire system is independent of the others.
This has come with challenges too. First, people require additional time to get familiar with Yelp Assistant, its architecture, and CHAOS, before they make effective contributions. As Yelp Assistant is a multi-platform product, this also requires careful thought and planning. It is reasonable to expect everything to behave consistently across platforms, but sometimes we discover that what works well in one platform is buggy in another.
This was the first iteration of Yelp Assistant and soon after it shipped, we realized that other teams in Yelp were thinking about building similar products. In one case, a team cloned and repurposed Android’s Yelp Assistant to bootstrap the pilot for Biz Ask Anything. We didn’t realize it at the time, but we laid the groundwork for what would become the next generation of Yelp Assistant.
Since then, Yelp Assistant has expanded from connecting users to local services pros to leveraging Yelp’s wealth of information to help users get things done. Across the company, teams can integrate their features with it to create new ways to connect with great local businesses.
Acknowledgements
Yelp Assistant is the product of years of collaboration across many disciplines. The list of people who have shaped Yelp Assistant is far too large to be able to thoroughly thank every person individually. But as this is a post about a narrow aspect of Yelp Assistant, I extend my thanks to the team behind CHAOS and to the people in the Services group who worked tenaciously to build the Yelp Assistant UI, especially Mario Stallone Villavarayan, Marjorie Figueroa, Nicholas Squire, Shaban Solakov, and Sunny Lin.
Become a Software Engineer at Yelp
Want to help us make even better tools for our full stack engineers?
View Job