coming soon

AI Has a UI Problem

Prompting made AI feel like magic, but it also flattened every task into the same shape: make a wish, wait, hope. As AI assistants take on real work, that's starting to cost us. What we'd like to see are interfaces that adapt to the task at hand and reach for language only when language is the best fit for it.

The Story so far

It has only been a few years since chat-based AI tools like ChatGPT, Claude, and Gemini found their way into our lives, and they have since become an integral part of how people engage with technology, shifting every major tech company's product roadmap to accommodate them.

The reason is easy to understand. Where people once had to tediously look things up themselves, comparing and judging different sources to piece together the bigger picture, an AI assistant answers the question quickly and easily, based on whatever its training data suggests.

One of the bigger problems we're seeing has less to do with what these tools can do, and more with how we engage with them. As OpenAI, Anthropic, and Google enrich the feature sets and capabilities for their intelligent helpers, we're not sure that natural language alone is the best way to express user intent efficiently.

The Almighty Textbox

The initial use cases for Large Language Models revolved mostly around answering questions, summarizing content, and generating text. Through 2024 and 2025, however, the narrative shifted, anchoring itself around coding assistance (vibe coding) and media creation, until images, video, and even audio could be generated from a simple prompt.

Prompting in natural language is a great way to state your initial intent in broad strokes, but it gets progressively more complicated and tedious the more precise the request becomes. Generating an image this way is fun; generating the image you actually had in mind can be daunting.

Language is grossly imprecise, and it leads to misunderstandings all the time. The graphical interfaces we're accustomed to are a direct reflection of our efforts to work around that, reducing the number of inputs needed to make a request as concisely and efficiently as possible. Indicating a location on a map is less likely to be misunderstood than explaining it, and dragging a slider to adjust the brightness of an image is more precise than asking three times to "make it a little brighter."

So there is a place for intelligent tools and features in pretty much every workflow imaginable, but we're skeptical that natural language and conversational UI is the best way to introduce those capabilities. Different scenarios call for different controls and flows to be efficient at all. That doesn't mean there's no place for language. It means language shouldn't be the only interaction paradigm we rely on when expressing intent.

Though comprehensive, an AI assistant's answer to a simple request is often impractical, and sometimes misses the point entirely.

The AI Superapp

AI tools have long moved on from the best lasagna recipes, summarizing longform articles into bite-sized pieces, and cheating on homework. Most models now support media creation, are an integral part of the development process at many companies, and handle all kinds of chores through third-party MCP server integrations.

All of that has the likes of OpenAI and Anthropic fantasizing about the world's first true superapp: an app that lets you simply do anything upon request, intelligently.

The trouble is the way we've grown accustomed to interacting with these tools: via language. It feels natural to us, but it's not the most efficient way of querying.

Take something as simple as a food order. The user has to settle on a cuisine, a price point, a search radius, an estimated delivery time, menu items, a delivery address, a payment method, and the list goes on. Or think of hailing a taxi: your pickup address, the time, your destination, the type of car you're after, and plenty more besides. Uber and Grab made us forget how tedious all of that querying can be, and how much of a hardship ordering a pizza or a taxi by telephone was right up until the 2000s.

Most of the use-case-specific UX/UI patterns we rely on have gone through decades of refinement, all of it aimed at reaching a given outcome with the fewest steps and the least input possible. Natural language is a worthwhile addition to that set of paradigms. Discarding the highly specialized GUIs we've refined over the years is not.

Imagine an interface that reads how precise your prompt is and assembles the controls to match, pulling live elements from whatever 3rd party service the task calls for. The vaguer the ask, the more it shows.

Adaptive User Interfaces

Don't misunderstand us. We'd love to see a company actually build an intelligent assistant that can help with drafting emails, editing photos, and ordering pizza all equally efficiently. But the interfaces it presents us with need to adapt to the use case, building on top of the already excellent foundation the current app ecosystem provides.

Want to book a flight? Imagine ChatGPT presenting you with the elements relevant to booking one. Want a new pair of trainers? Picture Claude laying out the selection criteria it needs to understand what you're after. Want to check on Google stock, and maybe short it? Envision Gemini generating an interactive price chart alongside the controls to place the trade through your brokerage.

Prompting is great, but it's become a skill in its own right, since not all prompts are created equal. We see a lot of potential in a hybrid approach that pairs the best of both worlds: the natural feel of a Conversational User Interface (CUI) with the fine-grained control of a Graphical User Interface (GUI). These Adaptive UIs have surely been experimented with before, but rarely with a model capable of deciding which controls to put in front of you. That's the shift: an assistant that works out which inputs a vague, incomplete prompt is missing, and asks for them in the form best suited to the task.

The first bits and pieces of this are already surfacing across a number of AI assistants. We've noticed Claude, for example, offering us multiple choices more and more often over the last year, and we'd like to see much more of it: a single tap sometimes saving us a dozen keystrokes.

Closing Thoughts

We're obviously not the first to express these thoughts. Quite a few people have made similar appeals about the UI paradigms we've settled on for interacting with AI, and the big labs seem aware of it too. OpenAI is working on partial integrations of services like Uber into ChatGPT. Google keeps improving how accessible its own services are from within Gemini. And Claude has been offering multiple-choice answers more frequently over the last six months.

We'd like to see more of it. More than that, we'd like the integrations to be far more consequential and comprehensive, so these assistants can perform to their full potential: a single tap sometimes saving us a dozen keystrokes.

Get in touch

Though headquartered in Hong Kong, our team and partner network extend internationally, enabling us to work with clients across borders and time zones.

Get in touch to discuss your next project. We are open to discussing scope, requirements, and next steps.