← Article directory

Tech enthusiasts never learn. Chatbots 2015–2018 as a manual nobody read

21. 4. 2026
Tech enthusiasts never learn. Chatbots 2015–2018 as a manual nobody read
Image from the original article on Médium.cz

The article reconstructs the 2015–2018 wave of conversational chatbots — from Facebook M and the Messenger platform through the Microsoft Bot Framework to concierge startups like Magic, Operator, and x.ai — and shows that it collapsed not on immature AI but on an economic equation: the cost of the human layer plus operations exceeded what users were willing to pay for the service, compounded by duplicated data and dependence on someone else's platform. The author argues that generative AI in 2026 (OpenAI, Anthropic) is repeating that same equation in a better suit, where the cost per resolved query still exceeds the revenue from the customer — only the patience of capital has changed, not the economics themselves. The only survivors, therefore, were models with a "captive" customer base and built-in monetization, like the banking assistants Erica and Eno.

The equation that buried the wave

A query handled by a trainer in an American call center cost somewhere around fifty cents to a dollar circa 2016. Booking a restaurant saves the user, say, three minutes. If the service costs the user nothing, they'll take it. If they have to pay a dollar for every use, you lose the subscriber. And between what the service cost the platform and what the user would pay for it, there was no gap through which a profit could flow.

This single equation contains the entire thesis of the article. The wave that between 2015 and 2018 convinced Silicon Valley that chatbots in Messenger would replace apps did not fall apart on immature AI. It fell apart on this arithmetic. And it took technically enthusiastic developers and their investors three years to take note of it.

The wave built its existence on one quiet assumption: that the human layer temporarily filling the holes in automation is a shrinking cost. By that picture, each round of training was supposed to reduce the share of queries a human had to handle. In the limit, the human layer was meant to vanish, leaving only software that scales at zero marginal cost. It did not decline.

The second problem was duplication. When ten users in a row work a bot over with requests of the type "book me a table at an Italian restaurant in the Mission," each arriving with a slightly different phrasing, the system learns only slightly more from them than from the first one. The long tail of unique requests — from ordering concert tickets through arranging a wedding to waiting on hold with the cable company — stayed long. Each additional trainer therefore didn't lower the unit cost, it just grew it by one more salary.

The third problem was the channel. A chatbot in Messenger or Kik was not itself the product; it was a thin layer over a platform that claimed the rules of the game for itself. Facebook could overnight restrict when a bot was allowed to message a user (in 2018 it limited this to 24 hours after the user's last active contact) and restrict which ads a bot was allowed to send. The owner of the channel could set the take-rate arbitrarily high, because no alternative distribution existed.

These three problems — the linear cost of humans, duplicate data, platform risk — were not a problem a transformer would fix. They were problems of an economic structure the category had chosen for itself. And the wave ran into them before anyone wrote the first BERT.

Zuckerberg's joke that summed up the wave

"Take 1-800-Flowers. I love this. I find it ironic — now, to order something from 1-800-Flowers, you'll never have to call 1-800-Flowers again." With these words, on April 12, 2016, at the F8 conference at Fort Mason, Mark Zuckerberg introduced an ambition that, over the following twenty-two months, fell apart so cleanly that little was left of it beyond that very quote. "I don't think I've ever met anyone who likes calling a business," he added that day, "and nobody wants to install a new app for every service." With that he uttered a sentence that became shorthand for an era of platform thinking: that messaging — a conversation with a bot in Messenger — would replace the app.

In the years 2015 to 2018, Silicon Valley went through a wave of conviction that this would indeed be so. It was set in motion by Facebook's acquisition of Wit.ai in January 2015, carried to its public zenith by Zuckerberg's keynote at F8, and buried by the quiet shutdown of the assistant M in January 2018. Between these three dates, several hundred thousand chatbots came into being, several dozen funds convinced that "messaging is the new runtime," and several dozen startups built on a thesis that none of them managed to make economically viable.

Why the view from 2026

The vantage point of 2026 makes it possible to reconstruct that wave as a closed episode. Conversational AI in the form we use today through ChatGPT, Claude, or Gemini has only a surface resemblance to the promises of 2016: the same user metaphor, an entirely different architectural backbone — and, something usually overlooked in the commentary, an entirely different business model. The products of 2026 are sold to users through direct subscriptions or to corporations through APIs; nobody monetizes them through chat bubbles in someone else's messaging platform. It is precisely that difference that makes hindsight meaningful — and at the same time reveals how exactly some of the mistakes repeat themselves in new clothes.

Wit.ai and M: the trap of the human layer

On January 5, 2015, Facebook announced the acquisition of the California startup Wit.ai, founded in 2013 by three French engineers. The sum was not disclosed; the startup had a $3 million seed round from Andreessen Horowitz, New Enterprise Associates, and SV Angel from October 2014 behind it and a community of six thousand developers who used its interface for voice recognition in apps, wearable electronics, and the smart home. The founders — Alex Lebrun, Laurent Landowski, and Willy Blandin — moved to Menlo Park. Before Wit.ai, Lebrun had built and, in January 2013, sold to Nuance the company VirtuOz, described as "Siri for the enterprise"; Wit.ai was his second attempt at the same thing, this time as a public interface in the style of Twilio or Stripe.

Facebook bought Wit.ai with a clear goal. The negotiations were entrusted to David Marcus, the former president of PayPal, whom Zuckerberg had poached in June 2014 to head Messenger. Marcus needed a natural-language processing engine that could power a more ambitious product than simple choices in chat. After seven months of development, on August 26, 2015, Facebook opened the beta of the personal assistant M — with the Wit.ai team as its language backbone.

The architecture that was supposed to phase out humans

On Wednesday, August 26, 2015, Jessi Hempel published a cover story in Wired about the launch of M. In his Facebook post, Marcus explained that M is "a personal digital assistant inside of Messenger that completes tasks and finds information," powered by artificial intelligence "trained and supervised by people." The key difference from Siri, Google Now, and Cortana was supposed to be that M can actually complete tasks: order flowers, book a restaurant, arrange a flight reservation, wait on hold with Comcast. The beta ran on "a few hundred users" in San Francisco and the surrounding area; Messenger had roughly 700 million monthly active users at the time.

An incoming query passed through the Wit.ai engine, which proposed a response or action; that went into a queue to an "M trainer" — a contractor with call-center experience, sitting at first directly with the engineers in Menlo Park — who approved the response, rewrote it, or replaced it entirely; the system then gathered training data from this teacher–student composition for the next rounds. In an interview with BuzzFeed, Marcus described the philosophy in these words: "Today's flowers are tomorrow's car rental." Automation was meant to spread gradually across verticals — with each additional category the human ratio was supposed to fall.

It did not fall. Fast Company later reconstructed that, by the time of the closed beta, the service had settled at roughly 2,500 users in the Bay Area; in April 2017, MIT Technology Review estimated it at 10,000, predominantly in California. Tom Simonite then wrote for MIT a sentence that became the epitaph of the entire project: "M is so smart because it cheats." When the algorithm didn't understand, it didn't return an error — a human imperceptibly took its place. According to later reporting, this amounted, over the long run, to roughly 70 percent of responses handled by humans — a ratio Facebook never officially confirmed or denied. As early as November 2015, Marcus, asked by BuzzFeed reporters whether M wasn't just an electronic Mechanical Turk, answered: "Vehemently [I deny it]. Vehemently."

The admission that bypassed the press

Only one person publicly described the economics of that layer. Mike Schroepfer, Facebook's chief technology officer, conceded in a November 2015 interview with Recode: "We can't afford to hire operators for the whole world to act as their personal assistants." It was a sentence that, under any normal assessment, should have ended the project. If it was true — and it was — M's original product thesis could never arrive at any threshold of profitability.

The product nonetheless continued for another year and a half. One reason was that Messenger needed a platform story for F8 2016 — and internally, at that moment, no one had anything better to put on the stage. The second reason was that, inside Facebook, M served as a research probe for Wit.ai — every query a trainer corrected was a training data pair. The product wasn't profitable in itself, but as infrastructure for generating data it was cheaper than commercially purchased datasets. In that sense, M was a success. Except Facebook could only talk about that internally.

The quiet retreat

On April 6, 2017, Landowski and product manager Kemal El Moujahid published a post on the Facebook Newsroom with the inconspicuous headline "M Now Offers Suggestions to Make Your Messenger Experience More Useful." It was a disguised departure from the original ambition. M Suggestions, rolled out across the United States, did none of what the original M was supposed to do — they didn't order flowers, couldn't book a table, didn't wait in a phone queue. They offered contextual links to existing Messenger features: split the bill, send a sticker, share a location, call a Lyft, write a birthday wish.

The phrasing in the announcement is remarkable for its passive aggression toward the original promise. "When we announced M over a year ago, it was a small, human-powered AI experiment that could fulfill almost any request. We learned a lot, and these interactions allowed us to build a fully automated version of M that suggests useful actions in chat." The assistant had become smart hints for the existing interface — that is, precisely the product that didn't need a human layer, because it did nothing that the AI of 2017 couldn't do.

Nine months later, on January 8, 2018, Facebook announced the complete shutdown of M; the service was switched off on January 19. The official statement read: "We launched this project to learn what people needed and expected of an assistant, and we learned a lot. We're taking these useful learnings forward into other AI projects." Romain Dillet at TechCrunch added what was missing from the announcement: "The secret sauce of M wasn't artificial intelligence — it was good old-fashioned humans." In May of that same year, Marcus left Messenger for a new blockchain group, from which the Libra project would emerge. At that moment, the wave lost its most visible salesman.

F8 2016 and the platform without a price

Between August 2015 and January 2018, there was one more moment without which the wave would have remained an internal Facebook experiment: Zuckerberg's keynote at F8. On the same day, April 12, 2016, Marcus unveiled the Messenger Platform (Beta). He announced three building blocks: the Send/Receive API, structured messages (images, buttons, receipt carousels), and the Wit.ai Bot Engine. The launch partners came on stage — CNN with a daily news digest, 1-800-Flowers, Spring, Poncho, Hipmunk, HP, Bank of America, and further eBay, Disney, Staples, Shopify, and Salesforce. Messenger had 900 million monthly active users at that moment, a billion messages flowed monthly between people and businesses, and Marcus framed the platform as "SMS on steroids for business."

One question was left unspoken on stage: how Facebook would make money on bots. The answer was silence — and that silence was the diagnosis.

Facebook introduced no transaction take-rate for bots. It didn't sell a single ad slot inside a bot conversation either. Sponsored Messages, which it launched in 2017 as paid distribution, were limited by policy to 24 hours from the user's last contact and could be targeted only at people who had already communicated with the business. Messenger CPMs were thin relative to the News Feed — users spent time in Messenger, but actively, not by scrolling; the ad inventory behaves differently here. Messenger ads themselves didn't ramp up until the end of 2017, and even then they were a commodity against the News Feed.

From Facebook's point of view, the platform made sense defensively: to stop Kik and Telegram before they could carve out a position as conversational interfaces, and to keep WhatsApp inside its own family. From the developer's point of view, though, the platform offered no monetization channel. The developer built a bot in Facebook's free-to-use environment, brought in customers from their own channels (web, email, SEO), and then ran their business on top of this foreign platform — without any control over the rules of access. Ben Thompson, on Stratechery in a piece dated March 31, 2016, named what was missing from the announcement: developers are commodity suppliers whom Facebook can throttle at will by changing the rules.

It throttled them. In 2018, Facebook tightened the 24-hour messaging window, and after the Cambridge Analytica affair added further cuts to access to user data. The entire category was built on a channel whose terms of access one player controlled absolutely — and that player began tightening the terms the moment it understood that developers on its platform weren't bringing in ad revenue commensurate with the cost of running Messenger.

Growth at first kept up with the promises. In September 2016, Marcus reported at TechCrunch Disrupt 30,000 bots and 34,000 developers; by F8 2017 it was already 100,000 bots and 2 billion messages between businesses and users per month; at F8 2018 Facebook stated 300,000 bots and 8 billion messages per month (four times the previous year's figure). The numbers looked like proof of growth, but they were deceptive — the count of registered bots is not the count of used bots, and the platform refused to disclose how many of them had more than fifty active users per month. Marcus himself admitted in September 2016 to Josh Constine of TechCrunch that "the problem was that it got so massively hyped so quickly… The basic capabilities we provided at the time weren't good enough to replace the traditional app interface." Developers reportedly had "a few weeks" to build the product. F8 2017 was a conference of remediation: a new Discover tab as a curated catalog of bots, a "hand-over protocol" for handing a single conversation between multiple teams, and Messenger 2.1 with built-in recognition of seven entity types — a quiet admission that third parties couldn't even do basic parsing.

Microsoft: ambition without a counter

In parallel with Facebook, Microsoft adopted the theory of the conversational platform. At the Build conference, from March 30 to April 1, 2016, Satya Nadella built the keynote around the phrase "Conversation as a Platform" and declared that this shift would have "as profound an impact as previous platform shifts — the GUI, the web, touch on mobile." Bots were to be "like the new apps"; the personal digital assistant (in Microsoft's interpretation: Cortana) was to be a "meta app, a browser over the bots."

Structurally, though, Microsoft was making a different mistake than Facebook. Facebook had distribution — 900 million Messenger users — and lacked a business model. Microsoft had Azure as an infrastructure business, and so it was selling tools instead of products. The Microsoft Bot Framework, released on March 30, 2016, contained an SDK, the Bot Connector as a routing layer for channels, and a public bot directory. Within a month, over 20,000 developers had registered; by Build 2017 there were 130,000. LUIS — the Language Understanding Intelligent Service — came out the same day as part of Cognitive Services and offered intent classification and entity extraction per API call. The developer paid Microsoft a marginal price per query for their project; Microsoft charged a margin for hosting and for the fact that the developer didn't have to maintain their own NLP infrastructure.

This monetization model worked. Except it worked on the infrastructure side, not on the consumer-bot side. Microsoft made money on developers who paid for Azure; it didn't make money on the fact that their bots communicated with users. Every dollar Microsoft poured into the platform as marketing or support had to come back through that developer's Azure bill — and only a small fraction of developers survived long enough for that bill to grow.

The wave, however, was struck a week earlier, already on March 23, 2016, by the episode with the chatbot Tay. Microsoft launched it on Twitter as @TayandYou with a target demographic of 18–24-year-olds; it was meant to be the Anglo-American sibling of the Chinese XiaoIce (which Microsoft had launched in China in 2014, with tens of millions of users) and the Japanese Rinna (July 2015). The conversational mechanics included online learning from tweets and a "repeat after me" command that literally repeated any input sentence. Users from 4chan and the /pol/ board coordinatedly poured racist, antisemitic, and misogynist content into Tay; within hours the account tweeted phrases about the Holocaust being "made up," that "Hitler was right," and that it "hates feminists." After sixteen hours and more than 96,000 tweets, Microsoft took the account down. Peter Lee, Corporate VP of Microsoft Research, published a post on March 25 titled "Learning from Tay's introduction" with an apology. Economically it was just a minor episode, but it made visible something that until then hadn't fit into the public debate: the human layer may not scale even in the domain of safety — and that raises the cost of every bot deployed to an unknown population of users.

In December 2016, Microsoft released the successor Zo — with a hard-coded refusal to discuss politics, religion, and race; the critic Chloe Stuart-Ulin described it in Quartz as "politically correct to the worst possible extreme." In May 2017 at Build, it launched the Cortana Skills Kit with 46 skills and the Harman Kardon Invoke speaker — behind Alexa, which already had over 10,000 skills. In May 2018 it bought Semantic Machines, a startup from Newton, Massachusetts, founded in August 2014 by Dan Roth (CEO) and Larry Gillick (CTO, former head of speech research at Apple Siri), with Dan Klein from UC Berkeley and Percy Liang from Stanford as the scientific backbone; financial terms were not disclosed. The ambition shifted from product to research — and the Microsoft Bot Framework later gradually morphed into Azure Bot Service, a B2B tool for customer support. From a conversational platform there remained an infrastructure line item.

The concierge startups: the purest fiasco

Beyond the giants, the wave spilled into hundreds of small experiments. The purest form was text concierge services — send an SMS, someone (or something) handles it for you. Of all the players of that era, the concierge startups demonstrated the equation of the human layer in its most legible form: without Facebook's platform complications, without Microsoft's B2B infrastructure alibi, just a founder, a team of trainers, and a product that was supposed to automate itself gradually.

Magic appeared at Y Combinator in the W15 class as a weekend project built on Twilio. After a viral launch on Product Hunt in February 2015, it raised, in March of that year, a $12 million Series A from Sequoia at a $40 million pre-money valuation — for a service that was itself nothing more than Twilio with a human layer behind it. In January 2016, eight months after launch, Magic launched the paid version Magic+ at $100 an hour or $3,000 a month. That was the first sentence in which the equation worked out: at a hundred dollars an hour, an operator could afford to handle complex tasks, because the margin covered their salary. But it was at the same time the last sentence in which the original thesis held — that a chatbot would replace apps. A product at $100 an hour doesn't replace apps; it replaces the personal assistant of a rich person.

Operator, the project of Uber co-founder Garrett Camp through his studio Expa, was led by CEO Robin Chan and raised a $10 million Series A from Greylock in April 2015 and a $15 million Series B from GGV Capital in September 2016. Unlike Magic, Operator tried to stay free — and to monetize through affiliate fees from purchases that users made through it. Chan, in an interview with TechCrunch, spoke of an average order value of $90 and a conversion rate of 10–15 percent, that is, dramatically above the e-commerce average. These were impressive numbers, but from a small base. And affiliate margins on purchases are single-digit in most categories — a user who ordered $90 shoes through Operator generated the company no more than $5–10 in commission, out of which it had to pay the expert who helped them with the selection. The equation worked out not much better than for Magic; it was just less visible until the product got out of beta.

GoButler, the Berlin equivalent, raised in July 2015 an $8 million Series A from General Catalyst with the participation of Lakestar, Global Founders Capital, and Ashton Kutcher's Sound fund. In January 2016 it moved its headquarters from Berlin to New York and gradually abandoned the original concierge model. In May 2016, under the new brand Angel.ai, it pivoted to a B2B NLP platform for third parties — precisely the type of pivot that revealed the economics of the original model didn't add up. In September 2016, Amazon hired founder Navid Hadzaad as "Product Lead, New Initiatives" — and officially denied that this was an acquisition of the company. Hadzaad, in an interview with Fast Company (2017), explained the shift toward a fully automated service in a single sentence that summed up the economics of the whole category: "My general view is that people will always want convenience, but they won't be willing to pay a premium for it."

Dennis Mortensen's x.ai (Amy and Andrew Ingram, a personal scheduling AI cc'd into email) raised in stages over $40 million. It lived for a long time in closed beta. Bloomberg, in April 2016, was the first to open the topic of the human layer behind Amy through a profile of a former trainer; half a year later, Business Insider, in an interview with Mortensen, gave concrete figures — out of 93 people at the company, 39 were "AI trainers" in the New York office at 25 Broadway, who reviewed practically every incoming email before Amy replied. The founder likened it to the safety drivers in Uber's self-driving cars — with the implicit promise that one day they would be superfluous. X.ai sold a subscription in the order of tens of dollars a month; at estimated costs of around $30–50 per active user per month (human editor plus infrastructure), this was a slightly negative margin throughout the company's existence. In May 2021, the company sold its assets to Bizzabo, and in October of that year shut down the original service.

The structural diagnosis across this category was identical, and it was summed up at the time, without illusions, by Ethan Bloch, the founder of Digit, in January 2018 for Inc magazine: "I'm not even sure we can say 'chatbots are dead,' because I don't know if they were ever alive." Bloch was speaking of financial chatbots, but the description fit the whole category. None of these startups managed to get unit economics into positive territory — not because the AI wasn't smart enough, but because even a hypothetical perfect AI wouldn't be able to bridge the gap between what the platform pays for a handled query and what the user is willing to pay.

Only those who already had cash flow survived

Above the field of dead startups, two small oases remained. But they were not demonstrations that the category eventually worked — they were demonstrations of the conditions under which a chatbot survives, conditions no consumer startup was able to reach.

Bank of America's Erica, presented at Money 20/20 in October 2016 and rolled out broadly in June 2018, rested on a team of more than a hundred people and claimed the position of one of the category's few genuine successes: by August 2025 the bank reported over 3 billion client interactions. Capital One's Eno, launched on March 10, 2017, at SXSW as the first SMS chatbot of an American bank, understood emojis (payment confirmations via thumbs-up were used by more than half of users) and gradually grew to 2,200 recognized ways of asking about a balance.

What both successes have in common is not the technology. It's the business model.

A bank doesn't need the chatbot to make money on its own. A chatbot in a banking app reduces the cost of the call center — every query Erica handles instead of a live operator in Charlotte is ten to twenty dollars saved. The bank has a captive customer base for the chatbot: the client already has an account with the bank, already went through KYC, already has a login in the app. They care about one thing — balance, transaction, payment — and they do it over and over, so the long tail of unique requests barely exists. And monetization is trivial: the bank makes money on fees, spreads, and interest; the chatbot is just another channel through which that money flows to it. No take-rate, no affiliate, no subscription needed.

That is precisely the opposite of what the concierge startups had at their disposal. A startup has no captive base — it has to steal it from Google or Facebook, at an acquisition cost that keeps rising. A startup has no repetitive queries — every new user arrives with a unique request. And a startup has no natural monetization channel — it has to invent one and defend it before a user who just typed a query in a messaging app where everything else is free.

The same rule also fixed the Microsoft Bot Framework — it transformed into Azure Bot Service for corporate customer support, that is, a tool for companies that already pay for Azure, already have their own customers, and need only a cheaper call center. LUIS and Dialogflow survived as moderately significant NLU products for narrow domains. From the Slack Fund — the $80 million fund from December 2015 co-financed by Accel, Andreessen Horowitz, Index, KPCB, Spark, and Social Capital — the startups that survived were those that pivoted from Slack-first bots to something else: Slack bought Howdy, Hugging Face abandoned the chatbot side entirely and became open machine-learning infrastructure.

Of the consumer concierge startups, Magic survived as a B2B platform of remote assistants — that is, again B2B billed by the hour, not a chatbot with an AI margin; Cleo and Lark as narrowly targeted fintech and healthcare apps with business models other than "converse a task." In all these cases the same rule holds: the survivors were those who stopped doing what Zuckerberg talked about at F8.

The mirage of WeChat: it was about payments, not conversation

The only argument the wave had at its disposal against these economic objections went: "In China it works." WeChat, with trillions of yuan flowing through the platform, and Line in Japan with two million active small-business accounts, showed, by that reading, that the conversational interface already works — you just have to replicate it.

Alongside Messenger, in 2015–2016 there existed several other environments that bet on bots — each with a different logic. Telegram was first; Pavel Durov launched the Bot API as early as June 24, 2015, ten months before Facebook, and on April 12, 2016, exactly on the day of F8, released Bot API 2.0 with inline buttons and location sharing — demonstrative timing, to remind the press of its primacy. Kik Interactive opened the Kik Bot Shop on April 5, 2016, with sixteen branded bots and 275 million registered users, predominantly American teenagers. The Line Bot API Trial, launched on April 7, 2016, plugged into the Line@ account system for small Japanese businesses, and by September 2016 roughly 20,000 chatbots had appeared there.

These two Asian examples (WeChat and Line) were the ones most frequently cited in Silicon Valley as proof that the conversational interface "works." As early as April 2016, though, Dan Grover, a WeChat product manager at Tencent, published an essay on dangrover.com that blew this reading wide open.

Grover's text of April 20, 2016 — published eight days after F8 — became the most-cited critique of the entire wave. He was writing from Canton and from inside WeChat: "The apparent success of messaging apps in fulfilling a surprising spectrum of tasks doesn't stem from the triumph of 'conversational UI.' It's a deft exploitation of the growing failure of mobile OS makers to fully serve users' needs — especially in other parts of the world." Grover then performed an experiment that became canonical: he compared ordering a pizza in the Microsoft Bot Framework demo (73 taps) with ordering it in the official Pizza Hut account on WeChat (16 taps, including dismissing a hint and a six-digit PIN). The difference was not in "conversation." The difference was that WeChat had removed app installation, login, entering payment details, and notifications.

This sentence is more economically loaded than it seems at first glance. WeChat succeeded not because it had a better conversational interface — it succeeded because it had WeChat Pay, QR identity, and a closed environment in which the user doesn't leave the app for any step of a purchase. When Tencent introduced Mini Programs in January 2017, these were not chatbots — they were full-fledged JavaScript applications inside WeChat, plugged into the same payment rail. Conversation with the store was just a surface layer; beneath it lay the payment rail and total identity, which no one in the American and European context had or could build. When Facebook tried to copy WeChat's surface (bubbles and bots), it copied exactly the part that was not the reason for China's success. The substrate that made the equation profitable — payments, identity, a closed ecosystem — it left unnoticed. Connie Chan of a16z summed this up on April 1, 2016, in a tweetstorm in a way that proved out: "Conversational commerce is unproven even in Asia. If typing takes more time than tapping a button in a webview, why is it better?" WeChat itself at the time made almost nothing from conversational commerce — it made money on payments, advertising, and games.

The critics who named the price

Alongside Grover, others spoke up around 2016, and only in retrospect can one see how precisely their critique was aimed at the economic core, not at the UX.

Benedict Evans of Andreessen Horowitz, in a March 30, 2016, piece "Chat bots, conversation and AI as an interface," noted ironically that "the challenge in connecting AI to a 'conversational' chatbot interface is that you don't have HAL 9000, but you're pretending to the user, in a certain sense, that you do." And he added that once a bot starts presenting choices, buttons, carousels, and cards, "you could call this, perhaps, a 'GUI.'" This remark was taken as aesthetic, but it meant something else: if a bot converges back toward a GUI, then it can be built directly as a GUI, without an NLU layer, without training data, without humans in the loop — and at a fraction of the cost.

Matt Hartman of Betaworks — who paradoxically co-founded Botcamp — formulated, in his 2016 and 2017 texts on the "Hidden Homescreen," a problem that sounded distributional but was commercial: "I don't have a link for MoveNight, because there is no website — the only way to refer to it is to tell you to add it in Kik by name. That's not just inconvenient. It has the potential for a viral coefficient of zero." A viral coefficient of zero means customer acquisition is 100 percent paid — and for a company with an affiliate margin of 5–10 percent, that means every new bot install is structurally loss-making.

Ben Thompson (as mentioned above) showed, via the theory of aggregation, that in the Messenger platform Facebook holds all the levers. Will Knight, in MIT Technology Review on April 25, 2016, put the question directly in the headline — "Is the chatbot trend one big misunderstanding?" — and pointed out that the successful Chinese chat services "avoid natural language in favor of more conventional input mechanisms — choices or buttons."

Standing apart from the product critics was Gary Marcus of NYU, who in the text "Deep Learning: A Critical Appraisal" (arXiv, January 2, 2018) enumerated ten structural weaknesses of deep learning. Marcus repeatedly tied the argument to the failure of bots: the promise that a system would "carry on a conversation" was unsustainable, because under the hood ran statistics, not understanding. And it was Marcus who expressed — without putting it that way — the thesis the article has been tracking from the start: if a model doesn't understand, the damage has to be made up with humans. And humans are expensive.

The counterargument: what if the AI had been better?

This analysis has one obvious weakness. It invites it. The whole thesis of the article rests on the claim that the problem was economic, not technological — but what if the transformer had appeared earlier? What if Vaswani et al. had published "Attention Is All You Need" not in June 2017 but in January 2015? Wouldn't that have solved the long tail of queries, wouldn't it have driven the share of human-handled responses to zero, wouldn't it have flipped the unit economics into the black?

This counterfactual has a peculiar advantage: it can be tested empirically. Ten years on, we are in an environment where transformers are in production. ChatGPT, Claude, and Gemini are massively deployed in 2026. The queries for which M needed a trainer, today's models handle autonomously. The cost of inferring a single query has fallen from a dollar of human labor to single-digit cents of compute time. The technological wall fell.

But the economic wall still stands.

The equation that merely got repackaged

OpenAI in 2025, according to projections reported by Sacra, is generating revenue of around $13 billion and burning over $8.5 billion in cash — with the projection that it won't be cash-flow positive before 2030. On inference alone, the company spent during the first three quarters of 2025 approximately $8.67 billion on Microsoft Azure — more than double what The Information estimated for all of 2024. In one quarter, ending September 30, 2025, a Microsoft SEC filing showed OpenAI, as an associated company, with a net loss on the order of $12 billion. Gross margin runs somewhere between 33 and 40 percent, structurally constrained by compute costs.

Anthropic, on a smaller scale, reported in 2025 costs of $2.7 billion against revenue of around $800 million. Perplexity spent in 2024 on AWS, Anthropic, and OpenAI together 164 percent of its full-year revenue. That is not a rate clients pay on top of the price — that is the amount the company paid through its suppliers.

In 2016 it held that for a query handled by the human layer, Silicon Valley paid around a dollar, while it saved the user on the order of three minutes. In 2026 it holds that for a query handled by a model, OpenAI pays roughly twice as much as it collects for it from the customer. The difference is that the query is qualitatively richer (code, research, explanation, a multi-step task), it can save the user hours instead of minutes, and the gap between input and output is financed by investors, not by operating cash flow. But the equation itself — the cost on the provider's side exceeds the customer's willingness to pay — is at its core unchanged.

What changed and what didn't

What changed: the absolute value AI creates for the user. Booking a table for three saved minutes wasn't worth a dollar; analyzing a twenty-page legal document for half an hour of saved work is worth ten dollars. The upper bound of willingness to pay has shifted. And the capital reserves have shifted too — OpenAI in February 2026 announced a $110 billion round at a $730 billion pre-money. A company that burns $17 billion a year (2026 projection, Sacra) can manage that from its balance sheet — whereas the concierge startup of 2016 burned out after $40 million and was done.

What didn't change: the economic structure. Every additional customer raises the inference bill linearly. A token generated at a longer context costs quadratically more (the KV cache scales with the square of context length). Frontier-quality models are priced below cost — aiautomationglobal.com summed it up in one sentence: "OpenAI, Google, Anthropic, and Meta are all selling inference below cost to gain market share." The moment the subsidies stop, either prices shoot up, or margins stay negative.

The 2015–2018 wave paid the deficit between price and value with the money of concierge startups and messaging platforms. The 2022–2026 wave pays it with the money of Microsoft, Amazon, Nvidia, SoftBank, and a few dozen other investors betting that Stargate-sized infrastructure will one day force cheaper compute. It is a different payment structure; it is not a different economics.

Two walls that remained

Two specific walls from 2016 are, meanwhile, still standing, and a model named GPT-5 has no leverage over them.

The first is the wall of monetization within someone else's platform. Messenger got repackaged into the browser, and from the App Store, for an AI agent, the channel became the Chrome browser and the Apple Intelligence interface. A developer who today builds an agent on top of ChatGPT Plus sits in the same role as a Messenger-bot developer in 2017: on a channel whose access rules are controlled by a single company. OpenAI can tomorrow change the API terms, raise prices, launch a competing product through its own GPTs store. Leaving the platform means moving the product to another channel with similar leverage — Anthropic, Google, or to one's own inference, which is usually more expensive. This is not a new problem, it is an old problem renamed.

The second is the wall of the UX paradox. For familiar, repeated tasks, conversation is slower than a button. ChatGPT's success doesn't rest on having pushed out buttons — it rests on tasks for which buttons don't exist ("write me an email," "explain this to me," "find claim X for me in a 200-page PDF"). That is true, but it also means that the category in which AI wins is structurally not customer-bound. No one has a banking model of a "captive base" for explaining abstract concepts or for creative writing. That is precisely why ChatGPT, Claude, and Gemini have to fight for the user anew every month — and precisely why the consumer AI segment has churn higher than subscription economics would warrant (17 percent monthly for ChatGPT Plus by estimates, while Netflix is at 2 percent).

The partial victories after 2022, then, are not convincing proof that the 2015–2018 wave was wrong only about timing. They are proof that if you have deep pockets and a market willing to subsidize a deficit for ten years, you can run an economically dubious business long enough for it to become part of the infrastructure. The chatbots of 2017 lacked three things: transformers (the technology), a hyperscaler with a vested interest (the channel), and ten years of a patient investor (the capital). The first thing the models caught up on. The second and third the capital market caught up on. And the economics as such waits for 2030, when it is supposedly meant to turn around.

Had the transformer in 2015 had enough investment and a Microsoft as a hyperscaler, the 2015–2018 wave would have looked like the 2022–2026 wave: billion-dollar losses, a public bet on the future cheapening of compute, endless rounds of financing. It would have looked more convincing, yes. But the economic equation would have come out the same — only no one would have demanded it come out right away.

The technological brake: what was really missing back then

The technological side, then, is not the central problem, but it's worth closing with a single paragraph, because it explains why even those who understood it could not have saved the category by improving the AI.

Production NLU in 2016 consisted of two cascaded supervised tasks: intent classification at the sentence level and slot filling / entity extraction at the token level. Wit.ai, LUIS, and API.ai all used variants of the same thing — a linear SVM or logistic regression over sparse n-grams or averaged dense vectors from word2vec and GloVe, and a linear CRF (conditional random fields) over hand-engineered features for slots. Rasa, founded in Berlin in 2016 by Alan Nichol and Alexander Weidauer, published this stack in its paper (Bocklisch et al., arXiv:1712.05181, December 2017) — and independent benchmarks of the time confirmed that the commercial interfaces were not algorithmically ahead, because everyone used variations of the same trio of SVM/CRF/BiLSTM.

The research frontier was racing ahead: Vinyals and Le in June 2015 showed a seq2seq LSTM in "A Neural Conversational Model," Huang, Xu, and Yu in August 2015 published the BiLSTM-CRF, Serban et al. in 2016 introduced HRED. The breakthrough came with "Attention Is All You Need" (arXiv, June 12, 2017), ELMo in February 2018, and BERT on October 11, 2018. Except no commercial bot platform during the peak of the wave was transformer-based; Rasa released DIET, its first transformer pipeline, only in 2020. In practice this meant that the 70-percent rate of uncompleted requests in Messenger, which The Information reported in February 2017, had no way to disappear other than through humans — and the economic equation therefore could not move other than by structural routes (a different business model, different distribution, different use), not by better technology.

What remained after the wave — and what came back

By the end of 2018 it was over. Facebook M ended in January 2018; Google Allo was discontinued by Google in March 2019; Microsoft absorbed Semantic Machines into a research Cortana, not into consumer bots; Cisco recoded the acquired MindMeld for corporate deployment for Webex and in May 2019 made it available as open source. Betaworks Botcamp, which in 2016 accepted 8 companies out of nearly 350 applications for $200,000 each (a SAFE note together with The Chernin Group), proved useful for history rather through one alumnus that switched from chatbots elsewhere: Hugging Face.

The pattern that survived was transparent: an existing captive base plus built-in monetization (banks), a B2B infrastructure product paid per API call (Azure Bot Service, LUIS, Dialogflow), or an escape from the conversational metaphor toward narrow fintech/healthcare deployment (Cleo, Lark). What remained were businesses that stopped chatbotting in the sense F8 2016 had planned — and began making standard software to which they gave a user-friendly text interface.

It came back differently, but with the same equation

The transformers that arrived late were absorbed elsewhere: first into search and translation, then into corporate assistants, and only in 2022 into consumer products. In November 2022, OpenAI released ChatGPT, and within five days it had a million users; within two months, a hundred million. The thesis that people would want to converse with a computer was revived in full force. And with it — more hidden, but identical — the equation that had buried the 2015–2018 wave. The human layer of trainers that M and x.ai paid for was replaced by an army of GPUs in data centers in Texas, Abilene, Virginia, and Iowa; the vocabulary of losses was softened from "cash burn" to "strategic investments in capacity." But the bill per query still exceeds the customer's willingness to pay.

The parallels between the waves are at times uncomfortably literal. Zuckerberg in 2016 promised that bots would replace apps; Altman in 2025 promises that agents will replace apps. Magic in 2015 operated with a human layer hidden behind an SMS interface; OpenAI in 2024 admitted that its "o1" model was trained on data labeled by cheaply contracted classifiers in Kenya. M relied on trainers generating training data that would one day render their own work superfluous; OpenAI relies on subscription users generating feedback that would one day make inference cheaper. In both cases the promise is essentially the same: pay us the deficit today, because tomorrow there won't be one. In both cases the empirical finding is that tomorrow arrives more slowly than planned.

Epitaph — or rather an interlude

The 2015–2018 wave did not end with the discovery that conversational interfaces don't work. It ended with Silicon Valley, for four years, mistaking one specific design — text plus NLU plus a bot in Messenger plus a human layer in the background — for a more general thesis about conversation as an interface. When Grover's critique was translated back into the American and European context, not enough was left of it for a business; not because the AI was weak, but because between the cost of the human layer and the user's willingness to pay for the service, no room was found.

What seemed at the time like a closed chapter, however, was not closed. The equation that buried the wave — that every query costs the provider more than the user pays for it — came back in 2022 in a better suit. It now has a transformer instead of a CRF, a data center instead of a call center in a New York office, and a hyperscaler instead of Sequoia. The deficit is larger in absolute numbers (OpenAI burns annually as much as the entire sector of concierge startups raised over its whole existence) and smaller in relative ones (a loss of $2 per dollar of revenue is harsh, but not as harsh as when Magic burned through everything it raised, and Operator raised so little it couldn't be counted).

The difference, then, is not in the equation, the difference is in the patience of capital. The chatbots of 2016 weren't given ten years — they got three. The chatbots of 2024 are, for now, holding on to patience, because they are embedded in a larger game about geopolitical standing in AI, and it pays the hyperscalers to feed them even at a negative margin for as long as it constrains the competition's room to maneuver.

Zuckerberg's sentence from F8 quietly fell apart in its twenty-second month — not because people didn't want to talk to a computer, but because around every conversation there had to sit a human who managed it. In 2026, a human no longer needs to sit around the conversation, the model handles it on its own. Except that model costs twice what the company collects from the customer. Three minutes for a dollar was a bad deal in 2016. An hour of expert assistance for the twenty dollars the company takes in, and the forty dollars the company pays, is a better deal — but still negative.

And this equation is only now awaiting its real answer.

Methodological note

The article is a reconstruction of the period January 2015 to January 2018 based on secondary sources published at the time (Wired, BuzzFeed, Fast Company, MIT Technology Review, The Information, TechCrunch, Recode, Bloomberg, Business Insider, Quartz, Stratechery, Mashable) and primary sources provided by direct participants (Facebook Newsroom, Microsoft Official Blog, company announcements, arXiv pre-prints). For the comparison with the current state of the AI economy (the "Counterargument," "What remained," and "Epitaph" sections), data from the 2024–2026 period is used, from the following sources: Sacra (sacra.com/c/openai/), Microsoft SEC filings, leaked internal documents reported by Ed Zitron (wheresyoured.at) and reprinted by The Register, The Information, and aiautomationglobal.com. All specific dates, numbers, financial figures, investors, and quotes come from sources captured in the research material. The figures for bot and user counts are reported as the platforms stated them (typically at the F8, Build, and TechCrunch Disrupt conferences) — that is, with the awareness that these are platform communications, not independently audited metrics.

• The unit costs of $0.50 to $1 per American task are period estimates from investigations by Bloomberg, Business Insider, and MIT Technology Review, not the companies' accounting data.

• The figure of "roughly 70 percent of responses handled by humans at M" comes from later reporting (TechCrunch, The Information) and was never officially confirmed or denied by Facebook.

• The specific ratio of 39 AI trainers out of 93 employees at x.ai comes from Dennis Mortensen's interview with Business Insider (late 2016), not from Bloomberg's original April text.

• The estimate of x.ai's costs at $30–50 per active user per month is the author's calculation from public figures (39 trainers out of 93 employees, estimated New York wages, the number of active users in the paid phase), not a company figure.

• The acquisition prices of Wit.ai and Semantic Machines were not disclosed; the article therefore does not state them.

• For some quotes (e.g., Connie Chan of a16z), the primary source is Twitter/X, whose archival availability has since changed.

Limits of the comparative part (2024–2026):

• The figures on OpenAI inference spend ($8.67 billion for Q1–Q3 2025) and revenue ($2.27 billion for H1 2025 vs. $4.3 billion reported by The Information) come from Microsoft internal documents obtained by Ed Zitron. They are consistent with the MS SEC filings, but not all of the component figures are audited.

• The $12 billion loss at OpenAI for the quarter ending September 30, 2025, comes from a Microsoft SEC filing in which OpenAI appears as an associated company; OpenAI itself does not disclose the details.

• The gross margin of around 33–40 percent derives from two different sources (Sacra and Ed Zitron's sources) of matching order of magnitude; absolute accuracy is not verifiable.

• The claim of $2 in cost per $1 of revenue is a rounded characterization of the ratio between the reported cash burn (~$8.5 billion in 2025) and the reported revenue (~$13 billion in 2025) after including infrastructure investment; different interpretations give ratios in the range of 1.3× to 2.5× depending on what is counted into the costs.

• The estimate of 17 percent churn at ChatGPT Plus is a published estimate from secondary sources; OpenAI does not disclose official retention.

• The figure of 164 percent of Perplexity's revenue spent on suppliers in 2024 comes from internal documents cited in Where's Your Ed At.

The concept, structure, and editorial line of the article are the work of the author, who developed the content outline, set the key theses, and directed the entire creative process. Generative AI (Claude, Anthropic) was used as a tool for research, fact-checking, and fleshing out the author's draft.

The author edited the outputs on an ongoing basis, verified the key findings, and approved the final wording. No part of the text was published without human review. All factual data were verified against the publicly available sources cited in the text.

The procedure complies with the requirements of Art. 50 of EU Regulation 2024/1689 (the AI Act) on the transparency of AI-generated content. #poweredByAI

Read the Czech original on Médium.cz.

AI · Claude — machine translation, may contain inaccuracies.