AI Economics / No. 025
AT&T Routed the Easy Tokens Off the Frontier Bill
Andy Markus has open models at 40 percent of AT&T's AI use and an 80 percent cost cut versus earlier this year. OpenRouter's 58 percent is a shopping mall, not a census. Closed labs keep the hard work if they can still get paid for it.
AT&T used to buy the closed package for the obvious jobs. Customer service, call transcription, coding help. Anthropic and OpenAI. Fees, no weights. Eli Tan’s New York Times story on September 4 is the update. Andy Markus, the company’s chief data and AI officer, said open models were 20 percent of AT&T’s AI use by May, 40 percent now, and maybe 60 percent in the coming months. Cost versus earlier this year is down as much as 80 percent. Airbnb and Deloitte show up in the same piece doing a version of the same move.
I take the 80 percent cut as a real number, not a slogan. A telco’s transcription and ticket macros are high-volume work with a grade you can actually compute. You do not need the model that wins a reasoning contest to turn a call into text. If Markus can park that layer on a downloadable model and keep Claude or GPT for the leftover, the invoice moves even if the leaderboard does not.
Jerry Tang, who runs Atlas Cloud, gave Tan the split that matches that invoice. Closed models still look better on heavy coding and on image and video generation. Open models often win on simple, specialized jobs. AT&T is running that split in production. Markus said the company takes open-weight models and customizes them for transcription and customer service. Llama did not replace the frontier stack. The frontier stack was sitting on a lot of tokens it did not need to touch.
OpenRouter’s figure will get quoted as if it were a market-share print. Tan wrote that last month open models accounted for 58 percent of AI use, up from 10 percent a year ago, according to U.S. user data from the routing platform. OpenRouter is a mall for people who already want to pick among models. It is a decent gauge of substitution among shoppers. It is a poor census of what a bank, a hospital, or a regulated telco spends on contracted APIs, on-prem GPUs, and Microsoft or Google bundles that never hit OpenRouter. Read 58 percent as evidence that the outside option is being used. Do not read it as 58 percent of corporate AI budgets walking out of the closed labs.
The geography of that outside option is the awkward part. Tan notes that many of the most popular open models come from Moonshot, DeepSeek, and Alibaba. Tang put some Chinese open models at 80 to 90 percent of closed-lab capability at as little as 20 percent of the price. AT&T is not taking that trade. Markus said the company researches Chinese models and does not use them. It is working with Google’s Gemma and Meta’s Llama instead. Last month Meta made its leading system an open-weights model and said it would price below other top U.S. models. Mark Zuckerberg’s accompanying line was that blocking foreign open-source models is not an effective solution, and that American open-source models should be the best globally.
So the Fortune 500 version of getting hooked is narrower than the headline. The buyer wants cheaper tokens and the right to fine-tune. The buyer still wants a U.S. name on the weights. Meta and Google pick up distribution they could not get as closed APIs. Nvidia’s $12.9 billion Hugging Face purchase, which Tan flags in the same story, sits on that stack. Someone has to host, rank, and operationalize the downloads. The inference bill does not vanish. It moves from a lab subscription onto servers, chips, and a catalog.
That movement is the listing problem. Tan’s piece says OpenAI and Anthropic are heading toward large offerings and have to charge more because they are spending billions on research and computing power. Neither company commented. A mix shift of the AT&T shape does not require the frontier product to fail. It requires the easy tokens to stop subsidizing it. If 40 percent of a large buyer’s volume leaves, and the rest is the hard work, the lab can still have a business. The business has to look more like scarce escalation capacity and less like a default meter on every employee prompt.
David Stout of WebAI told Tan that many open models now run on a phone or a laptop, while leading closed models are getting more complicated and more compute-hungry. That is a cost-curve split. It also explains why open does not mean free. Markus landed on models that can be downloaded and modified without payment or approval. AT&T still pays for the machines that run them, the people who customize them, and the evals that decide when a job is too hard for the cheap path. Dario Amodei’s case for tighter control of frontier weights is a different argument. It does not stop a telco from routing call summaries off the expensive API.
Three ways this can go. For the closed labs, the kind case is that AT&T-style routing caps the cheap work and leaves them a high-willingness-to-pay remainder that can fund the next training run. The middle case is a fork. More large U.S. shops copy the 20-40-60 staircase on Gemma and Llama, OpenRouter shoppers keep using Chinese models, and the two markets stop being the same market. The ugly case for the listings is that the remainder shrinks too. Once transcription quality is close enough, coding assistants get the same treatment AT&T already started, and the premium slice is a thin set of multimodal and high-stakes jobs.
Watch Markus’s next mix print, and watch whether the 80 percent cost cut shows up as a smaller Anthropic and OpenAI line or as a larger GPU line. Those two ledgers will not move together.