A brand new era of automation has supposedly arrived. Do you feel it? The leading AI companies clearly believe so, but ordinary people may feel nothing at all. On a rare slow news day, we had time to carefully ponder the significance of OpenAI's major announcement on Tuesday evening: its yet-to-be-released model made breakthroughs on hundreds of major mathematics and computer science problems, even solving some of them outright. For those of us who don't work through algebra in our daily lives, what does this actually mean? It is hard to separate these major technical achievements from the commercial hype designed simply to sustain AI's fundraising momentum. I noticed that while Tuesday's news was arguably a milestone, at the same time there were reports that OpenAI is in talks to raise funding and preparing for a $30 billion IPO; and a week earlier, its much-hyped developer day was viewed by outsiders more as playing catch-up with rivals.
At the same time, the latest AI craze — AI agents aimed at ordinary users — is starting to hit obstacles: a large number of websites plan to block AI crawlers. And in macroeconomic statistics, there is for now no sign that AI is boosting society-wide productivity. This round of AI gold rush is haunted by one core question: how should companies convert models' astonishing raw intelligence into real, verifiable business value? Not long ago, I witnessed this debate at Midway, a dimly lit hall on the San Francisco waterfront that often hosts electronic music festivals and art exhibitions. Last week, nearly 800 software developers gathered there for a conference hosted by Modal. Modal is not an electronic music DJ, but a highly valued company that provides developers with compute infrastructure and software tools. Modal handed its keynote to Scott Wu. He first became well known for his standout performances in math competitions; his code agent startup Cognition saw its valuation soar to $48 billion last month. Standing on stage, with a slide behind him reading "It's time to think big," he painted a not-too-distant future in which enterprises are filled with "virtual employees" that aim at outcomes rather than merely completing one task after another. He said that with enough data centers, memory, and sandbox environments, the industry will eventually realize this vision. "Agents just keep getting more capable, and that's the trend," he said.
The confidence behind his call to action came from breakthroughs in mathematics. "For me, the turning point was that AIME moment," he said. AIME refers to last year, when advanced AI models became so strong at reasoning that they outperformed humans in the high-level math competition he once competed in. "At that moment I knew the die was cast" — meaning humans no longer have a monopoly on intelligence. That same afternoon at the same conference, I heard a very different message. Diogo Almeida, a former OpenAI researcher who describes himself as a "co-author of several of OpenAI's hit models," took the stage in a pink fur coat with a smile, his tone carrying the zeal of an insider who had left a troubled country or belief system and could now tell everyone the truth. He now heads TypeSafe AI, which recently launched a widely noticed alternative to large models called Jev. He asked the audience: "Where did automation go?" He then rattled off a string of acronyms familiar to those in the field. "My one-sentence summary of the current state of AI: all large language models today are assistive tools optimized through reinforcement learning from human feedback (RLHF)." (For those unfamiliar: RLHF stands for reinforcement learning from human feedback.) In other words, large models are trained to please the humans involved. They are good at assisting people, but they cannot automate work in unattended settings.
At the same conference, several AI industry leaders were pushing for automation in software engineering while also acknowledging that this transformation faces heavy resistance. They also did not sound ready to hand everything over to machines. Cat Wu, product lead for Anthropic's Claude Code, said: "One direction we're focused on is thinking: which tasks am I always repeating? Why am I still stuck in this loop?" Anthropic is perhaps the company that believes most strongly in AI's potential worldwide. "A lot of automation gets stuck in semi-automation: Claude completes 80% of the work, and the remaining 20% still has to be done by a person, and you have to repeat those tasks over and over." Some software engineering leaders also feel that managing human engineers is a more troublesome problem than taming AI agents. Dax Raad, the developer of the well-known open-source project OpenCode, said their team's biggest current headache is that human engineers do not carefully review code written by AI. "The problem I want to solve now is how to get the team to ask the right questions; once unreviewed code goes live, they should feel guilty," Raad said. "This is a long-standing problem: how do you make people truly care?"