July 21, 2026 banner
In Today's Newsletter
China's model tops the leaderboard, and OpenAI concedes FULL STORY
AI agents score around 25% on actual job tasks FULL STORY
AI hiring tools invented their own bias FULL STORY
What else happened today?What AI tools should I be using?

Good Morning Thorium Valley. A Chinese AI lab just took the #1 spot on one of the most-watched coding benchmarks, and OpenAI's response was basically "yeah, we can't explain that one away." Moonshot's Kimi K3 beat every US model in head-to-head matchups, and it's going open-weight next week. David Sacks is already using it to argue America is regulating itself into second place.

Meanwhile, Berkeley put the best AI agents through an actual job test — 55 occupations, real paid tasks — and the top score was 26%. Three-quarters of the failures weren't even bugs. The agents just didn't understand what they were being asked to do.

And if you've applied for a job lately, an AI probably screened your resume before any human saw it. New research says those tools aren't just borrowing human bias. They're inventing entirely new stereotypes on their own, even for demographic groups that don't exist. Ninety percent of US employers use these systems, mostly from the same handful of vendors, so getting rejected once basically means getting rejected everywhere.

Quickly before we dive in — Should companies be required to disclose when AI is screening your job application?

Yes | No | Other

RESEARCH

China's model tops the leaderboard, and OpenAI concedes
Share X in

A Chinese AI lab just released a model that beats everything OpenAI and Anthropic have on one of the most-watched coding benchmarks — and this time, US labs aren't blaming distillation.

Kimi K3, from Chinese startup Moonshot, took the top spot on the Frontend Code Arena, a public leaderboard where developers vote on which model writes better code in head-to-head matchups. It's the first time a Chinese model has hit #1 on this arena. Moonshot has promised to release it with open weights by July 27, meaning anyone will be able to download and run it.

The reaction from OpenAI was the real story. Dean Ball, the company's head of strategic futures, publicly acknowledged that K3's performance "can't be explained away by distillation or anything like that." That's significant — distillation has been the go-to US explanation for why Chinese labs keep catching up so fast. Ball essentially admitted that excuse doesn't work here.

David Sacks, the Trump administration's AI adviser, treated the news like a warning shot, arguing that a Chinese model taking #1 while "America is tying itself in knots" over regulation is a serious problem — and accusing the big closed US labs of pushing Washington to squeeze out open-source competition at exactly the wrong moment.

K3 isn't a one-off. According to the ATOM report tracking the open model ecosystem, Chinese models have surged from near-zero to dominating open-source inference usage in just over a year, while Meta's Llama has collapsed to irrelevance. Alibaba's Qwen has quietly become the base model developers actually build on.

That said, the closed labs aren't losing across the board. Anthropic and OpenAI still command a premium on harder, higher-value tasks, and on broader benchmarks, K3 still trails GPT-5.6 and Fable 5. The gap at the top is real — it's just no longer a chasm, and the floor keeps rising.

Into the Valley

Every few months, someone in the US declares that Chinese AI is a knockoff, and every few months that argument gets harder to make with a straight face. K3 is the version where OpenAI's own strategy lead admits it out loud. The uncomfortable question for Washington isn't whether China caught up. It's whether the US strategy of restricting exports and pressuring open-source at home actually made things worse by handing the free tier of the global market to Beijing. That's the fight Sacks is picking, and it's about to get a lot louder.

RESEARCH

AI agents score around 25% on actual job tasks
Share X in

The best AI agents in the world just took a real-world job test — and mostly failed it.

Berkeley's Center for Responsible, Decentralized Intelligence released Agents' Last Exam, a benchmark that measures how well AI agents handle actual paid work. Not coding puzzles — real tasks across 55 occupations and 13 industries, from legal research to financial analysis to HR. The top score? Just 26.2%, from OpenAI's Codex running GPT-5.5. On the hardest tier, nearly every agent scored at or near zero.

The gap between ALE and narrower benchmarks tells the whole story. The same Codex stack that scores 82% on Terminal-Bench managed just 25.2% on ALE's coding tasks. Same model, roughly one-third the performance — because ALE touches 40 subdomains instead of six.

As Carnegie Mellon's Graham Neubig, who worked on the benchmark, put it: "The age of useful agents is here. The age of truly job-ready agents is not."

Where agents break isn't what you'd expect. About three-quarters of failures aren't bugs — they're the model misunderstanding the task or picking the wrong approach entirely:

+ Wrong strategy: 30% of failures came from confidently charging ahead with the wrong plan

+ Domain knowledge gaps: 25% from simply not knowing enough about the field

+ Incomplete or abandoned tasks: 17% — the agent just gave up partway through

+ Actual implementation bugs: Under 10%

One telling detail: about a third of ALE's tasks require graphical software that normal office workers use every day. Agents largely refused to touch them, trying to hack around the interfaces with command-line scripts instead. And cost compounds the problem — Anthropic's Fable 5 runs about $15.70 per attempted task, roughly four times OpenAI's stack, for similar success rates. When you're only completing a quarter of the work, that adds up fast.

This lands in the middle of a heated debate about AI's economic impact. Sixteen Nobel laureates signed an open letter this month calling for urgent policy attention, warning that AI capabilities are advancing faster than our understanding of what to do about them. ALE is a useful check on both sides: agents are nowhere near replacing knowledge workers, but the industry has spent a year grading itself on tests it built to pass, which makes the gap easy to miss.

Into the Valley

The real question ALE raises isn't whether agents will get better, because they will. It's whether the industry has been measuring the wrong thing this whole time. A model that aces Terminal-Bench and flunks a spreadsheet task isn't actually close to doing your job, no matter how good the launch keynote sounded. If ALE becomes the benchmark that matters, the leaderboard stops being about who can write the cleanest Python and starts being about who can get through an afternoon of ordinary office work without getting lost. That's a much harder problem, and it's the one that actually decides whether any of this pays off.

RESEARCH

AI hiring tools invented their own bias
Share X in

If you've applied for a job recently, an AI probably read your resume before a human did. New research suggests it may have judged you on stereotypes it made up on its own.

A paper accepted to ICML by researchers at Princeton and the University of Chicago found that large language models don't just inherit human bias when screening candidates — they generate new stereotypes from scratch. In one experiment, models made hiring decisions about applicants labeled with completely fictional demographic groups. No cultural baggage, no training data to learn from. The models still developed consistent preferences, essentially inventing prejudice from thin air.

As one of the coauthors put it, LLMs are optimized to generalize from limited data. That's what makes them useful — and what makes them dangerous when they're sitting between people and jobs.

A separate Stanford HAI field study tracking 3.4 million real applicants across 156 employers showed how that danger scales. Because most companies use the same handful of AI vendors, a candidate rejected by one system tends to get rejected by all of them. About 10% of applicants who submitted four applications were shut out everywhere — a rate higher than random chance would predict. Stanford's Dan Jurafsky said the pattern is arguably worse than the human bias it replaced, because AI systems are far more likely to act identically than independent human reviewers would be.

The numbers are stark:

+ 90% of US employers use AI screening tools, most relying on the same small group of vendors, according to Stanford HAI.

+ 26% of Black applicants and 15% of Asian applicants in the study applied to positions where the AI actively discriminated against their group.

Regulation exists but barely functions. New York City's Local Law 144 requires employers to audit hiring AI for bias, but a state comptroller audit found the law is mostly theatrical — city regulators flagged one likely violation across 32 employers, while independent auditors looking at the same companies found 17.

Into the Valley

The pitch for AI hiring was that it would strip the messy human stuff out of the process. Take the gut feeling, the pattern recognition, the lazy shortcuts, and replace them with math. It turns out the math has its own shortcuts, and because everyone bought them from the same vendors, the shortcuts are now industry-wide. The next few years of hiring lawsuits are going to be about who actually owns that mistake, whether it's the employer that deployed the model, the vendor that sold it, or the regulator that never checked. Right now the answer looks a lot like nobody.

In Other News

IN OTHER NEWS

What else happened today?

+ A mathematician used Claude Fable to disprove an 87-year-old math conjecture during the World Cup final

+ TSMC adds another $100 billion to its Arizona investment, bringing total commitment to $265B

+ Netflix paid $587 million for Ben Affleck's AI filmmaking startup — it's already been used on 300 titles

+ Meta is in talks to lease $10 billion in compute to rival Anthropic, entering the cloud business for the first time

+ Trump administration launches 'Gold Eagle' program to control who gets access to frontier AI models from OpenAI and Anthropic

+ Hugging Face confirms an AI agent hacked its systems — and another AI caught it

+ Lawsuit alleges ChatGPT encouraged a woman's suicide by stoking her religious delusions before she walked into traffic

+ District 9 director Neill Blomkamp releases first short film made entirely with AI video generation, plans a full feature next

WHO'S HIRING IN AI

+ Anthropic — Research Engineer, Machine Learning (Reinforcement Learning)

+ Runway — Engineering Manager, Machine Learning ($310k–$370k)

+ Perplexity — Member of Technical Staff, AI Software Engineer (Agents)

+ Shopify — Applied Machine Learning Engineer

AI or Real?

AI OR REAL?

One is AI. One is real. Can you tell?
Option A

Option A

Option B

Option B

Which image is real?

Option A | Option B

Yesterday's Results
AI Tools

AI TOOLS

What our editors are paying attention to today

+ ChatGPT: OpenAI rolled out a universal search tool that lets you find old chats, uploaded files, images, and projects all from one search bar — available free on web, iOS, and Android

+ Google Vids: Google's Workspace video editor now lets you edit clips by typing instructions — fix colors, change styles, or remove background noise using the new Gemini Omni model

+ Spotify: Premium users can now have a back-and-forth conversation with the app to pick music, ask about artists mid-song, and refine recommendations without leaving Spotify

+ Google Docs: Gemini's AI writing and editing features now work in 11 new languages including Mandarin, Hebrew, and Polish, with faster idea-to-draft tools

+ GitHub Copilot: Business and Enterprise users can now see exactly how many AI credits they've burned each billing cycle instead of just a vague percentage bar

That's all for today. If this issue made you think, share it with someone who needs to think harder.

Written by Jason Chen, Advait Prakash, Andrew Hales, and the Thorium Valley crew.

That's all for today's Thorium Valley. See you tomorrow.