Brendan Foody on Teaching AI and the Future of Knowledge Work
Brendan Foody — co-founder and CEO of Mercor, the marketplace that supplies AI labs with the domain experts who build the rubrics and evals models are trained and measured against — joins Tyler Cowen to argue that the scarce resource in AI is no longer text to train on but rubrics: structured ways of measuring success that let a model learn from repeated, graded attempts. The conversation ranges from why Mercor pays poets $150 an hour, through a Kantian challenge to whether taste can ever be reduced to a rubric, to Foody’s claim that most knowledge work is about to become a one-time teaching investment rather than a recurring task.
Key ideas
- Rubrics, not raw text, are the scarce input. Foody distinguishes two kinds of training data: curriculum — text a model reads and imitates once — and measurement — a rubric, a graded test, a unit test — that lets a model attempt a problem repeatedly and learn from the outcome. He calls the second kind far more valuable, and argues academic institutions sitting on rich measurement data (referee reports, seminar transcripts) are an underused resource the AI labs would pay for.
- APEX measures economic value, not academic difficulty. Mercor’s AI Productivity Index surveys hundreds of domain experts (ex-McKinsey, Bain, BCG consultants; doctors; lawyers) on how they actually spend their working time, then derives evals weighted by that time allocation — a deliberate correction to benchmarks like GPQA and IMO, which Foody says are disconnected from what customers actually want automated.
- Taste resists a rubric — Foody’s answer is RLHF, not a better rubric. Pressed with Kant’s claim in the Critique of Judgment that taste is precisely what cannot be captured in a rule, Foody concedes some domains (poetry especially) are too unbounded for a rubric and require human preference-ranking (RLHF) instead — a different data type from a rubric, not a refinement of one.
- Knowledge work becomes a one-time teaching investment. Foody’s central economic claim: instead of repeating the same analysis on every deal or ticket, an expert will increasingly build an RL environment once — encoding a task and how to judge success — after which an agent repeats it indefinitely, mirroring how software replaced recurring bespoke labour with a fixed cost copied at scale.
- Job survival tracks price elasticity, not task automatability. Roles where AI-driven efficiency gains expand demand (software engineering: make engineers 10x more productive and the market wants 10x more software) grow; roles with roughly fixed demand (accounting, some customer support) shrink under the same productivity gain. Foody expects a new job category — people who find where a model fails and turn that failure into a trainable RL environment — to employ a majority of high-end knowledge workers within five years.
Content
Measuring real-world AI progress
Foody opens by explaining why Mercor pays some of the world’s best poets $150 an hour: they write rubrics that let a lab measure and improve a model’s poetry once, at a skill that then applies across billions of users. Asked what Mercor’s group of hired experts — Larry Summers on finance, Cass Sunstein on law, Eric Topol on medicine — actually learned building APEX, Foody reports a 25–30% per-year improvement rate on economically valuable tasks, with GPT-5 scoring 64% [?]. The methodology surveys hundreds of professionals on how they spend their time (client meetings, research, analysis, deliverables) and converts each bucket into prompts and rubrics, using time allocation as a proxy for the economic value of that work. He is candid that mapping a score to real-world impact is domain-dependent: medicine tolerates almost no failure rate, while an imperfect legal draft or consulting analysis is already useful.
Why rubrics are the new oil
Cowen presses Foody on what data he would want if realism were no object — recorded academic seminars, referee reports, curated datasets. Foody’s answer distinguishes curriculum data (text the model reads and learns from once) from measurement data (a rubric, a graded test-and-answer pair, a unit test), arguing the second is by far the more valuable, because it lets a model attempt the same problem many times and learn from the scored outcome rather than absorbing an example passively. He suggests academic institutions — which already hold referee reports and seminar transcripts rich in exactly this measurement data — are an underused resource that AI labs would pay for, and speculates the reluctance is partly a lack of awareness of why evals matter and partly a more general anxiety about AI’s effect on jobs, even inside nonprofits nominally committed to abundance.
Enshrining taste in LLMs
Cowen invokes Kant’s Critique of Judgment — that taste is exactly what a rule or rubric cannot capture — against Foody’s claim that a rubric is the data Mercor most wants. Foody concedes the point for genuinely unbounded domains like poetry, where he says a rubric is harder to write and apply than for a well-scoped essay, and falls back on RLHF: generate two candidate responses, have someone with strong taste choose the preferred one, and repeat until the model has absorbed the preference pattern. The conversation extends into whose taste should be enshrined at all — Cowen notes that studies show ordinary readers often prefer AI-generated poems that expert critics judge worse, and that the very best poetry, by most people’s own standard, is centuries old (Milton, Wordsworth, Blake), so training only on the present generation’s taste risks enshrining a taste inferior to the past’s, with no way to resurrect history’s best judges to grade against. Foody’s guess is that models will eventually draw on the taste of every era and personalise which one a given user prefers, rather than the field settling on a single standard.
Turning society into one giant RL machine
Asked how much of society should ideally become a giant reinforcement-learning machine — taping every debate, every meeting — Foody’s answer is a specific economic mechanism rather than a surveillance fantasy: an investment banker who redundantly re-analyses a data room every few weeks will instead teach an agent to do that analysis once; a customer-support rep who fixes the same recurring mistake will turn that fix into a trainable RL environment, after which the agent solves the problem indefinitely. He draws the direct parallel to software’s economics — a fixed-cost investment, built once, used at effectively zero marginal cost thereafter — as the reason he expects a large share of the economy to convert into RL-environment-building. He does concede a limit: most people will not want their private, off-the-record conversations recorded and folded into a model’s training weights, whatever value that data might carry, and predicts brands with an existing trust advantage around privacy (he names Apple) will have a structural edge in collecting it anyway.
When AI will stump experts
Using Cass Sunstein as a concrete test case in law, Foody guesses it will take two to three years before Sunstein struggles to find an error in a model’s legal reasoning, because so much of legal judgement lives as undocumented tacit expertise rather than codified rule; Cowen guesses six months and the disagreement goes unresolved. On poetry, Foody expects a model to match the median Pablo Neruda poem within a year, but will not forecast when a model matches Neruda’s very best work, agreeing with Cowen that the long tail of quality is the hardest part of any domain to close. His standing heuristic: models are advancing fastest on the 50–75% of expert work that is codified or learnable from data, and will struggle for a long time with the last quarter — the residue of tacit, undocumented judgement that keeps human expertise economically necessary.
Why vibes-based hiring fails, and how AI-run labour markets might work
Turning to Mercor’s own hiring practice — the company has grown to just over 300 employees — Foody argues most conventional interviewing measures personality fit (‘do we want to hang out’) rather than the actual skill the job requires, and that the fix is to grade candidates on realistic project work instead, including work done with AI tools rather than banning them. He names the harder residual problem in talent assessment as measuring a candidate’s slope — how fast they will improve on the job — versus their easily measured current ability, their Y-intercept. Scaling this logic to the labour market as a whole, Foody argues the central inefficiency is that hiring is radically disaggregated (a jobseeker applies to a few dozen roles; a company considers a fraction of a percent of the workforce) and that the missing structural piece is not distribution (LinkedIn already solves that) but matching — reliably predicting who will perform well from a profile. He expects this to improve as full-time roles fragment into remote, fractional model-training work, though he partly concedes Cowen’s worry that as AI-optimised applications flood the market with plausible-looking candidates, hiring may revert to referrals and nepotism in some firms — his hope being that models evaluated against real outcome data can eventually discount unearned nepotistic signal rather than reward it.
Scaling the Thiel Fellowship, and a hypothetical gap year
Foody, himself a Thiel Fellow who dropped out of college at 19 to start Mercor, says the Fellowship faces the identical matching-and-scale problem as hiring generally, and reports Mercor has already worked with it on AI-assisted interviews and transcript analysis to improve fellow selection. Pressed on whether the Fellowship’s distinctive brand depends on staying a small, self-selecting, ‘ornery’ pool, Foody grants that referrals matter but argues broadening the aperture serves unconventional candidates who would otherwise never reach a venture capitalist, and expects the selection process could scale by orders of magnitude with the right technology — while conceding that an aggregate panel of domain experts will likely out-predict any single interviewer, including Peter Thiel himself, for a long time. Asked what he would do with a magic, cost-free year away from the company, Foody — who says he has worked roughly a hundred-hour week for three years — picks travel, specifically Japan, framing it as the individual analogue of a model that has absorbed the poetic taste of many eras: exposure to how different countries and cultures are thinking about AI.
On AI and employment
Foody’s frame for which jobs survive AI-driven productivity gains is price elasticity of demand, not raw automatability: making software engineers ten times more productive, on his account, will not shrink the profession but expand it, since the market wants far more software once it becomes cheaper to build; he expects the same dynamic in business-building and capital allocation more broadly. Domains with comparatively fixed demand — he cites accounting, some customer support, and education (there are only so many hours a day for a personal Sal Khan-style AI tutor) — are more exposed to net job reduction under the same efficiency gain, though he expects teachers specifically to retain an important role in the emotional and relational side of education that a model cannot fully replace. He expects a new job category to emerge and, within five years, employ a majority of high-end knowledge workers: people whose job is to find where a model makes a mistake and convert that mistake into a trainable RL environment, replacing the recurring version of their own former analyst-style work.
Donuts, debates, and dyslexia
In a lighter biographical stretch, Foody traces his taste for markets and scale back to an unlicensed eighth-grade doughnut resale business (‘Donut Dynasty’) — buying Safeway doughnuts at $5 a dozen and reselling them at school for $2 each, undercutting a rival on price before he ‘had learned anything about anti-competitive laws’ — and to a champion high-school policy-debate team where he met his Mercor co-founders at 14. He is dyslexic, and links this, via a correlation he says is well documented in the entrepreneurship literature, to early, forced practice at delegating tasks he could not do as fast as others (reading evidence quickly in competitive debate) and a bias toward unconventional approaches to problems — a frame he says he now applies deliberately at Mercor, encouraging staff to build on comparative strengths rather than patch weaknesses.
See also
- Brendan Foody — guest
- Tyler Cowen — host
- Brendan Foody on Evals, the Expert Labour Market, and Mercor's Rise — Foody’s earlier Lenny’s Podcast appearance, more focused on Mercor’s own growth and internal culture
- Evals — extended here with APEX’s economic-value-weighted measurement methodology, an alternative to academic-benchmark evals
- Reinforcement Learning from Human Feedback — extended here with the Kant/taste objection to rubric-based training and the ‘society as RL machine’ labour-market mechanism
- Sam Altman on Trust, Persuasion, and the Future of Intelligence — the wiki’s other Cowen conversation on whether a rubric can reach a poem’s true aesthetic ceiling