Silicon Valley Turned an Algorithm Into a Management Philosophy. The Algorithm Has a Known Failure Mode

Executive at a glass-walled San Francisco conference room whiteboard drawing an ascending performance curve that flattens at the top, city hills visible through the window

“Hill-climbing” is moving from AI labs to consulting calls to boardrooms. It names a search that improves whatever number sits in front of it and cannot see anything else.

Hill-climbing is a local search method from early AI research. You score your current position, test small nearby changes, keep whichever one scores higher, repeat. Marvin Minsky put the term into the AI canon in 1961 and warned in the same paper that the method stalls on false peaks and wanders without learning across flat ground. Silicon Valley has adopted the word and left the warning behind. For anyone running a company, the vocabulary carries a cost: a shared language that only names incremental improvement makes the exploration budget hard to defend in a meeting.

Satya Nadella used it in June to describe running many models inside infrastructure you control. Demis Hassabis used it in July to ask whether the AI industry has parked itself on a peak. Paul Graham uses it to tell founders to chase growth now and worry about ceilings later. Three men, three different jobs, one metaphor, and none of them are talking about hiking.

What Happened

The Wall Street Journal reported on Thursday that “hill-climbing” has become the phrase of the moment across San Francisco, spreading out of machine-learning research into ordinary business talk. Google Cloud’s Thomas Kurian and Palantir’s Shyam Sankar have both reached for it in public remarks. Engineers ask each other what they are hill-climbing this week. One AI executive calls his quantitative hires hill-climbers. A researcher asks his 12-year-old daughter whether she wants to climb the geometry hill tonight.

Nick Heiner, who runs reinforcement learning environments at the data company Surge AI, gave the Journal the flattest possible definition: “It’s just a universal term for the act of improving.”

That flattening is why the word travels so well, and it strips out the part that carries the meaning. Hill-climbing describes improvement along the axis you are already standing on. The axis is the whole idea.

Hanlin Tang, CTO of neural networks at Databricks, told the Journal he has heard non-technology clients use the term on calls, mostly the AI-forward consulting accounts. That is the transmission mechanism worth watching. Consultants carried “flywheel” into the boardroom the same way, and an Indiana University linguist quoted in the piece drew the same comparison.

The Backstory

Minsky described hill-climbing in “Steps Toward Artificial Intelligence,” published in the Proceedings of the IRE in 1961. You have a black box with inputs and a score. You cannot see inside it. You probe locally, find the direction of steepest ascent, move, and repeat until improvement stops.

He also wrote the caveats, in the same paper, in the same section.

The first caveat is the famous one. A climber that reaches a peak with no higher neighbor stops, whether or not that peak is any good. Minsky said the fix is to force larger steps.

The second caveat is the one nobody quotes, and it is the one that should worry operators. Minsky argued that for hard problems, getting trapped on a false peak is not the main obstacle. The main obstacle is finding any meaningful peak at all. Real score functions tend to produce what he called the Mesa Phenomenon: broad flat regions where a small change in a parameter produces no change in the score. On a mesa, a small-step climber wanders without learning anything.

His summary line is blunt. Hill-climbing requires structural knowledge of the search space, and without it, “hill-climbing may do more harm than good.”

Sixty-five years later, that sentence is the part that did not make it into the buzzword.

The Plan

Somebody noticed that a hill-climber needs a hill, and built a business.

Reinforcement learning runs on environments: simulated workplaces where an agent attempts a task and a grader scores the attempt. Mercor bought Sepal AI in February 2026 for expert-graded benchmarks, then bought Deeptune in July 2026 for enterprise environments that replicate Excel, Salesforce, Slack and Zendesk as sandboxes. Surge AI, bootstrapped to roughly $1.2 billion of 2024 revenue on about 110 staff and now valued by Forbes near $24 billion, ships its own environment suite. Prime Intellect raised $130 million at a $1 billion valuation in July 2026 on roughly $100 million of annualized revenue, pitching enterprises on owning their own optimization loop. Mechanize builds coding environments for Anthropic and pays engineers $500,000.

Across the wider training-data layer, press estimates put 28 startups at roughly $8.5 billion of combined annual revenue against something close to $100 billion of combined valuation, with Scale, Surge, Mercor and Handshake taking more than three quarters of the revenue.

Then look at what the environments line itself earns.

Bar chart comparing 2026 annualized revenue run-rates for Mercor at $2.0 billion, Handshake AI at $1.1 billion and Prime Intellect at $100 million, against Mercor's reinforcement learning environments unit target of tens of millions for September 2026

Mercor crossed $2 billion of annualized revenue in June 2026 and guides its RL environments business to tens of millions by September. Two acquisitions in five months, and the unit lands somewhere near 1% to 2% of group revenue. The scarce asset everybody names is, for now, a line item.

The Business Model Angle

James March named this problem for management in 1991, thirty years before anyone put it on a slide. Writing in Organization Science, he split organizational learning into exploration, meaning search and experimentation with uncertain returns, and exploitation, meaning refinement of what already works. His finding: adaptive organizations improve exploitation faster than exploration, which makes them effective in the short run and self-destructive over a longer one. Exploitation wins the budget argument because its payoffs arrive sooner, land closer to home, and come with smaller error bars.

Hill-climbing is exploitation with an equation attached. Importing it as management vocabulary hands the side that already wins a better sales pitch.

Three consequences follow for an operating company.

Exploration loses its name. An ambitious executive now has crisp language for “make the number go up” and no equivalent shorthand for “spend money to find out whether we are on the right number.” Nobody votes to cut the exploration budget. It shrinks because the person defending it has to explain the concept first while the person defending optimization gets a one-word noun.

Most corporate metrics behave like mesas. Brand equity, retention among an untargeted segment, developer satisfaction, the health of a channel you undermonitor: you can push hard on these for two quarters and read a flat line, which tells your team nothing about direction. Minsky’s warning applies with more force in a company than in a training run, because a business cannot rerun the quarter with a larger step size.

The hill is the deliverable. Enterprises buying agents will spend more defining the grader than buying the model, which is the argument we ran in AI is winning where the answer can be checked. The environments market says the same thing from the supply side, and its current revenue says the buyers have not fully arrived.

Marc Benioff already made this move in public. Salesforce swapped token consumption for “agentic work units” as its internal measure of AI success, a shift we covered in big tech’s token rationing reversal. Changing what you count picks a new hill. Hill-climbing has no word for that decision, which is why the decision keeps getting made by default.

The Risk

The capex arithmetic is where the metaphor gets expensive. The four largest hyperscalers have guided toward roughly $700 billion of combined 2026 capital spending, with source figures spread across a $660 billion to $725 billion range depending on which companies and which lease treatment you count; Morgan Stanley’s five-company estimate runs closer to $805 billion. That money is underwritten on continued hill-climbing, meaning that more compute and more RL keep producing gains along the current axis.

Hassabis, who runs one of the labs, asked in July whether climbing gets the industry the rest of the way or whether one or two more breakthroughs are required. A breakthrough is a jump across a valley. Hill-climbing algorithms cannot make jumps, which is why computer scientists invented random restarts and simulated annealing to work around them. Anyone lending against a 2026 data center is financing a climb and hoping for a leap.

The nearer risk sits inside ordinary companies. Roughly 90% of firms using AI report no productivity payoff yet, which we covered in why the AI productivity boom keeps not arriving. Teams that treat every problem as a hill will keep optimizing pilots that were never on the right slope. Our reporting on the chief of staff boom showed how that plays out: companies optimized against span of control because layers were countable, the proxy improved, and the cost reappeared somewhere the metric could not see.

Now the counterargument, because Graham’s version deserves better than a strawman. Founders who agonize about local maxima before they have customers tend to freeze, and early growth reveals options that no amount of upfront analysis would have surfaced. Climbing to the top of a small hill is often the fastest way to see the bigger one. Minsky’s own advice was to take larger steps, not to stop climbing. YC’s returns come from outliers, though, and outliers change function rather than parameter. Telling every founder to optimize locally is defensible advice for the median company in the batch and poor advice for the ones the portfolio needs.

One more wrinkle. Managers who reach for the term may sharpen their thinking rather than dull it, because “what are we hill-climbing” forces a team to name a metric, and most teams cannot. A word that exposes the absence of a measurable objective earns its keep.

Quick Questions

What does hill-climbing mean in AI? An optimization method that evaluates small changes to a current solution and keeps the ones that raise a score. Modern model training uses gradient-based and reinforcement learning methods that share the same logic: measure, adjust, measure again.

Why do people mention local maxima with it? A hill-climber stops at any point where no nearby move scores higher, even when a taller peak sits elsewhere. Getting stuck on a lesser peak is the textbook failure of the method.

Is hill-climbing the same as continuous improvement? Close enough for conversation, with one difference that matters. Continuous improvement programs like kaizen usually let workers question the process itself. Hill-climbing holds the objective fixed by definition.

What does this mean for my business? Split the budget explicitly. Fund the work that raises this quarter’s metric and fund the work that tests whether the metric is the right one, and hold the second line separately so it cannot lose an argument to the first.

Who sells the hills? Mercor, Surge AI, Scale, Handshake, Prime Intellect and Mechanize supply the environments and graded data that make reinforcement learning possible. It is a fast-growing layer, and the environments-specific revenue is still small against the human-data revenue underneath it.

The Business Model Analyst Take

Jargon reveals what a business finds easy. “Flywheel” spread because it flattered companies whose growth compounded on its own. “Hill-climbing” is spreading now because the AI industry has spent three years in a regime where more compute and better graders produced a better score, and that experience is contagious to executives who would like the same certainty.

Buy the word if you like. Buy the caveat with it. The people using it correctly are optimizing inside a system where somebody else already chose the objective function, and the labs pay Mercor and Surge real money to build those objectives, because defining a good one turns out to be the hard part of the work.

Inside your company, nobody sells you that service. You choose your own score function, you seldom revisit it, and the metric someone wrote on a whiteboard at a planning offsite three years ago now decides which hill your team is allowed to climb. That choice sits in no line of your budget. It sets your ceiling anyway, above everything the optimization work underneath it can reach.

Minsky’s advice still holds. When improvement stops, take a bigger step.

UNLOCK THIS FREE DOWNLOAD

DOWNLOAD NOW

Fill Your E-mail to Receive this Download Directly in Your Inbox.

RECEIVE OUR UPDATES

The Biz Model Club

Get daily, no-fluff insights on the latest business models, startup strategies, and trends delivered straight to your inbox.