The most capable agents we evaluated essentially saturated our Time Horizon 1.1 benchmark
Their measured horizon was "over two full-time-equivalent days", a figure METR said it was "increasingly uncertain about" as the suite saturates; only five of its 228 tasks are estimated to take a person longer than sixteen hours, which makes horizons in that range, in METR's words, "infeasible to precisely measure". In late June METR reported its attempt to measure OpenAI's GPT-5.6 Sol: three numbers, from about eleven hours to more than 270, none of which it was prepared to call robust. And on 29 September OpenAI began rolling out "dots", agents that its system card says run under "a new time-budget setting that guides how long they work".
Two questions sit underneath those events. What turns a language model that can call a tool into an agent? And what, exactly, is being measured when an agent is said to have a horizon of so many hours? The answers are older than the technology, and they explain why that number is so easily misread.
Two readings that mislead
The first misreading is that agents can now work unattended for two days. Three different clocks are being run together. METR's horizon is a property of tasks, not of running time: it is the length of task, measured by how long skilled people took to do it, at which an agent is predicted to succeed half the time. METR's own explanatory page is explicit:
It’s a measure of the difficulty of a task, rather than the time an AI spends to complete the task.
Agents, the same page notes, "are typically several times faster than humans on tasks they complete successfully". Thomas Kwa, first author of the paper that introduced the measure, put the point bluntly in a note of 22 January 2026: "Time horizon is not the length of time AIs can work independently." The second clock is how long an agent actually runs, which METR declines to report because it "varies greatly by inference provider and exact agent setup". The third is the time budget a product sets to guide how long an agent keeps working—dots' setting is one—and it is a different quantity from how hard a task the agent can finish.
The second misreading is that a doubling horizon means month-long work is around the corner. The number is a coin-flip point on a particular kind of task. METR's suite consists of "self-contained and well-specified" tasks, "primarily" in software engineering, machine learning and cybersecurity, each with an automatic success criterion; METR's page describes them as "much 'cleaner' than real economically valuable labor". Raise the bar from one success in two to four in five and the horizon shrinks sharply. For the best public agents in February and March 2026, METR's risk report put the 50% horizon at about twelve hours and the 80% horizon at about an hour and a half; across models, its paper finds 80% horizons "4-6x shorter". And the newest measurements sit where METR says horizon estimates from its current suite become unreliable. The paper's own extrapolation to month-long software tasks within five years is conditional—"If these results generalize to real-world software tasks"—and its authors say that possible changes in the trend and external-validity concerns account for most of their uncertainty.
Table view
| Measure | Value |
|---|---|
| 50% time horizon | ~12 hours |
| 80% time horizon | ~1.5 hours |
| How much shorter the 80% horizon runs | 4–6× |
What makes the loop
A model called once answers once. Even when it asks for a tool—a search, a calculation, a file—a program that runs the tool, returns the result and stops is following a path somebody wrote in advance. An agent differs in one specific respect: after each result comes back, the model reads it and chooses the next step, including the step of declaring the work finished. The working test is therefore a question about authorship. Was the next step written down beforehand, or chosen by the model after it saw what the last step produced?
The industry's own definitions converge on that test, though they use the word "workflow" in opposite senses. OpenAI's "A practical guide to building agents" (undated; its PDF metadata gives 7 April 2025) says that "Applications that integrate LLMs but don’t use them to control workflow execution—think simple chatbots, single-turn LLMs, or sentiment classifiers—are not agents", and describes the agent's "run" as "a loop that lets agents operate until an exit condition is reached". In OpenAI's usage the workflow is the job, which an agent runs; in Anthropic's essay "Building effective agents" of December 2024, a workflow is a "predefined code path", the opposite of an agent. The programmer Simon Willison, who collected 211 crowd-sourced definitions before settling on one, wrote on 18 September 2025:
An LLM agent runs tools in a loop to achieve a goal.
The words "to achieve a goal", he added, reflect "that these are not infinite loops—there is a stopping condition."
A single test applied step by step gives a sharp answer; applied to a whole system it gives a dial, because real systems mix fixed steps with chosen ones. OpenAI's researchers, in a December 2023 paper on governing such systems, likewise wrote that "there is no clear line along which to draw a binary distinction between “agents” and current AI systems like GPT-4", and preferred to speak of degrees of "agenticness". An older academic formulation offered a broader test of continuity. Stan Franklin and Art Graesser, proposing a taxonomy of software agents in 1996, excluded an ordinary payroll program because its output does not affect what it senses later: "It runs once and then goes into a coma, waiting to be called again." Their definition is wide enough to admit a thermostat—"A thermostat? Yes, a thermostat satisfies all the requirements of the definition"—so it marks the difference between a single call and a loop, not the narrower question of who chooses the next step.
Table view
| # | Stage | Note |
|---|---|---|
| 1 | A goal | set by a person, or by another agent |
| 2 | The model chooses the next step | reads the goal and every result so far; writes a request for an action, or gives a final answer |
| 3 | The harness carries it out | ordinary software runs the tool, or refuses; the model only asks |
| 4 | The result comes back | added to what the model will read before its next choice |
| 5 | Test: out of budget? | the harness's stop rule: a limit on turns, tokens or time; if none is reached, the model chooses again |
| 6 | Exit | an answer, a question for the user, or an error; the model saying it is finished is a claim, and in measurements such as METR's, success is judged outside the loop |
| From | To | Label |
|---|---|---|
| A goal | The model chooses the next step | |
| The model chooses the next step | The harness carries it out | a request |
| The harness carries it out | The result comes back | |
| The result comes back | Test: out of budget? | |
| Test: out of budget? | The model chooses the next step | not yet |
| Test: out of budget? | Exit | limit reached |
| The model chooses the next step | Exit | final answer |
An old idea: signals from the goal
The role of feedback in purposeful behaviour was set out in January 1943 by Arturo Rosenblueth, Norbert Wiener and Julian Bigelow, of Harvard Medical School and the Massachusetts Institute of Technology, in a short paper in Philosophy of Science titled "Behavior, Purpose and Teleology". They borrowed an engineers' word, "feed-back", for behaviour "controlled by the margin of error at which the object stands at a given time with reference to a relatively specific goal", and drew a distinction that maps closely onto the present subject. A snake, they observed, "may strike at a frog, or a frog at a fly, with no visual or other report from the prey after the movement has started"; by contrast "the behavior of some machines and some reactions of living organisms involve a continuous feed-back from the goal that modifies and guides the behaving object". Among machines they called intrinsically purposeful they named "A torpedo with a target-seeking mechanism". Their conclusion:
If a goal is to be attained, some signals from the goal are necessary at some time to direct the behavior.
Wiener named the wider field Cybernetics in a book of 1948, in which he dates the term to the summer of 1947 and credits the first significant paper on feedback mechanisms to James Clerk Maxwell's 1868 article on governors. The folk version, in which Wiener invented feedback in 1943, is wrong on both counts.
In 1960 three psychologists—George Miller, Eugene Galanter and Karl Pribram—made the loop the basic unit of behaviour in Plans and the Structure of Behavior. The old unit had been the reflex arc: a stimulus in, a response out, once. "The unit should be the feedback loop itself," they wrote, and called it the TOTE: Test, Operate, Test, Exit. Their worked example is hammering a nail. A plan that lists only lifting and striking "does not tell us, for one thing, how long to go on hammering"; it needs a test—is the head flush with the wood?—and the plan continues until the test is passed. They called the missing piece the "stop rule". The book cites Wiener and acknowledges material from Allen Newell, J. C. Shaw and Herbert Simon, whose General Problem Solver of 1958–59 worked by measuring the difference between what it had and what it wanted and choosing an operation to reduce it.
The mapping onto language models is direct, provided it is read as an analogy rather than a lineage: ReAct's authors cite the psychology of inner speech and earlier robotics work, not Wiener or cybernetics. A single call to a model is the snake's strike: nothing comes back. An agent receives the report, and, unlike a homing machine, chooses what to do next; feedback alone does not make a system an agent, since a thermostat has feedback and no discretion. Language models were put into such loops in the early 2020s. OpenAI's WebGPT (December 2021) issued browser commands and read back the pages it reached; Google's Inner Monologue work, posted in July 2022, fed descriptions of a robot's surroundings and of success or failure back into a model that planned its actions. In October a team from Princeton University and Google's Brain team published ReAct, in which a model interleaves written "thoughts" with actions such as a search and the observations they return; in its question-answering setup one of the permitted actions, finish, ends the task. In its decision-making tasks the thoughts appear only where the model chooses to write them. ReAct's distinctive contribution was the written reasoning between actions; its authors credit Inner Monologue as "the first work that demonstrates such a closed-loop system, which ReAct builds on".
The harness owns the exit
In an agent, the software around the model—the harness—does three jobs that a single model call does not need. It closes the loop: OpenAI's Agents SDK documentation describes the cycle as "If the LLM produces tool calls, we run those tool calls, append the results, and re-run the loop." It holds the limits the model does not set: the same page raises an error "If we exceed the max_turns passed". And it shapes what any measurement of the agent measures. METR's own description of its default scaffold, in a research note of 13 February 2026 by Nikola Jurkovic, is the loop in one sentence: "ReAct is a very simple scaffold where an agent takes an action, sees the results of the action, and repeats." In that note, two models measured inside commercial coding harnesses—Anthropic's Claude Code and OpenAI's Codex—did not do measurably better than inside METR's plainer scaffolds; the comparison covered two models, and METR's notes carry the caution that they "do not necessarily reflect the views of METR as a whole". (Claude Code is a product of Anthropic, whose Claude model is used to produce this publication.)
Dots show the same machinery built for duration. OpenAI's system card describes each dot as having "its own cloud computer and browser", delegating to subagents, and adds: "Multi-agent setups and persistence are not new, but dots use a new time-budget setting that guides how long they work." The year-long budgets that appear in the same appendix belong to OpenAI's safety evaluations, in which the model was given "a generous time budget of up to one year, and a clock tool with wait functionality"; OpenAI notes that at the one-year setting the model "spends less simulated time on tasks". A time-budget setting guides how long a loop keeps working. It is a different quantity from how hard a task the loop can complete.
Measuring the loop in human hours
METR introduced the time horizon in March 2025 in "Measuring AI Ability to Complete Long Tasks" (Thomas Kwa, Ben West and colleagues; arXiv 2503.14499, retitled "...Long Software Tasks" in later versions). The motivation was that benchmark scores "saturate increasingly quickly" and say little about real-world ability; the remedy was to express ability in the one unit every task already has, the time a skilled person needs. The method has four steps. METR times professionals—on average about five years' experience in software, machine learning or cybersecurity—on most tasks and takes the geometric mean of their successful times; where no reliable timing exists, as for most of the longest tasks, it uses expert estimates. Its usual protocol runs each agent on every task several times—six independent runs per task, as its page describes it today—inside a scaffold with token and time limits. It fits a logistic curve of the agent's probability of success against the logarithm of human time. And it reads off the task length at which the curve crosses 50%, or 80%.
The paper says the method is "inspired by" item response theory, the family of statistical models used to score human tests, in which the chance that a person answers a question depends on the person's ability and the question's difficulty. As in that theory, the fit finds the difficulty at which the test-taker succeeds half the time; unlike it, METR takes difficulty from human completion time rather than estimating it from the test-takers' answers. The paper's references for the idea are a textbook, a handbook and a paper applying the theory to machine-learning classifiers, not the theory's founders.
Table view
| # | Stage | Note |
|---|---|---|
| 1 | Estimate how long each task takes a skilled person | 228 tasks in the 2026 suite, from about a second to about 30 hours; geometric means of successful human times where measured; some durations, including most long ones, are expert estimates |
| 2 | Run the agent on every task | in the current protocol, six independent runs per task, inside a scaffold with a token and time limit |
| 3 | Fit a curve | probability of success against the logarithm of human time, a logistic curve as in item response theory |
| 4 | Read off the crossing | the task length at which the fitted curve crosses 50% (or 80%) is the time horizon |
| 5 | Plot against release date | the slope of the line through successive frontier models gives the doubling time |
| From | To | Label |
|---|---|---|
| Estimate how long each task takes a skilled person | Run the agent on every task | |
| Run the agent on every task | Fit a curve | |
| Fit a curve | Read off the crossing | |
| Read off the crossing | Plot against release date |
The horizon is a fitted point, not an observation. METR's FAQ illustrates what a two-hour horizon means in practice: on tasks taking people 90 minutes to three hours, a GPT-5 agent "succeeds 100% of the time for around one-third of the tasks, fails 100% of the time for around one-third of the tasks, and sometimes succeeds and sometimes fails on the remaining third". The paper adds that errors in individual models' horizons are strongly correlated, because sampling easier or harder tasks moves every model together; hence its authors' statement that they are "more confident in the slope of the time horizon trend than in the time horizon of any particular model".
Table view
| Item | Value |
|---|---|
| GPT-2 (Feb 2019) | 0.1 min |
| GPT-3, davinci-002 proxy (May 2020) | 0.1 min |
| GPT-3.5, turbo-instruct proxy (Mar 2022) | 0.6 min |
| GPT-4 (Mar 2023) | 4 min |
| GPT-4 1106 (Nov 2023) | 4 min |
| GPT-4o (May 2024) | 7 min |
| Claude 3.5 Sonnet (Jun 2024) | 11.4 min |
| o1-preview (Sep 2024) | 20.3 min |
| Claude 3.5 Sonnet, new (Oct 2024) | 20.5 min |
| o1 (Dec 2024) | 38.8 min |
| Claude 3.7 Sonnet (Feb 2025) | 60.4 min |
| o3 (Apr 2025) | 119.7 min |
| GPT-5 (Aug 2025) | 203 min |
| Gemini 3 Pro (Nov 2025) | 224.3 min |
| Claude Opus 4.5 (Nov 2025) | 293 min |
| GPT-5.2 (Dec 2025) | 352.2 min |
| Claude Opus 4.6 (Feb 2026) | 718.8 min |
| Claude Mythos Preview, early (Apr 2026) | 1,044.8 min |
The record so far
The first version of the paper, in March 2025, put "current frontier AI models such as Claude 3.7 Sonnet" (an Anthropic model) at "around 50 minutes", with the frontier horizon doubling every 212 days (95% interval 171–249) from 2019—"though the trend may have accelerated in 2024". Its July 2026 revision gives GPT-2 a horizon of two seconds and OpenAI's o3, released in April 2025, about 110 minutes, and a doubling time of 207 days (166–240). In January 2026 METR rebuilt the yardstick. Time Horizon 1.1 grew the suite from 170 to 228 tasks and the number of tasks of eight hours or more from 14 to 31, of which only five had measured human times. Measured with the new tasks, the post-2023 doubling time fell from 165 days to 131, and the post-2024 figure to 89; METR's explanation was that "it’s likely the new tasks are drawn from a slightly different distribution of difficulty". Part of the apparent acceleration, in other words, is a change of ruler.
The May risk report, which drew on non-public information and model access from Anthropic, Google, Meta and OpenAI, fitted only frontier models released after 1 January 2024, a date it chose as a break-point "between an older, slower trend and the current trend". That fit gives a doubling time of 105 days, about three and a half months, with an R² of 0.98. An independent measurer, using different tasks, reached a similar conclusion. Britain's AI Security Institute reported on 13 May that "The length of tasks frontier models can autonomously complete in our narrow cyber suite has been doubling every few months. This doubling rate has become faster over time, and recent models exceeded our previous trends"; its estimate of the 80%-reliability cyber horizon's doubling time had moved from eight months (November 2025) to 4.7 months (February 2026), and it said that two newer models, Anthropic's Claude Mythos Preview and OpenAI's GPT-5.5, "substantially exceeded both doubling rate trends", adding: "It is unclear whether this represents a new, faster trend." AISI's own post notes that its latest estimates "are close to those produced by METR".
Table view
| Item | Value |
|---|---|
| First paper, 2019 to early 2025 (March 2025) | 212 days |
| Original tasks, models since 2023 | 165 days |
| New tasks (January 2026), models since 2023 | 131 days |
| New tasks, models since 2024 | 89 days |
| Risk report (May 2026), models since 2024 | 105 days |
Then the ceiling. On 8 May METR added to its chart page the notice that "Measurements above 16 hrs are unreliable with our current task suite". The risk report of 19 May described the strongest agents as having left "only a handful of tasks longer than eight hours that they were still unable to solve", many of the failures "due to cheating rather than obvious inability". On 26 June, summarising its pre-release evaluation of GPT-5.6 Sol, METR reported that the result depended on how such attempts were scored: about 11.3 hours (95% interval 5–40) if they counted as failures, as its standard method requires; more than 270 hours if they counted as successes; and 71 hours (interval 13 to 11,400) if they were discarded, which removed the data for several informative long tasks. Its verdict: "we do not consider any of these numbers to represent a robust measurement of GPT-5.6 Sol’s capabilities". OpenAI's legal and communications staff reviewed and approved the post, which says so.
Table view
| Item | Value |
|---|---|
| Rule-breaking runs scored as failures (standard method) | 11.3 h |
| Rule-breaking runs discarded | 71 h |
| Rule-breaking runs scored as successes (lower bound) | 270 h |
Since 8 May no new time horizon for a publicly released model has appeared on METR's chart, and METR's summary of its pre-release evaluation of Anthropic's Claude Opus 5.5, published on 22 September, reports none; its capability testing used five other tasks. Anthropic, whose Claude is used to produce this publication, could review and edit that summary, which says so.
| Date | What happened | Source |
|---|---|---|
| January 1943 | Rosenblueth, Wiener and Bigelow: purposeful behaviour needs "signals from the goal" | Philosophy of Science 10(1) |
| 1960 | Miller, Galanter and Pribram: the TOTE unit and the "stop rule" | Plans and the Structure of Behavior |
| 1996 | Franklin and Graesser: a program that "runs once" is not an agent | ATAL-96 workshop |
| 12 July 2022 | Inner Monologue: language feedback in a robot's planning loop | arXiv 2207.05608 |
| 6 October 2022 | ReAct: thoughts, actions and observations in a loop | arXiv 2210.03629 |
| December 2023 | OpenAI: "no clear line" between agents and other AI systems | OpenAI paper |
| 18 March 2025 | METR introduces the 50% time horizon; doubling every 212 days | arXiv 2503.14499 |
| Undated (PDF metadata: 7 April 2025) | OpenAI's guide: an agent's "run" loops until an exit condition | OpenAI |
| 18 September 2025 | Willison's working definition: tools, in a loop, toward a goal | simonwillison.net |
| 29 January 2026 | Time Horizon 1.1: more and longer tasks; faster estimated pace | metr.org |
| 13 February 2026 | METR note: commercial harnesses no better than its plain scaffolds, for two models | metr.org |
| 8 May 2026 | METR: measurements above 16 hours "unreliable" | metr.org |
| 13 May 2026 | UK AI Security Institute: cyber horizons doubling faster | aisi.gov.uk |
| 19 May 2026 | METR risk report: suite "essentially saturated"; 105-day post-2024 fit | metr.org |
| 26 June 2026 | METR: no robust horizon for GPT-5.6 Sol | metr.org |
| 2 July 2026 | BRIDGE (revised): about six months by another method | arXiv 2602.07267 |
| 29 September 2026 | Dots: "a new time-budget setting that guides how long they work" | OpenAI system card |
Two readings of the trend
METR's position is that the slope is the robust quantity and the yardstick is the weak point. In January it wrote that even the new suite "has relatively few tasks that the latest generation of models cannot perform successfully", and that "We are prioritizing work on updates to our evaluations so they can measure the capabilities of very strong models." Its page states that it has seen no evidence of the exponential growth slowing, while warning that a logistic curve fitted to the early part of a trend can yield "wildly different asymptotes". The British institute, whose cyber results point the same way, is careful about the status of any such curve:
This is an imperfect model, and is not a future prediction, nor a fixed law.
The second reading questions the pace or the shape. BRIDGE, a method from researchers at Mila, McGill University and ServiceNow (first posted in February, revised in July and accepted at ICML 2026), estimates task difficulty from how many models perform across benchmarks, anchors it to METR's human timings, and "independently reproduce[s] METR’s exponential scaling results, with the 50% solvable task horizon doubling approximately every 6 months". Its authors present that as corroboration; a doubling time of six months is, on this publication's reading, closer to METR's long-run estimate than to its post-2024 one. A sharper dissent came from Haosen Ge, Hamsa Bastani and Osbert Bastani in February: "we argue that the data does not support exponential growth, even in shorter-term horizons", on the grounds that an S-shaped curve fits METR's data with its turning point already passed. They describe their aim as "to highlight the fragility of existing forecasts of exponential growth" rather than to offer a forecast of their own; their paper predates the two highest points on METR's chart.
The studies use different methods, periods and data, which accounts for part of the difference between them, and the disagreement is sharpest where the data is thinnest. New, longer tasks with measured human times, and today's agents run against them, would test the newest estimates and whether the recent pace holds; METR's page has added no measurement since May. Until then the summary that fits the evidence is METR's own: confidence in the slope, and growing uncertainty about the newest points.
What to watch
The first is the chart itself. METR's time-horizon page still read "LAST UPDATED May 8, 2026" on 7 October. A new date, a new task suite with tasks well beyond 30 hours, or a published horizon for any model released after April 2026 would each show whether the yardstick has been extended. METR's data file currently omits points above 16 hours from its fitted trend; a change to that note would be a sign of new tasks.
The second is METR's next risk report. The May report says: "We tentatively plan to run a similar process in late 2026." A second edition on metr.org by the end of December could bring new estimates of how far the companies' internal frontier runs ahead of the public one—estimates resting on non-public information—and a refreshed public-model trend would show whether the 105-day estimate still holds.
The third is the British institute's next estimate. Its May post says it will "continue to evaluate frontier autonomous cyber and software capabilities, and to update our estimates as the evidence develops". A new doubling time, faster or slower than 4.7 months, would be a published check on METR's direction from an evaluator using its own tasks.
The idea to keep
An agent is a model whose next step depends on what its last step produced, with a rule for stopping. The rule can belong to the model, which may decide it has finished, or to the harness, which counts turns, tokens and time. That is the whole of the difference between an agent and a model that answers once. Its nearest ancestor, by analogy, is the distinction Rosenblueth, Wiener and Bigelow drew in 1943 between the snake's uncorrected strike and behaviour guided by signals from the goal—with the addition that the model also chooses what to do next. How far such a loop can reach is measured, on METR's yardstick, in the human working time of the tasks it completes half the time: not how long it runs, not how long it is allowed to run, and not a promise about work that is messier than a well-specified technical task. A claim that an agent "works for hours" therefore invites three questions—whose hours, on which tasks, and at what odds of success—and in October 2026 the yardstick that answers them is waiting for longer tasks.
Sources
| Source | Date |
|---|---|
| Arturo Rosenblueth, Norbert Wiener and Julian Bigelow, Behavior, Purpose and Teleology, Philosophy of Science 10(1), pp. 18–24 | January 1943 |
| Norbert Wiener, Cybernetics, or Control and Communication in the Animal and the Machine (Technology Press, Wiley) | 1948 |
| Allen Newell, J. C. Shaw and Herbert A. Simon, Report on a General Problem-Solving Program, RAND P-1584 | December 1958, revised February 1959 |
| George A. Miller, Eugene Galanter and Karl H. Pribram, Plans and the Structure of Behavior (Henry Holt) | 1960 |
| Stan Franklin and Art Graesser, Is it an Agent, or just a Program?: A Taxonomy for Autonomous Agents, ATAL-96 | 1996 |
| Wenlong Huang and colleagues (Google), Inner Monologue, arXiv 2207.05608 | 12 July 2022 |
| Shunyu Yao and colleagues (Princeton University, Google Research), ReAct: Synergizing Reasoning and Acting in Language Models, arXiv 2210.03629 (ICLR 2023) | 6 October 2022 |
| Yonadav Shavit and colleagues (OpenAI), Practices for Governing Agentic AI Systems | December 2023 (undated; PDF metadata 18 December 2023) |
| Anthropic, Building effective agents | 19 December 2024 |
| Thomas Kwa, Ben West and colleagues (METR), Measuring AI Ability to Complete Long Tasks, arXiv 2503.14499 v1 (v4, 10 July 2026: …Long Software Tasks) | 18 March 2025 |
| OpenAI, A practical guide to building agents | undated; PDF metadata 7 April 2025 |
| Simon Willison, I think "agent" may finally have a widely enough agreed upon definition to be useful jargon now | 18 September 2025 |
| Thomas Kwa (METR), research note on the limitations of time horizon | 22 January 2026 |
| METR, Time Horizon 1.1 | 29 January 2026 |
| Haosen Ge, Hamsa Bastani and Osbert Bastani, Are AI Capabilities Increasing Exponentially? A Competing Hypothesis, arXiv 2602.04836 | 4 February 2026 |
| Nikola Jurkovic (METR), Measuring Time Horizon using Claude Code and Codex | 13 February 2026 |
| METR, Task-Completion Time Horizons of Frontier AI Models (live page and data files) | last updated 8 May 2026; read 7 October 2026 |
| UK AI Security Institute, How fast is autonomous AI cyber capability advancing? | 13 May 2026 |
| METR, Frontier Risk Report (February to March 2026) | 19 May 2026 |
| METR, Summary of METR's predeployment evaluation of GPT-5.6 Sol | 26 June 2026 |
| Fengyuan Liu, Jay Gala and colleagues (Mila, McGill University and others), BRIDGE: Predicting Human Task Completion Time From Model Performance, arXiv 2602.07267 v2 | 2 July 2026 |
| METR, Summary of METR's predeployment evaluation of Claude Opus 5.5 | 22 September 2026 |
| OpenAI, GPT-6 Astra System Card, section 12 ("Appendix: dots"), and Introducing dots | 29 September 2026 |
| OpenAI, Agents SDK documentation, Running agents | read 7 October 2026 |