AI Perspectives

Seventy Years On, AI Still Evaluates Itself the Way Dartmouth Did

Correction, 5 September 2026. This article has been rewritten. The previous version dated the Dartmouth proposal to the summer of 1956; it was written on 31 August 1955. It credited Marvin Minsky with establishing the MIT Media Lab, founded in 1985 by Nicholas Negroponte and Jerome Wiesner, when Minsky co-founded the MIT AI Lab in 1959. A closing section of predictions written in 2023 has been removed.


Dartmouth College marks the seventieth anniversary of its Summer Research Project on Artificial Intelligence this autumn, with a conference in Hanover in October. The workshop ran for eight weeks in 1956 and produced no collective report and no agreed research programme. McCarthy missed the deadline for the summary and the papers he meant to place in that September’s IRE symposium proceedings, though Newell and Simon’s Logic Theorist paper and one by Rochester and colleagues did appear there. The name and the claim were both a year old by the time the workshop opened.

Both were written on 31 August 1955, in a funding request to the Rockefeller Foundation. Seventy years later we still measure AI systems by what they can be shown to do, and almost never by what happens when somebody works to break them.

Nothing in the request says how anyone would know whether it had worked.

Notes from that summer describe a machine that was evaluated by deliberately disturbing it, and a warning that one badly chosen example could derail a brittle reasoning system for good. Decades afterwards, one of the participants named a third problem he remembered from the room, which was researchers choosing the problems that made their techniques look good to whoever was paying. Anyone doing security work now would recognise all three on sight.

Two failure modes recorded in 1956, and one bad incentive identified in hindsight by somebody who had been there. Neither the proposal nor any evaluation method the field went on to build mentions any of them.

A grant application, August 1955

McCarthy first approached the Rockefeller Foundation in February 1955. He met Robert Morison, the foundation’s director of biological and medical research, that June with Claude Shannon, to work out what kind of proposal Rockefeller would fund. The document they eventually produced is dated 31 August 1955 and went to Rockefeller on 2 September, signed by McCarthy, Marvin Minsky, Nathaniel Rochester of IBM and Shannon. The original typescript ran to seventeen pages plus a title page, and copies are held at Dartmouth and Stanford. Stanford’s scan is the whole thing including the mailing list; AI Magazine reprinted the proposal itself in 2006.

They asked for two months, ten people, and $13,500, itemised in the proposal’s own budget table. The study would proceed, the first paragraph says, on the conjecture that “every aspect of learning or any other feature of intelligence” could be described precisely enough for a machine to simulate it. Seven problems follow: getting machines to use language, forming abstractions and concepts, neuron nets, self-improvement, randomness and creativity, and two others. The four proposers thought significant progress on one or more of them was available to a well-chosen group over a single summer.

The neuron-net section reports partial results. The proposers’ biographies describe machines already built, among them Minsky’s simulator of learning in nerve nets. Partial results on nerve nets and the machines already built do not amount to every aspect of learning or any other feature of intelligence, and the document never tries to close the distance. Morison halved it. He offered $7,500 for five weeks and wrote back that the whole area of mathematical models for thought was still hard to grasp clearly. He called what he was doing a modest gamble and declined to risk a substantial amount at that stage. The Rockefeller Foundation still calls the project five weeks long, which is the length it paid for. The man holding the cheque was the only person in the exchange who priced the claim against the evidence.

Machines now use language, form usable concepts, and solve problems that were reserved for humans in 1955. Against the seven problems the proposal named, the research agenda has held up better than most forecasts from that decade. The universal conjecture underneath it has not been established and probably cannot be. My argument is with how the document is built. It claims a capability, defers the evidence, and says nothing at all about how the thing would fail. A model launch in 2026 follows the same sequence: capability numbers at the top, a safety section near the bottom, and nothing at all quantifying what the system does under attack.

The name was chosen to escape cybernetics

McCarthy coined “artificial intelligence” for this proposal. He was explicit later about why. Cybernetics was the obvious existing label. He rejected it for two reasons. Its concentration on analog feedback seemed misguided to him. And he did not want to “accept Norbert Wiener as a guru or having to argue with him.” Ronald Kline worked from unpublished archives for his history of the split between the three fields, in IEEE Annals of the History of Computing 33(4), October–December 2011.

Wiener had defined cybernetics as the science of control and communication in the animal and the machine. Control is the operative word. The cyberneticians asked what a system does when something disturbs it, which is a question about failure and recovery. McCarthy’s field asked what a machine could be made to do, which is a question about capability.

Both are legitimate questions. Only the first one leads anybody to ask what happens when the disturbance is somebody’s doing.

Through that summer most of the participants still described their work as cybernetics, automata theory, or complex information processing. Allen Newell and Herbert Simon disliked “artificial intelligence” and preferred the third of those. The label won anyway. Departments took the name and so did the funding lines. We now need four separate words for secure, safe, responsible and trustworthy AI, and we need them because the name McCarthy chose describes what a system can do and says nothing about whether it can be trusted to do it.

What happened in Hanover

The workshop opened around 18 June 1956 and ran to about 17 August, roughly eight weeks. Sessions were held on the top floor of the mathematics building. Grace Solomonoff built her account of the summer from her husband Ray’s handwritten notes and Trenchard More’s. It is the closest thing to a daily record that survives, and what it shows is a much smaller event than the one in the textbooks.

Three people attended throughout: McCarthy, Minsky and Ray Solomonoff. Donald MacKay and John Holland, both on the list McCarthy sent Morison in May, never came. Rochester was too busy with the IBM 704 and sent More in his place for three weeks. Newell and Simon were down for the first fortnight. On the days More took attendance, between three and eight people were in the room. During the whole of the fifth week there were four.

McCarthy then lost his list of who had actually been there. Ray Solomonoff’s notes give twenty names and More’s list gives thirty-two, ordered by rough degree of involvement. A list McCarthy circulated afterwards runs to forty-seven, mixing participants with visitors and people who had merely expressed interest. McCarthy also proposed on 7 August that the group work jointly on chess. Minsky was against it, Shannon doubted its relevance, Rochester would go along if nothing better turned up. No group project came out of the summer.

Ray Solomonoff’s own summary in his notebook is four lines long and blunt. He rated the project not very suggestive. The things of value he lists are getting his report written and reproduced, meeting people in the field, and forming a view of how poor most of the thinking in it was.

Seventy years of retelling has left that verdict out of the standard history. So has almost everything else in the notes.

Ashby’s homeostat, 23 July 1956

W. Ross Ashby was a British psychiatrist, and he worked in cybernetics, which is the tradition McCarthy had picked a new name to avoid. He spoke on the Monday afternoon of 23 July, and both More and Solomonoff took notes.

He described his homeostat, which he had built in 1948. Four units, each with a magnetic needle. When the system was stable the needles stayed centred. When it went unstable a needle would swing out and hit a stop, and the machine would then switch one or more of its uniselectors, rewire its own input connections, and search the new configuration space until it found stability again. You tested it by disturbing it. That was the only way to learn anything about it, because a homeostat sitting undisturbed in a stable state tells you nothing at all.

Ashby made a second point that summer which I would put in front of anyone writing an AI security assessment in 2026. A simple machine judged psychologically can look extraordinary, and it looks that way most strongly when the observer cannot see part of its mechanism. Concealment inflates apparent capability. An evaluator who cannot see the mechanism will overrate the machine. That is the security argument for interpretability, and Ashby made it in a Dartmouth classroom in 1956. Some of what we now call emergent behaviour may be the same effect, though Ashby on his own does not establish that.

More and Solomonoff both recorded the objections. The homeostat kept no memory of previous solutions, so having learned to handle one disturbance and then a second, it took as long as ever to handle the first again. Its solution time grew exponentially with the number of units. Both were correct, and both concerned the machine’s capability. The surviving notes do not record anyone saying that the method of evaluation was the interesting part.

Solomonoff wrote down single-example brittleness

Most of the participants were pursuing deductive, logic-based systems. Solomonoff was the only one working entirely on induction with probabilistic measures. McCarthy had told him the search space would explode.

Solomonoff’s notes record his objection. A machine built on non-probabilistic methods, he wrote, could “be seriously disturbed forever by one ill-conceived example,” where a probabilistic machine would not be much disturbed by a single counterexample.

That is a statement about worst-case input, written in August 1956. It is not an adversarial example in the modern sense, and I do not want to dress it as one. Solomonoff was describing brittleness under a natural distribution and arguing for probabilistic induction. Nobody in that argument imagined a person choosing the input on purpose. But the failure mode is the one attackers later learned to trigger deliberately: a single input producing a persistent failure that no volume of testing on representative data would surface. Treating classification as a game against an adversary who manipulates inputs arrives formally with Dalvi and colleagues in 2004, and the adversarial attacks literature most people mean starts with Szegedy and colleagues in 2013.

He made the argument to win a methodological dispute about induction, and he lost it. Deductive, logic-based work set the agenda for the next two decades, and probability came back into the mainstream of AI only in the 1980s.

The demo-to-sponsor problem, named fifty years later

Solomonoff wanted the group to work on many easy problems at once, looking for what they had in common. Most of the others wanted hard problems in narrow domains. He gave the disagreement a name much later, in a note Grace Solomonoff dates to 2005, prompted by the film-maker Wendy Conquest’s recollection of where he had stood. Researchers choose the problems that will show their technique working, because a sponsor is watching. He called it the demo-to-sponsor problem.

The label is retrospective, and the split it describes was in the room in 1956.

Dartmouth had a demo-to-sponsor problem of its own. Rockefeller had funded about half the request, and Morison had put his hesitancy in writing.

I have written before about AI security theatre, meaning products and claims aimed at threats the vendor has described but not demonstrated. It is the same incentive. Choose the evaluation that shows the technique working. Report the result. Do not construct the case that would embarrass it. In 1956 that incentive decided which problems anyone worked on. In 2026 it decides which numbers appear in launch posts. An attack success rate quoted without the model, the access level, the query budget and the defence in place is not a measurement, and it never was.

The program that was and was not running

The best-known thing about Dartmouth is that Newell and Simon turned up with the Logic Theorist while everyone else had proposals. The outline is right and two of the details are not.

Newell and Simon were down for the first fortnight only, and when they presented, the Logic Theorist had never executed on a computer. They had run parts of it by hand in January 1956, giving index cards to Simon’s wife, his three children and some graduate students so that each person acted as a component of the program. The RAND report of that July describes a specified, hand-simulable system with machine realisation still ahead of it.

Roberto Cordeschi, writing in Applied Artificial Intelligence in 2007, dates the first JOHNNIAC-printed proof to August 1956, which puts it inside the workshop’s own eight weeks. Jeff Shrager’s reconstruction of the original IPL-V code, an arXiv preprint from March 2026, puts the IPL-II implementation on the JOHNNIAC in late 1956 and early 1957. Simon’s own memoir called the Logic Theorist a running program in the summer of 1956. I would rather show that disagreement than settle it. Every account dates the famous result to 1957: 38 of the first 52 theorems in chapter two of Principia Mathematica, plus a proof of theorem 2.85 shorter than the one Whitehead and Russell published.

The gap between the claim and the machine was therefore weeks, not years. That is a good deal better than I expected to find when I started checking, and I would sooner report it than inflate the case.

Shrager’s own introduction says Newell and Simon brought a system that actually ran, and his historical section then says it was not yet running on a machine. A footnote in the same paper concedes that Arthur Samuel’s checkers program predated the Logic Theorist, was also heuristic, and did run. Samuel was at Dartmouth too, and he talked about checkers on 6 August. Nobody wrote “the only working program in the room” in 1956. That sentence appeared later, once the summer needed to have meant something, and seventy years of repetition turned a hand-simulated specification into a running machine.

The incident got simpler and more impressive with each retelling, nobody went back to the notes, and the version that survived is the one that makes the better story.

AI security folklore behaves the same way, and this is the earliest case of it I know.

The reception was lukewarm, which nobody involved expected. Minsky later explained his own casual response by saying he had sketched a heuristic search for a geometry machine and hand-simulated it himself in an hour or so.

What the sheet still does not measure

In April 2026, a review of agentic AI evaluation in Artificial Intelligence Review examined fifteen major AI agent benchmarks. Not one of them fully implements security testing, safety constraints, adversarial robustness evaluation or cost tracking, and coverage of those four is sporadic where it exists at all. What the benchmarks measure consistently is task completion. That is a capability metric, and it asks the question the seven problems in the 1955 proposal asked: what can the thing do?

A separate arXiv preprint from the same month catalogues forty agent safety benchmarks published between 2023 and 2026 and finds robustness covered by no benchmark as a primary dimension, which the authors call a complete evaluation blind spot. Preprint, so treat the counts as provisional, but the direction matches the peer-reviewed finding.

Ashby’s machine had one evaluation protocol and it was disturbance. Push a needle off centre on purpose and watch what the machine does to recover. Nothing in the 2026 tables corresponds to that test.

By default an evaluator runs the system on inputs it has curated and publishes the score. Red teaming exists and it is improving. Article 55(1)(a) of the EU AI Act requires providers of general-purpose models with systemic risk to conduct and document adversarial testing, and that obligation has applied since 2 August 2025; the Digital Omnibus, in force from 27 July 2026, deferred parts of the high-risk timetable and left Articles 51 to 55 alone. Checked 5 September 2026. Adversarial testing still happens after the capability numbers are in, and the tables people quote in procurement meetings and press coverage are capability tables.

The Ashby question

Ask this of any AI capability claim anyone hands you.

What was demonstrated while somebody was actively trying to make this fail, and who was that person?

The follow-ups

What access did the tester have, white-box or black-box? Was a white-box result reported as though it were black-box? What was the query budget? Which defences were running during the test and which were switched off? Was the finding at the level of the model, the deployed system, or an agent holding credentials and tools? A model, a system and an agent have different blast radii, and a result at one level does not transfer to the next.

Then the fifth, which almost nobody asks. Did anyone with an incentive to find a different answer repeat the test?

Why it took seventy years

Ashby was testing by disturbance in 1948, with four magnetic needles and a set of uniselectors, and he explained the method to the people in the room at Dartmouth on a Monday in July 1956. McCarthy wrote down the capability conjecture and nobody wrote down the evaluation method.

The founders were not wrong about what machines would eventually do. Against the seven problems in their proposal they have been closer to right than almost anyone predicted. What they never wrote down was how you would know whether any of it worked. Seventy years of expectation running ahead of results later, systems still pass every test anyone thought to run and fail the first one somebody designs to break them. That is most of what I write about here.

222fb9d292e3d0111656a33900e24a27cfb6a36eb7b202a94a66bb84766154b4?s=120&d=mp&r=g
[email protected] | About me |  Other articles

In the early 2000s, running emerging-technology risk labs at CyberAgency, a defence client asked my team to break the AI systems they planned to put into weapons. We did. That is where my work on AI security started, two decades before the current wave of attention. I kept at it through risk labs at IBM, Accenture, PwC and KPMG. In 2016 I co-wrote a book on AI and leadership. My commercial work today is quantum, at Applied Quantum, which is why this site sells nothing.

Related Articles