News

From Dundee to Claude’s constitution: philosophy and AI

Dr Oisín Keohane, Lecturer in Philosophy, is interviewed by The Irish Times about AI labs recruiting philosophy graduates

Published on 27 August 2026

Writing in The Irish Times on 17 August, Joe Humphreys asked whether the AI labs recent enthusiasm for hiring philosophers amounts to anything. He put the question to Dr Oisín Keohane, Lecturer in Philosophy at Dundee: can philosophy make AI ethical, or is that overselling the discipline?

He replied that “anyone selling philosophy as a guarantee is overselling it. But the question is no longer merely academic. Earlier this year Anthropic refused to let the US government use its AI model, Claude, for the mass domestic surveillance of Americans or for fully autonomous weapons. In March the Department of War designated the company a supply chain risk to national security. Anthropic sued. Whatever else that was, it was a written set of values meeting real institutional pressure.

The better question is narrower: can philosophy change how these systems behave? I would say yes, and here, unusually, we can say something about how. Philosophy supplies the content whilst the training loop supplies what philosophy has never had with people: a way of checking whether it took, and of trying again if it didn’t. You can iterate, evaluate and retrain a model. You cannot straightforwardly do that with a person.

Anthropic calls its approach ‘constitutional AI.’ Rather than paying people to label harmful outputs, the model was handed a written list of principles and taught to critique and revise itself against them. The list of principles itself, Claude’s first constitution, was published in 2023. The constitution which was published this January does something else. It is addressed to the model itself and explains at length why its principles hold rather than stipulating that they do. That is an Aristotelian gambit, and it rests on an empirical bet: a system that grasps the reasoning behind a rule will adapt to cases the rule never anticipated, whereas a system holding only to the rule will follow it off a cliff. The example given in the constitution, whose primary author, Amanda Askell studied philosophy at the University of Dundee, is instructive: tell a model always to hand a distressed user a fixed list of support services, and it will do so even when those services are useless to that particular person. Worse, the model infers from the instruction what kind of agent it is meant to be: one that follows a fixed procedure rather than attends to whoever is in front of it.

There are however serious objections. Aziz Huq, an expert in constitutional law at the University of Chicago, told the New Yorker that companies are taking words like “constitution” – words connoting publicness and answerability – and transplanting them into settings where the mechanisms that give them meaning do not exist, then benefiting from the confusion. They are right. A constitution with no court behind it is not a legal constitution; it is a company document, revisable by the company that wrote it, and it should be read that way. Anthropic held its line against the Pentagon this year; nothing external obliges it to hold that line next year.

But that is a claim about accountability, not efficacy. Claude’s constitution is not packaging wrapped around a finished product. It is used in training, which is to say it shapes the model itself rather than being attached to it afterwards. How much, and how reliably, Anthropic has not disclosed. Its authors say more will be published in time; that reticence is a fair target in its own right.

Two questions, then, not one. Does the constitutional analogy earn the authority it borrows? No. Does moral philosophy, written carefully enough and fed into training, change what these systems do? Yes.”

Read the interview here (there is a paywall).

Story category Public interest