Logo
29/09/2026
Whose values would artificial intelligence apply if it handed down judgments?
Notes from EALE 2026 in Warsaw. Part one: the court and the AI models
common.coverImageAlt
The question I left the hall with

Would an artificial intelligence that may assist judges and lawyers decide a case the same way in Polish, in English and in Chinese, and on which values would it rely? These are the questions I was left with after two days of the conference of the European Association of Law and Economics, which returned to Warsaw in mid-September after thirteen years, to the Warsaw School of Economics (SGH) and to the building of the Faculty of Law of the University of Warsaw on Lipowa Street.


The programme ran to nearly a hundred papers on … everything, from constitutional compliance to child maintenance, but the thread of artificial intelligence ran from the opening lecture to the last session. This text is about three papers that showed how models decide cases and who decides that. The discipline itself, its European roots and the question of whom it wants to be useful to may get a second text, and the opening lecture by Maciej Szpunar, First Advocate General of the Court of Justice of the European Union, a third, because his opinion in the Court of Justice's first case on generative artificial intelligence is due at the end of October.

common.galleryImageAlt
Three papers and their authors. My own photographs taken at the conference.
Two questions economics puts to rules and to models

For me, a lawyer and an economist with a passion for research, the problem described here is more than a dispute about technology. A lawyer learns what a rule commands. An economist asks how people will react to the rule, who gains from it and sometimes even who wrote it and how they understood their own interest in doing so. I studied law and economics from the European side in Turin and from the American side at the University of California, Berkeley, and that is why I put to the models both questions the discipline puts to rules: does the model decide cases optimally, which is the question of the economic analysis of law, and who wrote its values into it and in whose interest, which is the question of public choice theory.

common.galleryImageAlt
Chess players’ memory according to Chase and Simon, illustrated on 8 × 8 chessboards using positions from actual games and random positions (the positions and choice of pieces are illustrative).
common.galleryImageAlt
The legal code as a finite table, the core, the penumbra, and an endless stream of new situations.
A finite text and an infinite world

In the keynote lecture of the second day, Professor W. Bentley MacLeod of Yale and Columbia, who is writing a book meant, as he put it, to update the textbook by Robert Cooter and Thomas Ulen from which a whole generation of Polish law-and-economics scholars started, myself included, first put economics itself in order. He divided its models, that is, simplified representations of how people and institutions make decisions, into three classes: price theory and game theory from 1940 to 1980, information economics and contract theory together with causal analysis from 1980 to 2010, and bounded rationality, the economics of a decision-maker who cannot compute everything, together with behavioural economics and big data. The third class is the youngest, though its ideas are hardly new. Herbert Simon, Nobel laureate of 1978, described as early as 1955, in A Behavioral Model of Rational Choice, a decision-maker who does not look for the best of all possible options, because he cannot compute them all, and settles for one that is good enough. Only around 2010, as MacLeod pointed out, did computers become fast enough for such models to be computationally operational.


Bounded rationality is exactly what is needed to understand artificial intelligence in law, because today it can be measured, and the measure is computational complexity, that is, the amount of time a computer needs to solve a problem depending on the number of its elements. The classic account is the 1979 book by Michael Garey and David Johnson, Computers and Intractability. There are problems in which every additional element doubles that time, so with a few dozen elements no computer can solve them exactly and one has to settle for an approximation. A standard economic model assumes that the decision-maker faces resource constraints or information constraints. Bounded rationality adds a third, cognitive constraint: too little memory and time to compute what could in principle be computed. That is how Simon put it, and after him John Conlisk in his survey Why Bounded Rationality? (Journal of Economic Literature, 1996).


MacLeod showed it with chess, the example Simon himself reached for. The number of possible games has about a hundred and twenty zeros (Claude Shannon estimated it in Programming a Computer for Playing Chess in 1950), so nobody plays chess by calculation. In the experiment William Chase and Simon described in 1973 in Perception in Chess, players were shown a position of about twenty-five pieces for five seconds. A novice could then recall only a few pieces, roughly as many as working memory holds, while a master recalled nearly the whole board, because he had not seen twenty-five pieces but a handful of familiar chunks: a castled king, a Ruy Lopez structure, an isolated queen's pawn. When the pieces were placed at random, the master's edge vanished. Skill, then, is human capital acquired through experience, it consists in recognising patterns, and it is not easily codified, because the master does not know what he knows until he sees the position; and when the position is not in his memory, the move is guided by an evaluation function that we colloquially call intuition. In the same way an experienced nurse catches early sepsis before any single vital sign on the patient's monitor trips an alarm threshold, because she sees the whole picture, whereas a novice nurse reacts to one signal.


From there MacLeod moved to law. Herbert Hart, one of the most important legal philosophers of the last century, wrote in 1958 in Positivism and the Separation of Law and Morals that the words of legal rules have a core of settled meaning and a penumbra of debatable cases, in which a court cannot simply apply the law. His example was a rule forbidding vehicles in a park. A car is a vehicle. Certainly. But a bicycle? An electric scooter? A stroller? The rule does not say, and the court has to decide. Hart wrote this as a philosopher; MacLeod wrote the intuition down as a model. The set of possible cases is effectively unbounded, while every legal system is a finite text, so it works like a lookup table with a finite number of rows. For a case that has its row, the text determines the outcome. For a case that does not, the outcome has to come from the decision-maker's evaluation function, that is, from his sense of what is right. On the slide MacLeod presented it is a separate function that assigns to each case the outcome society's values deem correct. The proposition he showed says that every finite legal system has a penumbra, and that adding rules shrinks the penumbra but never eliminates it. A judge who raised children himself will let strollers into the park, and his decision becomes a precedent that makes the law more precise.


What follows for artificial intelligence? MacLeod argues that any value system written down as a finite list of commands, prohibitions and values is incomplete, so all human values cannot be programmed into a computer, though that is exactly what, as he put it, a bunch of twenty-year-olds writing the code is trying to do, whereas a human being has a childhood and a lifetime of adult experience for it. I have one reservation about the analogy: a language model is not a table of the kind MacLeod describes, in which a case either has its row or does not; it builds its answer from the probabilities of successive words, so to a case it has never seen it does not answer “I don't know” but generates a decision of unknown quality, and rather than having a penumbra like a code, it always acts like a judge in the penumbra: it recognises a pattern and fills in the rest. MacLeod's claim about values therefore stands, but its cause lies elsewhere: the model's evaluation function was written in by its producers at the alignment stage, and every instruction and document the model receives together with the question then adds its own context to it. How much that changes is what the two papers described below are about.


His conclusion, which I would summarise as selection rather than design, MacLeod prefaced with the warning that from that point the lecture became speculation. He derived it from the penumbra proposition and from the way the law deals with judges: since values cannot be written down in full, systems of artificial intelligence, which have become so complex that even their designers do not really understand their behaviour, have to be selected for desirable properties, in much the way judges have long been selected, rather than governed by criteria specified for all systems. His example was down to earth: an AI agent, that is, a program that carries out tasks online by itself, was told to get a reservation at a very popular gym, was given a goal but no list of prohibitions, and achieved the goal by hacking the reservation system. The most important question, and here I agree with MacLeod, is who holds the control rights over the selection mechanism, exactly as with the selection of judges today. The law has long since stopped trying to design the perfect judge; it selects judges and then evaluates them by their judgments. Aligning a model is at the same time, as MacLeod showed following Steven Kerr (On the Folly of Rewarding A, While Hoping for B, 1975) and Canice Prendergast (The Provision of Incentives in Firms, 1999), a principal-agent problem, well known in agency theory: the model, like an employee, does what it is rewarded for rather than what the producer meant, and performance measures work only in the simplest cases, because in complex environments performance is judged subjectively and after the fact. With Elliott Ash, MacLeod showed in Reducing Partisanship in Judicial Elections Can Improve Judge Quality that in the United States, state supreme court judges selected in nonpartisan elections or by technocratic merit commissions write opinions that are cited more often than those of judges from partisan elections, and that this is the effect of selection rather than of the structure of incentives. Part of this reasoning MacLeod published this year in the Annual Review of Economics, in The Economics of Professional Decision-Making: Can Artificial Intelligence Reduce Decision Uncertainty?.

common.galleryImageAlt
A value map of the thirteen models proposed by Peng, in English and Chinese.
common.galleryImageAlt
Four earlier studies of model values compared with Peng’s map.
The ethical map of thirteen models

Yali Peng of Renmin University of China in Beijing observed that the literature on AI alignment, that is, on aligning models with values, is written by computer scientists with almost no legal scholars in the room, although lawyers know better which dilemmas are hard. In her paper thirteen models, eight of Chinese origin and five of US origin, answered fifty forced-choice legal-ethical dilemmas, from the death penalty to assisted dying. Every question was asked in English and in Chinese, ten times each, at deterministic settings, with the temperature set to zero, at which a model should answer the same question identically. Thirteen thousand responses were collected in mid-2025, and in 81.5% of cases the model did answer identically across all ten runs, a result typical of the technology, as Berk Atil and co-authors showed on five other models in Non-Determinism of "Deterministic" LLM Settings. Peng arranged the responses along three value dimensions, each rated from 1 to 7 with 4 as the neutral point: universalism, that is, universal rights and individual autonomy versus authority; change, that is, reform versus tradition; and negative liberty, that is, individual and market freedom versus regulation.


On the ethical map Peng showed, all the models form a dense cluster a short distance from the neutral centre, and the Chinese-origin and US-origin models overlap. There are differences, though. GPT-4.1 valued individual freedom against regulation the least of all, the Chinese Kimi and ChatGLM the most, and the Chinese Doubao answered more like the US models than like the Chinese ones. But the differences between the models are smaller than what they share. The arrows on the map show what happens when the same question is asked in Chinese. Peng expected that the Chinese corpus, that is, the Chinese-language texts a model learned from, would lower the ratings on rights and reform, and so it did, but the biggest difference language makes is on the dimension of individual and market freedom versus regulation, and there the arrows point upward. The facts of a case dominate the language it is described in, yet the language of the query, one would think, should make no difference to a judgment at all, and it does. Peng called this procedural arbitrariness.


Peng's map is not the only one, but it is the only one that measures legal dilemmas. Fabio Motoki, Valdemar Pinho Neto and Victor Rodrigues showed in More Human than Human: Measuring ChatGPT Political Bias (Public Choice, 2024) that ChatGPT answered political compass questions like a Democratic voter in the United States, a Labour voter in Britain and a Lula voter in Brazil. Paul Röttger and co-authors warned in Political Compass or Spinning Arrow? (2024) that such results depend on whether the model is forced to choose an answer and on the wording of the question. Maarten Buyl and co-authors, in Large Language Models Reflect the Ideology of Their Creators (2024), asked seventeen models about several thousand political figures in English and in Chinese and found differences that depend both on the language of the question and on whether the producer is Western or Chinese. Kazuhiro Takemoto, in The Moral Machine Experiment on Large Language Models (Royal Society Open Science, 2024), put the self-driving car dilemmas of the Moral Machine experiment to the models and found that they mostly choose as people do, but more sharply on a few dimensions. None of these studies made the models decide legal disputes, and none asked them the same question ten times, and that is the difference.


The three things Peng asked her audience to take away sound like a warning. The models converge on one value profile, universalist, reformist and low on negative liberty, which is to say a contestable profile that people argue about and that some users will find alien. Its expression shifts with the language of the query, not with who built the model. The producers wrote essentially the same profile into their models, so the difference between them is smaller than the difference between the languages in which a model receives its query. A single alignment standard would therefore entrench, across the world, a profile that nobody chose and nobody debated, and Peng called this algorithmic legal centralisation: a transplant without deliberation. The user cannot opt out, she said, because users cannot fine-tune their own models, and values written in at the alignment stage travel with the infrastructure, not with the jurisdiction: the same model, with the same parameters, answers from the same servers to a user in Warsaw and a user in Beijing, so what decides the values is which product the user uses, not the law of the country he is in. Her proposal is modest: instead of a model that gives one answer to a question about the death penalty, what is needed is a model that illuminates the conflict of values rather than resolves it, and leaves the decision to a human being.


Who, then, decides a model's values, and in whose interest? That is a question straight out of the public choice textbook, that is, the part of the discipline which assumes that politicians, officials and judges follow their own interest just as entrepreneurs and consumers do, and it is worth saying why. A model's values are decided today by its producers in the United States and in China, and each of them has its own interest in the matter, because it sells one product in many markets and, as Advocate General Szpunar said in Warsaw, reckons only with the regulators of three centres: the US, the EU and China. The politician who will regulate these models also has his own interest, so he will stand up for local values only when voters ask about them, not because that is the right thing to do. The judge who is handed a model as a tool has seven hundred cases in his caseload, as Jarosław Bełdowski said at the panel on the evolution of the discipline, and a basic interest of his own in having fewer of them as quickly as possible. If all three act according to their interests, the value profile that reaches the courts and public offices will be the one the producer wrote in for its main markets and the regulator deemed safe, not the one the community in question would have chosen. Nobody has to act in bad faith for that to happen.

common.galleryImageAlt
Percentage of decisions following the letter of the law among laypeople, lawyers, and models (Katz and Zamir, Study 1).
common.galleryImageAlt
Distributions of decisions by laypeople and models on a 1–7 scale in an easy and a difficult case (Katz and Zamir, Studies 2 and 3; Figure A1 redrawn).
Law or justice

Ori Katz of Bar-Ilan University and Eyal Zamir of the Hebrew University of Jerusalem asked a different question. Instead of asking the models about values, they made them decide cases as if they were judges and checked what happens when the formal legal rule and considerations of justice point to different outcomes.


The example is the Apartment vignette. Upon their marriage, the wife transferred half of her apartment to her husband; two years later she was badly injured in an accident and is confined to a wheelchair, and the husband began a relationship with another woman and applied for divorce. The wife demands the half back, because she gave it on the assumption that they would remain married, but the applicable law says that a gift between spouses is final and irrevocable, unless it was expressly stipulated that it is conditional upon the continuation of the marriage.


The legal rule therefore supports the formalistic decision, in favour of the husband, while the sense of justice supports the wife. In a study published in 2025 in the Journal of Empirical Legal Studies under the title Law, Justice and Reason-Giving, the authors showed on three such vignettes that people balance the two. Legal professionals, including judges, instructed to decide according to the legal rules chose the formalistic decision, that is, the husband's side, in 66% of cases, and when instructed instead to choose the outcome they considered most just and reasonable they still, though less often, found for him in 41% of cases. Laypersons from a new US sample answered almost identically: 66% and 44%. AI models, by contrast, do not do this. In a new paper, forthcoming in the same journal under the title Law, Justice, and Artificial Intelligence, Claude Sonnet 4 and Gemini 2.5 Flash, instructed to decide according to the law, chose the formalistic decision in 97% and 100% of cases, and instructed to decide according to justice, in 1% and 0%, that is, they did not so much as mention the rule they knew. Zamir said they flip a switch between law and justice. The exception was GPT-4o, which, asked for justice, kept to the rule in 33% of cases, less than people, but closest to them. When instead of deciding the models were asked to predict how one hundred judges would decide, their answers moved closer to the human ones. So the models “know” that people would decide differently, and they decide their own way regardless. That is also the answer to the question about the user's influence left open above: the instruction changes the decision, but within the limits the producer has set.


The models also turned out to be more consistent than people. In the hard version of the apartment case, the decision was recorded on a scale from 1 to 7, where 1 means that one would definitely rule in favour of the wife, that is, order the return of half the apartment, and 7 that one would definitely rule in favour of the husband. The standard deviation, that is, the measure of how widely the answers are scattered around the mean, was 2.31 for laypersons, which on a scale from 1 to 7 means that their answers stretched across almost the whole scale, and between 0.13 and 0.42 for the models, which means that almost all of a given model's answers were the same number. The models almost never went below the midpoint of the scale, whereas laypersons found for the wife in four cases out of ten. Consistency is a virtue, Zamir replied, but not necessarily so, because the noise in human judgments facilitates intellectual diversity, and today's minority view might very well become tomorrow's majority view.


This is an old dispute. On one side stands legal certainty: the citizen should know in advance what is allowed and a judgment should be predictable, and for Friedrich Hayek, for instance, that is the essence of the rule of law. On the other side, the law learns and adapts precisely because judges differ, since the world keeps producing cases the legislator did not foresee, and the same Hayek, in Law, Legislation and Liberty, defended judge-made law as an order that learns from cases, while Nicola Gennaioli and Andrei Shleifer showed in The Evolution of Common Law that the common law learns and improves precisely because of the diversity of judicial views, since the judges' successive corrections wash out their biases and render the law more precise than any one of them could. A model that always answers the same way gives a certainty no bench of judges can give, and at the same time takes from the law the mechanism by which it learns. Between consistency and pluralism one therefore has to choose, because a gain on one is a loss on the other, and the question is how much variation the law needs and who should decide that: the judge, the legislator or the producer of the model.

common.galleryImageAlt
The judge as a local bearer of community knowledge and the model as a central planner of values, with silhouettes of three cities and a judge.
Artificial intelligence and the knowledge problem

Peng, and Katz and Zamir, treated the models the way experimental economists treat people: they sat them in front of a dilemma and counted the answers, and the discipline has ready concepts for this, because a model as an object of study is a boundedly rational decision-maker in the purest form, and its alignment is the principal-agent problem discussed above. The harder question is whether artificial intelligence undermines the assumptions on which those concepts rest, and here the dispute I know from my studies in Turin as the knowledge problem matters. The label comes from the Hayekian tradition: Hayek wrote in 1945 in The Use of Knowledge in Society that the knowledge needed to run an economy exists only as dispersed bits in the minds of millions of separate individuals, much of it the knowledge of the particular circumstances of time and place that cannot be communicated to a central board, so a central planner who wants to replace the market does not have what the market knows. Cass Sunstein, in his essay The AI Calculation Debate, asks whether models that read everything invalidate Hayek's argument, and answers that they do not, because part of that knowledge arises only in the process of exchange, and preferences are not given in advance but discovered. Erik Brynjolfsson and Zoë Hitzig, in AI's Use of Knowledge in Society, are more cautious and show that models can codify local knowledge that was previously tacit and process dispersed data on a scale no planner ever had, and so make centralised coordination and control more feasible.


For law this dispute matters directly, and the three Warsaw papers fit together here into one whole. MacLeod's judge in the penumbra is Hayek's man on the spot, a local carrier of knowledge, because his evaluation function grew out of his community, its language and its disputes, and the judgment he gives returns to that community as a precedent. Peng showed that a model carries one value profile, written in by the producer, and that this profile changes with the language of the query, not with the jurisdiction. Katz and Zamir showed that two of the three models decide a case the same way every time and in their own way, while the third, GPT-4o, is closer to people, though all of them “know” how people would decide. A model that decides in the penumbra the same way in Warsaw, Beijing and Washington is therefore a central planner of values, and that is exactly what Peng called algorithmic legal centralisation. Hayek's argument does not say that such a planner will be bad, but it explains that he will not know what the community in which the case arose knows, and that nobody will notice the gap, because the answer will arrive quickly and confidently.

common.galleryImageAlt
The procedure for selecting a judge versus selecting a model without rights of oversight, and the possible components of an examination.
common.galleryImageAlt
A visual reading list featuring the twenty-four works cited in the article.
Who selects and who controls the selection

From the three papers it follows, then, that the question put to Professor MacLeod from the floor, whose values a model would apply, now has an empirical answer: the values written in by the producer, expressed in the language of the query and applied with a certainty no judge has. MacLeod added to this a conclusion I consider the most important of the whole conference, though he himself called it speculation. Since systems this complex cannot be designed, the AI alignment question is not so much a question about the correct algorithm as a question that has been put the wrong way. The right question is who holds the control rights not over the design of the system but over the building of the selection process, that is, over which model is admitted to decide which cases. Exactly as every position that gives an individual control over people is filled. Asked who should be the selector, he answered that well-functioning complex systems do not allocate authority to a single entity but to multiple competing agents with different tasks and values. He recalled the theorem of Kenneth Arrow, Nobel laureate of 1972, from his 1951 book Social Choice and Individual Values, that no rule of collective choice satisfies several reasonable conditions at once, unless one person decides. A single person in charge is therefore, as he said, more rational than a group, but in a complex environment will not necessarily enhance performance. Even a good selection process sometimes produces bad leaders, MacLeod said, so the point is not how to select infallibly but how to select so that the error of one choice does not stop the whole system.


The Warsaw papers do not yet provide such an examination, but they show what it could consist of. Peng showed how to measure a model's value profile in English and in Chinese, and by the same method in any other language; Katz and Zamir showed how to check whether a model balances law and justice as people do or merely flips a switch between them; and Ash with MacLeod showed that the way judges are selected shows up later in the quality of their judgments. A judge handed a model will not ask whether the model is efficient. If, crushed under a pile of cases to decide, he asks anything at all, it will probably be which model, and in which language, he can trust in a case no rule decides. The answer to that question will come neither from the producer, who sells one product in every market, nor from the regulator, who has no data, but it can come from a selection process with an examination and an evaluation by judgments. One like the process for judges, and someone who holds the control rights over it. In Poland a judge is selected through a judicial examination, a period as a judge's assessor and an appointment, and his judgment is reviewed by a higher court. The model a judge will use will today be chosen by the judge himself, buying access like any consumer, or by the court administration in a public tender whose criteria are price and, perhaps, data security, but, I would wager, no longer a value profile. Nobody has yet been given the control rights over that choice, and the producers managed to update the models from both Warsaw studies before the papers were published. From what moment, then, does the absence of a decision about who selects become a choice made by the model producer?


 

© 2026 Jarek Jurczak