What Exactly Is Our Real Problem with AI?
On AI-assisted writing, academic credentials, and the nature of expertise.
[This piece is dedicated to my friend and co-author, Prof. Walter E. Block.]
On 5 August 2026, Jason Arday resigned from his chair in the Sociology of Education at the University of Cambridge and from his fellowship at Jesus College. The resignation came hours after the university announced an investigation into his academic qualifications and honorary appointments, following months of allegations, many of them pressed by the former Cambridge research affiliate Nathan Cofnas, that passages of his doctoral thesis and subsequent papers reproduced the work of others without attribution, and that parts of his personal biography did not withstand scrutiny. Liverpool John Moores University, which awarded the doctorate, had previously examined complaints about the thesis and found no evidence of misconduct. Arday has conceded that he made mistakes, has rejected the characterization of himself as a fraud, and has said the campaign against him was racially motivated. His resignation statement, published through the Good Law Project, described a level of public scrutiny he was no longer willing to endure and insisted that stepping down should not be read as an admission.
Within hours, the story acquired a coda. A post circulated on Facebook, echoing a post by the biologist Colin Wright on X, reporting that the AI-detection service Pangram had classified the resignation letter as “Fully AI Generated” with “High Confidence.” Others ran the text themselves and reported the same result. So we arrived at a scene worth pausing over: a controversy about academic authenticity, resolved by a resignation, whose final act was a piece of software being asked to certify the authenticity of the resignation.
I want to spend very little time on Arday himself, partly because the misconduct allegations are being investigated by people with access to evidence I do not have, and partly because the detector question is more interesting than the man. So assume the detector is right. Better still, assume that detectors of this class are almost always right, which is more generous than the evidence warrants. What exactly have we learned?
We have learned something about textual provenance: on this assumption, we know which system produced the sentences. The detector says nothing about the decision to resign, the choice of grievance over contrition, what was conceded and what denied, or the judgment that the statement should go out through a campaigning legal organization rather than a university press office. Those are much closer to what we mean when we ask whether a document is “his.” The machine answered a different question, answered it well, and we treated the answer as though it settled the one we cared about. That substitution, made almost unconsciously and at enormous scale, is what this essay is about. We need to examine what, exactly, lies behind our intuitive anxiety about students, dissertations, peer review, credentials, and expertise itself.
What a Detector Can Actually Tell You
AI detectors are not nothing. The best of them, on independent audit, catch most fully machine-written text while rarely flagging human text, and that is a real capability. As evidence about a particular person, though, they belong with the polygraph. Half a percent of false positives is negligible until a faculty runs ten thousand essays through the machine and produces fifty students who must now prove they did not do something. Performance varies substantially across detectors and settings. Earlier systems showed serious bias against writing by non-native speakers of English, and mixed human-AI authorship, which describes most real academic writing, creates a separate classification problem again. The training data and the thresholds are proprietary, so the accused can neither reproduce the judgment nor cross-examine it. Turnitin says as much to its own customers, warning that the score “should not be used as the sole basis for adverse actions against a student.” Vanderbilt switched the feature off altogether.
Set every one of those limitations aside. Imagine a detector that is never wrong: no bias, no false positives, immune to paraphrase, telling you with certainty that a sentence emerged from a language model. It answers the question “was this text generated by an AI system?” and it answers nothing else. It does not tell you whether the person who submitted it formed the intention behind it, believed its conclusion, understood its argument, rejected three earlier versions on grounds he could articulate, or could defend it under hostile questioning. It does not distinguish the student who prompted once and pasted from the researcher who spent three weeks in argument with a model and used it, at the end, as a typist.
Academic authorship has never been reducible to who moved the pen. Anyone trained in law understands this intuitively, because legal practice runs on the principle that the person who signs is responsible for a document four other people touched. Responsibility, judgment, contribution and understanding are the currency. Provenance is evidence about those things, sometimes strong evidence, and we have just built a machine that measures provenance beautifully while telling us nothing about the rest.
University rules do distinguish permitted from impermissible use, but they cannot settle the question, because they are themselves attempts to answer it. They are written by committees responding to liability, workload and whatever their software vendor is selling, and they have moved twice since 2023. Citing current policy to decide whether current policy is sound runs in a circle. What I want to know is what the rules are trying to protect.
Where Did the Thinking Happen?
Here is a process I recognize, described honestly.
I have an intuition about a legal doctrine. It is half-formed and probably wrong, but it itches. I put it to a model in conversation, badly. The model pushes back and I discover the objection I had not seen. I restate the thesis more carefully. I add an observation from a case I once worked on. I ask for a map of the competing positions in the literature, then ask a second system to check the first one’s account, because I have learned that the friction between two models is where errors surface. I find that a version of my argument was made in 1978 by someone I had never heard of, and made better. I narrow my claim to the part that is still standing. I ask for the strongest objections. Two are decisive and I abandon those branches. One is a misunderstanding of the doctrine and I say so. Over days or weeks the thing acquires a shape. At the end I ask for the material to be organized and rendered into clean academic prose, and I edit that prose because the model’s instinct for emphasis is not mine.
Where did the intellectual authorship occur?
Certainly not in the typing, which I outsourced, and not much in retrieving the citations. The initial intuition mattered, though intuitions are cheap and most of mine are worthless. More of it lived in knowing which questions to pursue. Some in recognizing that the 1978 paper killed the original version of the argument and left a narrower one worth having, where a reader without domain knowledge would have skimmed the same summary and registered nothing. Some in knowing which objections were real and which merely sounded good. And at the end there is the plain question of whether I can defend the thing with the machine out of the room.
The bundle
I do not think there is a tidy answer, and I distrust anyone who offers one. What I am confident about is the structural point: the process above pulled apart a set of things that used to travel together. Possessing knowledge, finding knowledge, understanding knowledge, exercising judgment, producing something original, writing it down, demonstrating persistence, satisfying an institution, acquiring a credential and being recognized as an expert were, for most of the history of the university, so tightly correlated that nobody needed to separate them. You could not produce a competent literature review without having read the literature. You could not write elegant academic prose about a field without having spent years inside it. You could not locate the obscure judgment without knowing the reporter series existed. The final artifact was therefore excellent evidence of the entire invisible process behind it, and institutions could evaluate the artifact and be confident they were evaluating the process.
That inference is what AI has damaged. Someone can now produce polished prose without being a good writer. Someone can identify the controlling authority without knowing how to run an old-fashioned literature search. Someone can grasp the architecture of an unfamiliar debate in an afternoon rather than a semester. And, in the other direction, someone can generate a document of impeccable surface sophistication while understanding nothing whatsoever, including whether its citations exist. Both of those are now cheap. Any honest account has to hold them in view simultaneously, because the same tool produces both, and the visible output no longer discriminates between them.
What a credential was pricing
Signaling theory earns its place here, though it has to be handled more carefully than it usually is. Michael Spence’s 1973 model showed how, under asymmetric information, employers can rationally use a costly observable achievement to sort candidates by an unobservable trait, provided the achievement costs less to obtain for those who have the trait. Bryan Caplan pushed that framework hard in The Case Against Education, arguing that something like eighty percent of the private return to education is signaling rather than skill formation. The figure is heavily contested and you do not need it to use the insight.
A credential signals more than knowledge. It signals intelligence, conscientiousness, persistence, tolerance for tedium, willingness to submit to institutional authority for several years, and membership of a selective group. The strength of that signal depends on its cost. A degree that anyone could obtain in a weekend would signal nothing about anything, which is precisely why difficulty is not incidental to credentialing but constitutive of it. AI has reduced the cost of performing many of the observable tasks by which those underlying traits were inferred. It has not reduced the cost of possessing the traits. The signal weakens even where the person is unchanged.
Some of this split was visible well before AI and we managed not to notice. The content of the best courses at the best universities has been available free for years, and almost nobody confuses having watched them with having the degree. Access to information became dramatically cheaper long before AI, even if acquiring expert knowledge never did. What stayed expensive was demonstrating that you had submitted to a costly competitive process in order to acquire it, which tells you which of the two the credential was really pricing. AI is now doing to the demonstration what open courseware already did to the content.
Four accusations in one sentence
“He did not really do the work” accuses a person of very different things. Perhaps he never acquired the knowledge the credential attests to. Perhaps he cannot exercise the judgment it licenses. He may be claiming somebody else’s intellectual contribution as his own, which is fraud in the ordinary sense. But sometimes the complaint means something else entirely: he escaped work his examiners once had no choice but to endure. The first three are charges of incompetence or dishonesty. The fourth is a complaint that someone got a lower price than you did, and it is not obvious why that should be an academic offense.
The third charge deserves separating out, because our vocabulary for academic wrongdoing was built around it and does not fit the new case. Traditional plagiarism ordinarily involves identifiable prior work whose authorship has been misrepresented. It presupposes an author whose work was taken, which is why the remedy has always been attribution: naming him restores what he lost. When a student submits AI-generated text, there need not be an identifiable wronged author in the immediate transaction. Whatever is wrong with it lives in a misrepresentation made to the assessor rather than a misappropriation from a creator, and those are governed by different principles. A student who discloses at the outset how the work was produced has answered the charge of deception, though he may still have broken the terms he was assessed under, which is a different failure again. The Arday story happens to contain both wrongs, and nearly every account of it has run them together, because both arrive as the same visible symptom: text on a page that did not originate in the mind of the person whose name is on it.
There is an analogy that gets us further: AI as the newest research assistant. It breaks in one important place. Delegation has always been legitimate in academia because responsibility remained located somewhere: the principal investigator directs the project, decides which questions matter, and answers for the result. The assistant is a person with independent agency, his own expertise, his own judgment about whether the task is sound, and an interest in his own reputation. He can refuse. He can say the data do not support this. He can be wrong in ways that a competent supervisor learns to anticipate. AI has none of that. It does not refuse, it does not have a reputation to protect, it cannot hold responsibility, and its characteristic failure mode is to be confidently and plausibly wrong in a register that mimics competence. So the analogy holds where it matters legally, in that delegated cognitive labor is legitimate and responsibility stays with the principal, and fails where it matters practically, in that the delegate cannot share any part of the burden of being right.
I am not proposing that universities are merely signaling machines. They train people, and training works. They produce knowledge. They socialize novices into disciplines in ways that no amount of reading replicates. They maintain standards that protect the public from confident incompetents. They sustain communities in which arguments can be had properly. All of that is real, and all of it was bundled into a single process along with the signaling. The demand AI is making is that we say which of these functions we are actually performing when we award a degree, and that we defend each one separately.
The Value of Difficulty
Underneath the institutional question sits a moral intuition, and it is worth naming rather than sneering at. We believe that a hard-won achievement deserves more than an easy one. That belief is not irrational. Effort is often the only visible evidence of discipline. Repetition genuinely does build competence. Struggle, in a well-documented range of cases, is the mechanism by which learning happens rather than an obstacle to it.
But effort can also come loose from the value it was supposed to produce. Suppose one researcher spends six months reaching a conclusion and another reaches the same conclusion, or a better one, in six hours with better tools. What exactly makes the six-month version intellectually superior? The six months may have mattered. They may have given the researcher an ear for a tendentious source, a stock of half-remembered cases that will collide with something years later, or the accidental encounter with contradictory material that reframed the question. But six months can also mean interlibrary loans, photocopying, transcription, formatting, three days chasing a citation to a journal that turned out not to hold the volume, and a paper the researcher could not read in the original. The distinction I want is between productive difficulty and historical inconvenience, and it will not be drawn correctly by anyone who is committed in advance to either romanticizing or dismissing the old process.
Earlier technological changes
History suggests we have drawn this line before, badly at first and then well. When cheap calculators arrived in classrooms in the 1970s, the argument was about formation rather than convenience: whether a child who never learned to compute could be trusted to delegate computation. The math teaching bodies moved first, with the NCTM’s An Agenda for Action calling in 1980 for calculators at all grade levels and for “basic skills” to mean more than arithmetic. Teachers resisted for years, and in 1986 some of them picketed an NCTM meeting over a statement urging that calculators be integrated into classwork and assessment. Their concern was not stupid, and the eventual settlement vindicated part of it: we still teach arithmetic, because number sense turns out to be load-bearing for later mathematics, and we stopped teaching long division of six-digit numbers by hand as a test of seriousness.
The pattern repeated with word processors, electronic legal databases, statistical software and Google Scholar. Each time, a task entangled with competence became cheap, institutions worried that removing the task would remove the competence, and the task was eventually reclassified: either the underlying skill was found to matter and taught some other way, or it quietly disappeared and nobody mourned it. No law faculty today requires students to prove their seriousness by searching bound volumes of a citator by hand. The skill that mattered, knowing which authorities control and how to read them against each other, was retained. The manual labor of finding them was not.
Generative AI is not the same as a calculator and I have no interest in pretending otherwise. A calculator automates a step whose correctness is verifiable in principle by anyone who understands the operation. AI assists with synthesis, explanation, critique and composition, which are much closer to what we think of as thinking, and it produces output whose correctness is frequently not verifiable by the person receiving it. That is a genuine difference of kind and not merely of degree, and it is why the historical analogy is a starting point rather than an argument. What the history establishes is only this: the question “does removing this task remove a capacity we care about?” is an empirical question with a different answer for different tasks, and our first instincts about it have historically been unreliable in both directions.
Possession and Judgment
For most of the history of scholarship, an expert was partly defined by holding things other people could not easily reach. The private library was not decoration. The expert might not remember every page, but he held a map: he knew which book contained the answer, which authority was superseded, which controversy had already been settled and by whom. That map was expensive, it took years, and it could not be borrowed.
Search eroded that advantage. AI has taken a much larger piece of it. A person can now obtain, in an afternoon and in ordinary language, the leading authorities on a question, the history of a doctrine, the schools of thought and their quarrels, the standard objections, the terminology, and an honest account of what remains contested. Expertise survives this. What seems to me to move is the location of the scarce part of it. Access is cheap; discrimination is not. The scarce part is deciding what to trust. An apparent contradiction may be nothing more than two authors framing a question differently. An authority may be relevant without controlling. One factual variation changes the analysis and another does not. A paragraph that sounds entirely authoritative may be catastrophically wrong. Knowing which of those you are looking at, and what to ask next, remains expensive.
It is worth being concrete about the second capacity, because stated abstractly it sounds like a consolation prize. Henry Ford did not need to master the thermodynamics of internal combustion, and he did not invent the motor car, which already existed and was being built by people who understood it better than he did. What he had was the ability to see what was missing in a field whose technical foundations other people had laid. The knowledge he needed only had to go far enough, and it had to point at the gap. Something similar is true of a great deal of legal scholarship: the useful article is usually written by someone who noticed that a doctrine produces an absurd result in a case nobody had thought to run it against, and who knew enough to see that the absurdity was real rather than apparent. Deep field knowledge often produces that perception, which is why the two capacities were rarely distinguished. But the perception is the contribution, and it is not the same thing as the coverage.
The democratization of critique
The consequence I find most interesting is sociological rather than epistemic. AI has lowered not only the cost of acquiring information but the cost of entering a specialized argument. Serious critique of an academic paper used to require years of field-specific preparation, most of which was spent acquiring the ability to understand what was being said at all: the notation, the assumed background, the implicit commitments, the history of who had already tried the obvious objection and why it failed. An intelligent outsider might sense that something was wrong and be entirely unable to say so in a form a specialist would recognize as an argument. Now that outsider can ask for the terminology to be explained, the premises made explicit, the contested points identified, the strongest objections mapped, and the disciplinary dialect translated. He is not thereby an expert. He has become considerably cheaper to make dangerous to an expert’s argument.
This weakens something real: a practical monopoly over who may participate in specialized conversations. A curious lawyer can push into physics, and an independent researcher with no affiliation can take apart a published paper from home. Most such interventions will be bad, and the ratio will be dismal. Some will be excellent, and the excellent ones will come from people who would never have been admitted to the conversation under the old arrangement. Democratizing and destabilizing are the same event described from two positions.
Kuhn, expertise and apprenticeship
Thomas Kuhn is useful here if he is handled carefully. He is often read as saying that experts close ranks to keep outsiders out, which misses him badly. His claim in The Structure of Scientific Revolutions was that mature sciences work through shared paradigms, which include theories but also exemplary problem-solutions, methods, standards of what counts as a legitimate question, and a common language, all transmitted through long apprenticeship rather than stated as rules. Normal science is the productive puzzle-solving such a framework makes possible, and Kuhn thought its narrowness was a feature. It is precisely that framework, absorbed through years of professional socialization, that AI can now partially render into explicit prose for someone outside it. A hazard follows from Kuhn’s own account, or can at least be read out of it. What can be handed over is the map. What cannot be handed over is the tacit judgment that specialists acquire by working problems inside the paradigm until certain moves feel wrong before they can be shown to be wrong. Giving someone the map without the judgment produces a person who can formulate questions in the right vocabulary and cannot tell which of them are worth asking. Whether that person is a nuisance or a contributor probably depends on facts about him rather than about the tool, which is an unsatisfying conclusion and I think a true one.
Cheap production, expensive judgment
There is an economic version of this whole argument, and it is short. AI has collapsed the marginal cost of several inputs to intellectual work: information retrieval, basic synthesis, competent prose, routine critique, translation, summarization, first drafts of technical sections. When an input becomes abundant, the value migrates along the chain to whatever remains scarce. The candidates are not mysterious: judgment, verification, taste, the selection of questions worth asking, reputation, access to data nobody else holds, the ability to tell plausible nonsense from truth, and willingness to be answerable for a decision.
If that is right, the institutional implication is uncomfortable. Universities can attempt to preserve scarcity in tasks that technology has commoditized, which is a losing position defended at increasing cost, or they can shift toward certifying the capacities that have become more valuable rather than less. The latter is harder to assess, which is exactly why the former is tempting.
Nothing about this justifies optimism about the volume of what is coming. AI makes good research cheaper and makes bad research cheaper by the same factor, and there is far more bad research to be made. It produces citation-shaped objects that are not citations. It generates arguments that are plausible and false, in quantity, at no cost, and in a register indistinguishable from competence. If anyone can generate fifty objections to a paper in thirty seconds, the scarce good becomes determining which objection deserves five minutes. Cheap critique can bury serious work as effectively as no critique at all, which is the same argument arriving from the other side. As production becomes cheap, evaluation becomes the bottleneck, and the future expert will be distinguished less by the ability to produce material than by the ability to discriminate among an overwhelming supply of it.
The Case for the Old Process
The strongest case for the old process is that it manufactured capacities that cannot be separated from the process that produced them.
The evidence behind it is not soft. The learning literature on desirable difficulties, associated with Robert and Elizabeth Bjork, shows repeatedly that conditions which make acquisition feel slower and more effortful, spacing, interleaving, generation, produce better retention and transfer than conditions that feel fluent. Retrieval practice, documented by Roediger and Karpicke among many others, shows that the act of pulling something out of memory strengthens it in a way that re-reading does not, so that studying a text and testing yourself on it are not substitutes even when they feel equivalent. Research on cognitive offloading, surveyed by Risko and Gilbert, shows that we readily externalize cognitive work when a tool is available and that we do so even when internalizing would have served us better, and Sparrow’s work on search engines found that people remember where to find information at the expense of remembering the information. Polanyi’s point about tacit knowledge, that “we know more than we can tell,” states the objection in its purest form. Chase and Simon’s chess studies and the broader expertise literature suggest that expert intuition is largely stored pattern recognition, built from tens of thousands of encounters and not available on request from any external source, because the expert cannot fully specify what he is recognizing.
Push that further and it becomes genuinely uncomfortable for my position. If original ideas emerge from the collision of things simultaneously present in a mind, then a mind that has offloaded its contents has fewer things to collide. The person who spent three days reading cases may end up with something the person who read a synthesis in ten minutes does not have, and neither of them will be able to see the difference at the time. That is the shape of the worry: the loss is invisible at the moment it occurs and shows up years later as an absence of judgment nobody can trace.
AI and learning
The empirical work on generative AI and learning points in both directions, which is what one should expect. Bastani and colleagues, in a field experiment with nearly a thousand Turkish high school students, found that access to a standard GPT-4 interface improved performance substantially while it was available and left students worse off on unaided exams afterwards, and that a version with pedagogical guardrails, giving hints instead of answers, largely eliminated the damage. Kestin, Miller and colleagues at Harvard, with a carefully designed tutor that withheld full solutions and released one step at a time, found learning gains roughly double those of an excellent active-learning classroom, in less time. The two results sit together more comfortably than they first appear to. What both studies measure is whether the tool was configured to let the student skip the difficulty or to keep the student inside it.
So the objection survives, and it forces a reformulation. Asking whether effort is good gets us nowhere. The useful questions are these: which forms of effort create judgment; which create memory that remains necessary even when lookup is free; which train intellectual independence; and which merely reflect the limitations of the tools available to a previous generation. Reading a hundred cases in a field you will practice in for thirty years is probably the first kind. Manually reformatting footnotes to a house style is certainly the last. Between those poles lies almost everything that matters, and the honest position is that we do not yet know where most of it falls, because the question was never asked while the answer was not needed.
What Academic Credentials Certify
My friend and co-author, Prof. Er’el Granot, used to tell me about his problems with exams as a method of testing academic knowledge long before generative AI entered the picture. He is a physicist, and physicists have the advantage of teaching a subject where the difference between understanding something and reproducing it is unusually visible.
The criticism long predates both of us. A timed written examination measures some blend of memory, speed, stress tolerance, quality of preparation, ability to predict what a particular examiner rewards, and willingness to perform under conditions that resemble nothing the graduate will ever encounter again. These correlate with understanding. They do not constitute it. The complaints are old and some of them are documented. Miller and Parlett’s study of the examination game in 1974 described students they called cue-seekers, who worked out what a particular examiner rewarded and studied that instead of the subject. Essay grading is unreliable in ways that would embarrass any other measuring instrument: in one small study of examiners in a UK dental school, scripts re-marked by the same nine people produced the same grade category only once. Retention is the one place where the folk story overshoots, and it is worth being accurate about it: the best review I know of finds that roughly two-thirds to three-quarters of examined material survives a year, falling to somewhere under half by the end of the second, which is better than the cynical version and still means that a good deal of what a degree certified is gone within two years of certifying it.
If a take-home essay can now be produced by a machine, one conclusion is that students are cheating. Another, which I find harder to dismiss, is that the take-home essay was always a weak proxy for the thing we said we were measuring, and that its weakness was tolerable only because the cost of exploiting it was high. The assessment problem is decades old. AI removed the friction that had been keeping it out of view.
The return of the oral defense
The obvious response is to make people defend their work in person, and universities are doing exactly that. Institutions across the UK, US and Australia are reinstating proctored exams and pairing written work with oral assessment, and instructors keep describing the same telling pattern: flawless submitted work followed by blank incomprehension when the student is asked to explain it.
The instinct seems to me sound. If you claim to have written a thesis, you should be able to defend it: to say why you relied on that authority rather than the obvious alternative, to state the strongest objection to your own position and why it fails, to handle a hypothetical you have not seen, to say what evidence would falsify your claim, to explain a source you cited. An argument that a candidate has genuinely made his own survives pressure applied at an unexpected angle. An argument he has merely received does not.
I have to be honest about the limits of this, because my training is in law and I have spent a good deal of time around the failure mode. Verbal fluency creates an extraordinarily convincing illusion of depth. A quick, articulate person can enter an unfamiliar topic, absorb its structure in a few hours, seize on one genuinely interesting point, speak with the cadence of authority and leave a room of intelligent people convinced they have met an expert. I have watched it done. I have done it. Charisma, rhetorical intelligence and an accurate read of what an audience wants to hear will distort an oral examination as reliably as memorization distorts a written one, and they will distort it in favor of exactly the personality type that already does well in academic life. A serious oral defense therefore has to be designed to probe depth rather than to reward performance: novel hypotheticals rather than invitations to summarize, pressure on the weakest joint rather than the strongest, questions whose answers cannot be improvised from general intelligence. Which is to say that oral examination is another proxy, better suited to this moment, and not an escape from the problem of proxies. And that leads to the obvious question, the one everything so far has been circling: what is the thing we are actually trying to certify?
What a doctorate attests to
Put that question to any doctorate and the answers come out in a heap. That the holder knows more than almost anyone about a narrow subject. That he has produced an original contribution to knowledge. That he can conduct independent research. That he can teach the field. That he can evaluate other people’s research critically. That he has demonstrated unusual persistence. That he has survived a competitive institutional filter. That he has been admitted to a professional community. That he can be trusted to speak with authority in public.
All of these are true of some doctorates, and the system has never had to say which one it means, because until recently they arrived together. Now they come apart in awkward combinations. Original insight no longer guarantees exhaustive command of the literature, if it ever did. Someone with deep command may produce nothing original in a lifetime, which was true long before AI and is simply more visible now. Another researcher may direct AI systems with real skill while lacking most traditional research techniques, and an impeccably trained scholar may use the same tool badly and accept a fabricated citation because he never learned to distrust a fluent paragraph. One candidate may write almost none of the final prose and defend every intellectual decision in it. Another may type every word himself and add nothing that was not already in the sources. The credentialing system now has to say which of these differences it cares about, and it has no settled vocabulary for doing so.
An idealized past
A comparison with an idealized traditional education does not settle this. Something close to the nirvana fallacy is at work, in academic dress. The honest baseline is a system in which many students memorized and forgot, many credentialed professionals were mediocre, many assignments were formulaic exercises in producing what the grader expected, and intellectual labor was distributed across other people constantly. Doctoral candidates have always leaned on supervisors, statisticians, librarians, editors, translators and friends who read the draft and said the third chapter does not work. Scientists work in teams whose contributions cannot be disentangled, and lawyers sign documents drafted by clerks. The relevant baseline has never been wholly unaided individual cognition. The real comparison is between one configuration of distributed intellectual labor and another. That does not excuse fraud, and it is not meant to. It removes a fictional standard that is being used to avoid the actual question.
Where the novelty entered
That leaves originality, which deserves better than it usually gets. Academic originality almost never means producing an idea from nothing. It means reading widely, mapping what has been said, noticing a tension or an unasked question or a case the accepted explanation cannot handle, and building something from ingredients that all came from elsewhere. No one says a legal argument is unoriginal because it uses other people’s cases. So what changes if a machine helps assemble the map? If the researcher is the one who notices a relation between two positions that nobody had connected, the mere fact that the positions were retrieved by software does not obviously destroy the contribution. But the converse must also be conceded: it does not follow that every AI-assisted synthesis is original. The contribution still has to be identifiable, and the question to ask of any piece of work is where the novelty entered, and whether the person claiming credit can say where.
There is a related argument I find genuinely strong, and it rests on a distinction philosophers of science have used for most of a century. Hans Reichenbach separated the context of discovery from the context of justification, and Popper built on it: how a hypothesis was arrived at, by inference, by dreaming, by accident, by a chemist staring at a fire, has never been what confers scientific standing on it. What confers standing is whether it survives the tests we apply afterwards. Kekulé’s snake, Archimedes in the bath and Fleming’s contaminated plate are all part of the folklore precisely because the route was so obviously irrelevant to the result. Nothing about that changes if the route runs through a language model. A claim arrived at with AI assistance is justified or unjustified on exactly the terms that applied before, which are whether the sources say what they are said to say, whether the reasoning holds, whether the objections have been met and whether the thing can be defended. The distinction has been attacked, by Kuhn and others, as too clean, and they are right that the two contexts leak into each other. But it survives well enough to carry the weight I am putting on it, and it cuts both ways. If provenance is irrelevant to justification, then the detector result is beside the point; equally, the person who cannot verify what the machine gave him has failed at justification, which is the part that was always doing the work.
The second version of the argument is narrower. Millions of people have access to broadly similar systems. If one researcher nevertheless produces something new and valuable, the mere availability of the tool cannot by itself explain the result. Something differentiated that interaction: the initial intuition, the domain knowledge that made an ordinary-looking answer significant, the persistence to keep going after the first plausible output, the strange question nobody else thought to ask, the willingness to reject an answer that sounded good. This does not entitle the human to all the credit, and I want to be careful not to inflate it into that. It establishes only that universal access to a tool does not produce uniform outputs, and that selection and interpretation of what a tool generates can itself be a real contribution.
The one-prompt problem
Now the case that gives my whole position trouble. Suppose someone writes one exceptional prompt. The system returns an excellent argument in a single pass. The person reads it closely, understands it fully, agrees with it, could defend every step under examination, and submits it. Has he done no intellectual work? Counting iterations is not a measure of contribution, and I would not accept it as one against myself, so I cannot accept it here. But press the other way. If his entire contribution was recognizing that a machine’s argument was good, how much authorship should recognition generate? Recognition is real. Editors and juries and appellate judges make their living on it and we do not think they are doing nothing. It is also thin, and there is a point at which it becomes too thin to support a claim of authorship. My theory needs a limiting principle and this is where it goes. Choosing, understanding, directing, transforming, originating and taking responsibility are different relations to a piece of work. A credential that purports to certify an original scholarly contribution cannot rest on choosing alone.
There is a limit on the other side that I should state plainly. Understanding someone else’s brilliant article does not make me its author, however completely I grasp it and however well I defend it in conversation. Competence and defensibility cannot be the whole test. Something has to have entered the work that would not have been there without this particular person: a question, a distinction, a rejection, a connection nobody else drew. Where that threshold sits I cannot say in the abstract, and I doubt it can be settled in the abstract, which is an uncomfortable place to leave a criterion that credentialing bodies will have to apply case by case.
AI and academic competition
In sport, restrictions on technology make obvious sense, because the activity is partly constituted by comparing unaided human performance under agreed conditions. A runner on a motorcycle has not won the race; he has left it. Scholarship is different insofar as its purpose is the production of knowledge. If a tool helps us find a legal inconsistency or solve a scientific problem, the value lies in the inconsistency being found, not in the quantity of unassisted human labor consumed along the way. But universities are not solely in the knowledge business. A doctorate is simultaneously a contribution to knowledge, a training program, a personal achievement and a competitive credential, and AI touches those functions differently. It can plainly help produce knowledge while damaging training, when it lets the learner skip the difficulty. Its effect on personal achievement is harder to describe. As a competitive credential it creates a more serious problem. Much of the incoherence in academic reactions comes from people arguing about one function while their interlocutor is defending another.
The competitive function is where intuitions about fairness turn out to be least stable, and an argument I have from my friend and co-author, Prof. Walter E. Block, is the quickest way to see why. He applies it to performance-enhancing drugs. If nobody dopes, someone finishes first, second and third. If everybody dopes, someone still finishes first, second and third. The podium survives, so where is the harm? It lies in the internal redistribution. Athletes do not respond identically to the same drugs, and some have better physicians and more money. The ordering shuffles, and it shuffles along a partly different axis from the one the sport claimed to measure.
Academic competition looks to me to have been reshuffled in much that way. Baseline access is now widespread in well-resourced academic environments, and the distribution of outcomes has not collapsed: there are still top grades, still funded doctorates, still accepted papers. What has changed is who occupies which position, because people do not respond identically to the tool either. Advantage flows to whoever has the domain knowledge to see that an ordinary-looking output is significant, the persistence to push past the first plausible answer, the instinct to verify, and, awkwardly, the willingness to use the tool at all where the rules are unclear and the conscientious are penalized for observing them. It flows also to whoever can afford the frontier model rather than the free one. The ranking still sorts people, and we have lost our grip on what it sorts them by. Judged as a competition that is corrosive, since the ordering changed for reasons unrelated to what it was meant to represent. Judged as a knowledge-producing enterprise it may be an improvement, because the traits now rewarded look more like the ones that produce good research than facility at locating a source in a library ever did. Nothing in the current machinery tells us which reading to apply, which is the honest statement of the problem: we no longer know how to allocate authoritative qualifications, because we no longer know which of the qualities the new distribution rewards are the ones a qualification was meant to attest.
Why is academia so attached to the arrangements it has? I want to resist the cheap answer that older academics are protecting the value of their own suffering. The plausible causes coexist: sunk cost, professional identity built over decades, well-founded worry that students are not learning, the sheer difficulty of redesigning assessment at scale, and ordinary conservatism about practices that governed a profession successfully for generations. Alongside them sits an interest in the scarcity of the credential itself. I have gone back and forth about whether the word cartel does any work here, and strictly it does not, since nobody is fixing prices. But professional communities do control entry, define the qualifications for entry and benefit collectively when those qualifications stay costly, which is a structure with cartel-like incentives even where everyone inside it is acting in good faith. The strongest response is also true: standards protect genuine competence, and the public has an interest in not being represented by lawyers whose credentials were cheap. A restriction can serve the public and the incumbents at once. What follows is only that the burden of justification falls on the restriction, each time, in specific terms.
What Exactly Was Lost?
I want to end with the test I keep coming back to, because it is the one that has actually changed my mind about particular cases.
When AI removes or compresses an intellectual task, ask what precisely was lost. Sometimes the answer is serious: the student never acquired the doctrine, never built an internal model of the field, cannot defend the argument he submitted, cited sources that do not exist, failed to verify a claim he was responsible for, or skipped the specific difficulty that would have built the judgment he will need in ten years. Those are real losses and they should be treated as such. Sometimes the answer is that nobody spent six hours locating articles, or formatted citations by hand, or personally translated a paper in a language they do not read, or spent three days constructing a summary a machine produced accurately in ten minutes. If that is the whole of the loss, someone needs to explain why it matters.
Then ask why the loss bothers us. Sometimes the problem is epistemic: we can no longer tell whether a claim was checked. Sometimes students are failing to acquire capacities they will need later, or taking credit for work they did not do. There are economic and competitive interests in play as well. Existing credentials may lose value, and people who were ranked under one set of rules may find themselves competing under another. Institutions may simply no longer know how to assess the work. And sometimes, less comfortably, the issue is status: authority is harder to defend when the things that used to demonstrate it can be bought for twenty dollars a month.
For centuries, effort, knowledge, authorship, expertise and credentials were correlated tightly enough that institutions never had to define them separately, and the visible artifact could stand in for the invisible process behind it. Our discomfort now is the discomfort of losing an inference we relied on without knowing we relied on it. That is a larger problem than students using ChatGPT. It goes to the evidentiary structure by which academia decides who knows, who worked, who deserves credit, who gets credentialed and who is entitled to speak with authority.
So when a professor objects that AI has removed something important from academic work, the first question is what exactly was removed. And when he answers, the harder question is waiting behind it: was the thing that disappeared valuable because it produced knowledge and judgment, or because its difficulty gave us a convenient way to decide who deserved to be called an expert? I do not think that question has a single answer, and I am not sure the two purposes can be cleanly separated even in principle. But we are now going to have to try, in public, case by case, and with the awkward possibility in view that some of what we have been defending as the integrity of scholarship was also, quietly, a rule about who is allowed in.
Sources
Cited in the order they appear.
• Inside Higher Ed, on Cambridge opening an investigation and the resignation, 7 August 2026.
• CBS News, on the Liverpool John Moores review of the thesis, August 2026.
• Jason Arday, resignation statement, published through the Good Law Project.
• Facebook post reporting the Pangram result, and the earlier post by Colin Wright on X.
• Brian Jabarian and Alex Imas, Artificial Writing and Automated Detection, NBER Working Paper 34223 (September 2025).
• Weixin Liang and others, GPT detectors are biased against non-native English writers, Patterns 4:7 (2023).
• Turnitin, guidance on the AI writing report.
• Michael Spence, Job Market Signaling, Quarterly Journal of Economics 87:3 (1973).
• Bryan Caplan, The Case Against Education, Princeton University Press (2018).
• National Council of Teachers of Mathematics, An Agenda for Action (1980).
• Christian Science Monitor, on teachers picketing over the calculator statement, 9 May 1986.
• Thomas Kuhn, The Structure of Scientific Revolutions, University of Chicago Press (1962).
• Robert Bjork and Elizabeth Bjork, Making Things Hard on Yourself, But in a Good Way: Creating Desirable Difficulties to Enhance Learning (2011); Henry Roediger and Jeffrey Karpicke, Test-Enhanced Learning, Psychological Science 17:3 (2006).
• Evan Risko and Sam Gilbert, Cognitive Offloading, Trends in Cognitive Sciences 20:9 (2016); Betsy Sparrow, Jenny Liu and Daniel Wegner, Google Effects on Memory, Science 333 (2011).
• Michael Polanyi, The Tacit Dimension (1966); William Chase and Herbert Simon, Perception in Chess, Cognitive Psychology 4 (1973).
• Hamsa Bastani and others, Generative AI without guardrails can harm learning: evidence from high school mathematics, PNAS 122:26 (2025).
• Greg Kestin, Kelly Miller, Anna Klales, Timothy Milbourne and Gregorio Ponti, AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting, Scientific Reports 15:17458 (2025).
• C.M.L. Miller and Malcolm Parlett, Up to the Mark: A Study of the Examination Game, Society for Research into Higher Education (1974).
• Adam Hasan and Bret Jones, Assessing the assessors: investigating the process of marking essays, Frontiers in Oral Health 5:1272692 (2024).
• Eugène Custers, Long-term retention of basic science knowledge: a review study, Advances in Health Sciences Education 15 (2010).
• Times Higher Education, on the return of in-person and oral assessment; NBC, on oral exams in American universities.
• Hans Reichenbach, Experience and Prediction (1938), for the distinction between the contexts of discovery and justification; Karl Popper, The Logic of Scientific Discovery (1959).


