Prepare students for post-graduation AI judgement challenges
In a preregistered field experiment published this year in Organization Science, 758 consultants at Boston Consulting Group worked through realistic tasks with and without a generative AI assistant. On work within the assistant’s competence, results were strong throughout: more tasks finished, about a quarter faster, and at higher quality.
One task was designed to sit just outside its competence. There, 84.5% of the consultants working without AI reached the correct answer, while those using it reached it 70.6% of the time. The group that also received a prompt-engineering overview did the worst, at 60%. The better prepared they were to use the tool, the further they fell.
The consultants had no way of knowing in advance which side of the line a task fell on, and the authors report that the frontier is not obvious even on work people know well. Graders who had not been told the correct answer also rated the AI-assisted recommendations as more coherent and more persuasive, whether or not those recommendations were right.
The AI-assisted answers read better, including the wrong ones.
This study has travelled mostly as a caution about adoption: train people better, keep a human in the loop. Its own data offers little support for the first.
Nor do three preregistered experiments with 1,372 participants, reported by Steven Shaw and Gideon Nave, suggest an easy fix: time pressure, per-item incentives, and immediate feedback each moved baseline performance without eliminating the pattern the authors call cognitive surrender. Confidence rose even after errors.
So the failure is quiet and well presented, and nothing in the exchange itself warns the person that it has happened.
Where my own argument stopped
In May I argued in these pages that graduate programmes have a demand-side problem and that the window between admission offer and first class is high-leverage and currently empty. In August I argued that the act programmes fail to train is abduction and asked which cognitive acts a programme reserves for the student. Both pieces assumed a student inside an institution, and both stopped at the programme’s edge.
The people who need this most are working while they study and will keep working long after they finish. Fifty-three working teachers on a masters course in the Mekong Delta described in nine focus groups what generative AI was doing to their writing.
More than four in five worried that leaning on it would erode their ability to adjust as circumstances changed, because the tool could not know what was happening in their lives. Nearly three in four worried about becoming indecisive, and as many about feeling less responsible for work the tool had checked.
What they were doing about it matters more. More than nine in 10 were deliberately dividing the work, deciding which parts the tool could take and which it could not. More than half had turned to handwritten drafts and to explaining and defending their ideas aloud.
Some of it their lecturers had prompted. Most of it they had worked out for themselves, and after the degree there will be no lecturers to prompt them.
The second empty window
Writing here in July, James Yoonil Auh argued that the scarcity of this century is no longer knowledge but judgement, and that institutions must decide which forms of it may never be delegated. The question becomes harder one level down, where it carries a date.
Consider a graduate two years out. Promoted, or waiting to be. Deciding what matters, unsupervised, with an assistant open in the next tab. This is the period when professional judgement is actually formed and when no institution is watching.
What does a business school offer that person? A newsletter. An invitation to a paid short course, usually technical, usually bought by an employer. A network. Nothing that touches how they decide, and nothing that leaves a record that they decided anything at all.
The pre-programme window is thin and unoccupied. This one is wide, lasts years, and is occupied by marketing.
Even the most thorough response stops at the line
MIT published the strongest example in August. Its ad hoc committee on AI use spent five months on the question, with faculty from every school, undergraduate and graduate students, and staff from units including the libraries.
The report is candid about what AI is doing to the life of the institution and serious about remedies: assessments redesigned rather than AI-proofed, funded pilots for deliberately AI-free work, early drafting done in class by hand, in-person evaluation preferred to detection software, which it judges unreliable and a source of distrust.
Its eighth guiding principle is to think beyond the classroom and the campus, and it applies that principle to residence halls, clubs and living groups. The report speaks of preparing students for their lives beyond MIT, but life after the degree receives a single concrete proposal: that the institute track post-graduation feedback.
That is no criticism of a committee which answered the question it was given. It is the shape of the field. Even at its most ambitious, higher education reads the AI problem as one bounded by enrolment.
Other professions settled this long ago. Medicine, law, accounting and engineering treat qualification as a beginning, require documented records of practice afterwards and treat their absence as a professional failure rather than a private choice. Management education makes the larger claim that it forms judgement – and then ends at the ceremony.
What occupying the window would cost
Less than the pre-programme version. One email to graduating cohorts in the final term. A one-page instruction. One scheduled hour, 18 months later, in which the graduate reads their own record back. No platform, no curriculum, no headcount.
Two of the three disciplines on that page were described in the August commentary: write your position before consulting any tool, and record what you expect, what would prove you wrong, and the date you will check. The third is new and built for this problem.
Keep a delegation log for one working week. Write down what was handed to a model, the actual task rather than the category. Then mark every entry accepted without checking, and beside each mark write why: no time, no expertise, or it looked right.
The entries marked because they looked right are the frontier, in the graduate's own handwriting. That log is also the only baseline a school can realistically obtain from a working manager before intervening, and unlike a satisfaction score, it still means something 18 months later.
The obvious counter
The consultants' study used a 2023 model, and a dean will fairly ask whether the finding has survived. The authors are explicit that the frontier is not fixed but expands and shifts as models improve. What has not changed is that its location is hard to see from inside the work, and that fluency rises whether or not accuracy does.
What is being protected is not a task the machine cannot yet perform. It is a capacity that exists only in the person exercising it, and a judgement the graduate did not make trains nobody, however good the output that replaced it.
In May the argument was that programmes should occupy the window before the degree begins. This is the harder one. The window after it is longer, emptier, and it is where the judgement a school claims to have formed is either exercised or quietly handed away.
Fifty-three teachers in the Mekong Delta are already doing this work by hand, with a programme still around them. When it ends, nobody will keep the record but them.
Dr Phuong Nguyen (MBA, PhD) draws on three decades of commercial leadership and talent development, along with the wider literature on adult and professional learning, in his teaching at CFVG – Centre Franco-Vietnamien de Formation à la Gestion – in Vietnam. The author on LinkedIn: linkedin.com/in/drphuongnguyen.
This article is a commentary. Commentary articles are the opinions of the author and do not necessarily reflect the views of University World News.