Client
A UK university group spanning several faculties and tens of thousands of enrolled students across undergraduate and postgraduate programmes
Sector
Higher Education
Engagement
Assessment integrity case management platform, AI-use evidence workflow and academic misconduct panel tooling - multi-quarter programme.
What the client needed
Individual academics had spent two years making their own judgement calls about suspected generative AI misuse in coursework, using whatever combination of AI-detection software, stylistic instinct and prior familiarity with a student's writing they happened to have. Detector output alone was not treated as sufficient evidence by the institution's own misconduct panels, correctly, given how unreliable a single detector score is in isolation - but no consistent process existed for what should sit alongside it. Two students submitting near-identical circumstances could end up with entirely different outcomes depending on which module leader raised the case and which panel happened to hear it, and the students' union had begun formally raising natural justice concerns about the inconsistency. The client needed a process that treated a suspected AI misuse case as a structured evidence-gathering exercise rather than a single detector score, produced a consistent record any panel could rely on, and could scale to a caseload that had grown several times over in two academic years without proportionally growing casework staff.
How we worked
- Worked with academic registry, students' union representatives and a sample of module leaders to define what a defensible evidence package for a suspected AI misuse case actually needed to contain, before writing a line of the platform.
- Built a case intake workflow that combines detector output, submission version history and drafting evidence where a student's own tools captured it, and a structured similarity comparison against the student's own prior submitted work, rather than treating any single signal as conclusive.
- Designed the evidence package explicitly to be interpreted by a human panel, not to render an automated verdict - the platform assembles and structures evidence; it does not output a guilty or not-guilty determination.
- Built a case management workflow that tracks every case through investigation, student response, panel hearing and outcome on a consistent timeline, replacing a patchwork of email threads and spreadsheets held by individual module leaders.
- Delivered an outcomes dashboard that lets academic registry see whether similar evidence profiles are producing consistent outcomes across faculties and panels, surfacing genuine disparities rather than assuming consistency by default.
- Ran the platform through supervised trial across two faculties for a full assessment period before wider rollout, tracking panel feedback and adjusting the evidence package format directly from that feedback.
Measured results
All figures verified with the client. Individual case details, student identities and specific outcome data withheld in line with our standard confidentiality terms and the institution's data protection obligations.
- Time for a module leader to assemble a complete evidence package for a suspected case fell from several hours of manual work to a structured process measured in minutes.
- Outcome consistency across faculties for cases with comparable evidence profiles improved substantially, addressing the specific disparity the students' union had raised.
- The institution can now demonstrate, case by case, that a detector score was never the sole basis for a misconduct finding - a distinction that mattered directly to two appeals raised during the trial period, both upheld in the institution's favour on process grounds.
- Casework staff report materially reduced time spent reconciling inconsistent detector outputs against their own judgement, freeing capacity as case volume continued to grow.
- Average time from case referral to panel hearing shortened, reducing the period a student spends with an unresolved case affecting their studies.
- Academic registry now holds a consistent evidence base across the full caseload for the first time, supporting a policy review of the institution's overall approach to generative AI in assessment.
"We were making decisions that affected a student's degree on the strength of a percentage score from a detector we didn't fully trust ourselves. What changed wasn't that we got tougher or softer on AI misuse - it's that two module leaders looking at similar evidence now reach the same kind of decision, and we can actually show our working if a student appeals."
Working on something similar?
If this engagement looks like the kind of problem you are facing, we would be glad to compare notes by email.
Context and constraints
Generative AI misuse in assessment moved, within two academic years, from an occasional edge case to a routine part of casework for most UK universities, and the client's existing process had been built for the occasional case, not the routine one. AI-detection tools improved unevenly over that period and remained prone to false positives, particularly for non-native English speakers and neurodivergent students whose writing patterns detectors sometimes flagged as suspicious for reasons entirely unrelated to AI use. The client's own policy correctly refused to treat a detector score as sufficient evidence on its own, which was the right call, but left a genuine gap: nobody had defined what should sit alongside that score to make a case decision defensible. Individual academics filled that gap with instinct, and instinct is not a process that produces the same answer twice.
The natural justice concern raised by the students' union was the forcing function for the engagement, and it shaped an important constraint from the outset: the platform could not become, or be perceived as, an automated accusation system. Any design that looked like it was replacing academic judgement with an algorithmic verdict would have made the fairness problem worse, not better, however much more "efficient" it might have looked on paper.
Designing an evidence package that survives an appeal
The core design decision was treating a suspected case as an evidence-assembly problem rather than a detection problem. The platform pulls together detector output, submission version history where the institution's own coursework platform captured it, and a structured comparison against a student's prior submitted work in the same module, presenting all of it to the reviewing academic and, ultimately, the panel, without collapsing it into a single score. We deliberately built in prompts that surface exculpatory evidence as readily as inculpatory evidence - version history showing genuine iterative drafting, for instance - because a system that only ever surfaces evidence supporting a misconduct finding will produce exactly the biased outcomes the students' union was already concerned about, just with better paperwork behind them.
We also resisted pressure from one stakeholder group to add an automated risk score ranking students by likelihood of misuse. It would have been straightforward to build and would have looked impressive in a steering committee update, but it would have reintroduced the exact single-number-drives-the-decision problem the engagement existed to fix, this time with a Halfteck-built score standing in for the detector score academics already didn't fully trust.
Rolling out without losing academic trust
Academic staff had reasonable scepticism about a new system touching student misconduct decisions, some of it shaped by frustration with the detection tools that had come before. Running a full assessment period as a supervised trial across two faculties, with a direct feedback channel to the delivery team rather than a generic support ticket, let us catch and fix genuine usability problems - an evidence package format academics found hard to scan quickly during a busy marking period, in particular - before the platform reached the rest of the institution. That trial period cost real calendar time, and we'd make the same trade again: a misconduct tool that academic staff don't trust enough to use properly is worse than no tool at all, because it adds process overhead without improving consistency.
Lessons learned
The first lesson was that "AI detection" and "assessment integrity" are not the same problem, and treating them as the same problem is what produces bad outcomes. Detection is a signal; integrity is a fair, consistent process for interpreting that signal alongside everything else relevant to a specific student's case. Institutions that buy a better detector and stop there are solving the wrong half of the problem.
The second lesson was that resisting the more "automated" version of a system can be the harder and more valuable engineering decision. An automated risk score would have been simpler to build and easier to demo. Structuring evidence for human judgement, in a way that surfaces both sides fairly, took more design work and produced a system people actually trusted enough to rely on when a student's degree was on the line.
The third lesson was that a fairness problem raised by a body like a students' union is usually a genuine signal worth treating as a design requirement, not a communications problem to be managed separately from the technical build. Building the evidence-package concept directly from that concern, rather than after the fact, is most of why the appeal outcomes held up.
If your institution or organisation is handling AI-related integrity or compliance decisions with more consistency and evidence than an individual reviewer's instinct alone, we would be glad to discuss what a programme like this might look like for you. Email sales@halfteck.com.