Take-Home Exams and AI: How Schools Are Adapting
By AI Detector 360 Editorial Team · · 8 min read
A take-home exam is an assessment you complete on your own time, without a proctor watching, usually with your notes and books available. That definition is accurate, widely used, and the reason most institutions misdiagnosed what happened to the format. It puts supervision at the center, which makes the obvious fix "watch them again," and the obvious fix is mostly the wrong one.
Take-home exams did not break because supervision disappeared. They broke because the format assumed that producing a competent written answer was itself evidence of understanding. That assumption held for a century and stopped holding in about eighteen months. Restoring supervision addresses the symptom; the durable fixes change what the answer has to demonstrate.
Key takeaways
- The take-home format relied on effort as a proxy for understanding, and generative tools broke the proxy rather than the proctoring.
- Four replacements dominate in practice, and each trades one kind of validity for a specific equity cost.
- Detection can support an assessment design but cannot rescue one, because error rates are highest exactly where exam prose is most formulaic.
- The formats that hold up best make the reasoning visible rather than making the writing harder to produce.
What actually broke
Assessment design has always run on proxies. You cannot observe understanding directly, so you observe something correlated with it and grade that. For written exams, the proxy was the answer itself: composing a well-organized argument about Reconstruction, or a coherent analysis of a case, took knowledge plus time plus effort, and the finished text carried the fingerprints of all three.
Generative models decoupled the three. Producing text that looks like the output of knowledge, time and effort now takes none of them. The proxy is not weakened; it is severed. That distinction matters because it tells you which repairs work. Any fix that makes cheating harder while leaving the proxy severed buys a term or two. Any fix that restores a connection between the artifact and the student's reasoning lasts.
Notice which faculty complaints turned out to be tractable. "Students can generate an essay" is not solvable by policy. "I cannot tell whether this student can reason about the material" is solvable, by a dozen different designs, some of which are cheaper than the exam they replace.
Where take-home exam AI policy actually stands in 2026
Institutional policy as of mid-2026 has settled into a rough pattern, though it varies enormously by country, discipline and even by department within a single university. Three positions are common.
The first is prohibition with detection: AI use is banned on unsupervised assessments and submissions are screened. This is the most brittle position, and the most likely to generate the accusation cases that consume department time. Vanderbilt's decision to disable Turnitin's AI detector in August 2023 remains the cleanest public statement of why: at a 1% false-positive rate, its roughly 75,000 papers a year would produce about 750 wrongly flagged submissions, and the institution judged that trade unacceptable.
The second is permission with disclosure: AI use is allowed and must be declared, sometimes with a required appendix showing prompts and edits. The third is format substitution, which sidesteps the question by changing what gets assessed.
Most departments now run some blend, and the blend is usually incoherent in a way students bear the cost of. A single term can include one course banning AI outright, one requiring a disclosure appendix, and one that assumes you will use it. Reading each syllabus individually is not pedantry; it is genuinely the only way to know the rule.
Picture what that means for one person. A second-year politics student takes four courses in the same term: a seminar that bans AI on all written work, a methods course that requires a prompt appendix with every submission, a survey lecture whose syllabus says nothing at all, and a writing-intensive elective that permits grammar tools but not generative rewriting. She uses the same editing extension in all four, because it is installed in her browser and she has used it since first year. In one course that is fine, in one it must be declared, in one it is a violation, and in one nobody has decided. She has not done anything different in any of them.
The four replacements, and what each one costs
Every substitute for the unsupervised take-home fixes something and breaks something else. The equity column is the one that gets skipped in faculty meetings and shows up later in complaints.
| Format | What it restores | What it costs |
|---|---|---|
| In-class handwritten writing | Direct observation of the work | Penalizes disabilities, slow writers, non-native speakers |
| Oral exam or viva | Reasoning made visible in real time | Heavy faculty time; anxiety and accent bias |
| Project portfolio with drafts | Process evidence over end product | Requires sustained access to tools and time |
| Open-AI exam with disclosure | Tests judgment rather than production | Advantages students with paid model access |
| Locked-down browser proctoring | Restores supervision at scale | Surveillance burden; unequal home environments |
The pattern in that table is not accidental. Every format that increases validity does so by increasing what the student must demonstrate in conditions the institution controls, and controlled conditions are exactly where students with unequal circumstances diverge most. A student sitting an in-class handwritten exam with a documented writing accommodation, or a student giving an oral defense in their third language, is being assessed partly on something the course does not teach.
Check your essay before you submit
See your AI likelihood score, sentence-level flags and confidence level — so a detector never surprises you.
Open the AI essay checkerOral checks, and the arithmetic nobody does
Oral assessment is the single most effective response to generative tools, and the reason it is not universal is a spreadsheet.
Run the numbers on a mid-size lecture course. Take 180 enrolled students and an eight-minute individual oral check, which is short for the format. That is 24 hours of contact time before scheduling gaps, no-shows, rescheduling and grading calibration, spread across whoever is available to conduct them. In a department where the same instructor teaches three such sections, the format is arithmetically impossible without new staffing.
What works instead, at scale, is the partial oral: a three-minute follow-up conversation on a random subset of submissions, announced in advance as a possibility for everyone. It changes the incentive structure without requiring 24 hours, and it produces something detection cannot, which is a direct observation of whether the student can talk about their own argument.
There is research support for why humans do well at this. A 2025 ACL study by Russell, Karpinska and Iyyer found that expert annotators who frequently use ChatGPT identified AI-generated text with 99.3% accuracy, and remained robust against evasion techniques that defeat automated tools. Notably, that is a finding about experienced human judgment, not about software. A faculty member who has read four of a student's in-class paragraphs holds a comparison no classifier has.
Open-AI exams, and the design that actually holds
The most interesting redesigns run in the opposite direction from prohibition. The assignment assumes model access, supplies or requires the transcript, and grades the things a model cannot supply: identifying where the output is wrong, connecting it to specific course material, defending a choice the model advised against.
A workable syllabus clause looks something like this, and you are welcome to adapt it:
For this exam you may use any AI tool. You must submit your full prompt-and-response transcript as an appendix, and your answer must include a section identifying at least two points where the tool's output was incomplete, misleading or wrong, with corrections grounded in course readings. Answers without the appendix will not be graded.
Two honest caveats. This design advantages students who can afford the better paid models, which is a real equity problem that departments should solve by supplying institutional access rather than by pretending it does not exist. And it takes considerably longer to grade, because you are assessing reasoning rather than scanning for correctness. Faculty designing along these lines will find more patterns in our guide to assignments that resist AI and still teach.
What detection can honestly contribute
Detection has a role here, and it is narrower than either its critics or its vendors suggest.
It is genuinely useful as an early, private signal, and least useful as evidence. The reason is specific to this context: exam prose is short, formal, formulaic and often written under time pressure, which is the exact profile that inflates false positives. The 2023 Stanford study in Patterns found seven detectors flagged an average of 61.3% of TOEFL essays by non-native English speakers as machine-generated, with one flagging 97.8%, while performing nearly perfectly on native-speaker eighth-grade essays. Timed exam answers by international students sit squarely in that failure mode. The Washington Post's April 2023 test of Turnitin's detector, which got more than half of 16 mixed student samples at least partly wrong, points the same direction.
So the defensible institutional use is screening that triggers a conversation, never a score that triggers a penalty. That is also why AI Detector 360 reports an explicit confidence level and a sentence-level heatmap rather than a bare percentage, and why our methodology page publishes how those confidence bands are set. A tool that will not tell you when it is unsure should not be anywhere near a grade. How instructors combine these signals in practice is covered in how professors actually check for AI.
For students sitting one this term
Practical, in order of value. Read the specific assignment brief rather than assuming the course rule, and get any ambiguity answered in writing before you start. Write somewhere that keeps version history, because a composition trail is the strongest artifact you can produce if questions ever arise, and it costs nothing to have.
If you want to know how your own writing reads to software before you submit, scan it yourself. AI Detector 360's AI essay checker gives sentence-level detail and a downloadable report, and the free scanner covers 5,000 characters with no account. Use it as information, not as a target: rewriting honest prose into worse prose to chase a lower number is a bad trade, and our guide to checking an essay before submitting explains where that habit goes wrong.
One last thing worth saying plainly. Nobody can tell you what share of institutions have actually changed assessment format, because the available surveys are self-selected and inconsistent, and no one collects this systematically at national scale. Treat confident percentages about "most universities" with suspicion, including ours. What is observable is the direction: assessment is moving toward formats where you have to show your reasoning, and that is a change worth preparing for regardless of what any survey says.
Check your essay before you submit
See your AI likelihood score, sentence-level flags and confidence level — so a detector never surprises you.
Open the AI essay checkerFrequently asked questions
Are take-home exams being phased out entirely?
No, and the evidence for a wholesale retreat is thin. What is visibly changing is the weighting, since many departments now cap how much of a final grade an unsupervised written component can carry and pair it with something observed. The format survives as one input among several rather than as the whole assessment.
Can my professor require me to use AI on an exam?
Some instructors do design assessments where AI use is expected and the grading criteria reward how well you direct, verify and critique the output. Where that happens it should be stated in the assignment brief, along with what you must disclose. If it isn't stated, ask before you assume either way.
What if I have an accommodation that makes handwritten exams difficult?
Raise it with your disability or accessibility office as soon as the format is announced, not after the exam. Assessment redesign has moved faster than accommodation practice in many institutions, and formats such as timed handwriting or live oral checks can create barriers that the previous format did not.
Does using Grammarly or a spell-checker count as AI use on an exam?
It depends on the course rule, and this is the single most common source of honest confusion. Basic spelling and grammar correction is usually permitted; generative rewriting that produces new sentences usually is not. Ask your instructor to state which side of that line their tools policy falls on.
Sources & further reading
Fair-use note: AI detection scores — from any tool, including ours — are probabilistic estimates, not proof. Never make academic, employment or legal decisions on a score alone.
Related reading
Designing AI-Resistant Assignments That Still Teach
How to design AI-resistant assignments that still teach: process-visible tasks, local data, in-class components, oral defenses and rubrics that reward thinking.
Sep 9, 2026 · 10 min read
How Professors Actually Check for AI in 2026
How do professors check for AI? Turnitin scores, writing-style comparisons, version history and oral follow-ups: what really happens after you submit.
Aug 3, 2026 · 6 min read

How to Check Your Essay for AI Flags Before You Submit
How to check your essay for AI flags before you submit: a responsible workflow to understand your risk, fix voice issues and keep proof you wrote it.
Jul 27, 2026 · 6 min read