A Practice Guide for Educators · First Edition, 2026

A Calibrated Framework for AI-Resilient Assessment Design and AI-Ready Skills Development

Designing assessments that keep learning visible and build AI-ready skills.

J. M. Shalani Dilinika First edition · 2026 Six sections · five worked examples

Introduction

This guide introduces the AI-Resilient Assessment Design Framework. This is a structured, evidence-informed model for designing assessments that keep students’ understanding, reasoning, and capabilities visible, whether or not generative AI is used in their work. It also shows how assessment design can protect assessment validity while helping students develop AI-ready skills.

The framework was developed through a systematic review and synthesis of higher-education assessment research. The framework has two main purposes: to support valid assessment of student learning when generative AI may be used, and to help students develop the critical, evaluative, and responsible AI-use skills needed in academic and professional settings.

Who is this framework for

This guide is written for instructors redesigning assessments, programme and department leaders shaping assessment policy, and curriculum designers developing new courses in higher-education settings.

01 The problem

Generative artificial intelligence now enables students to produce essays, code, analyses, and multimodal work with varying degrees of independent effort. Its availability complicates long-standing assumptions about what an assessed artifact reliably indicates about a student’s own thinking and competence. Current detection tools remain unreliable and inconsistent, producing results that cannot be used as dependable evidence in assessment decisions. Restrictive policies may limit some forms of use but cannot fully prevent it, while a return to handwritten or tightly invigilated examinations risks narrowing assessment to formats that do not reflect the technological environments many students will encounter in academic or professional settings.

Assessing learning and building AI-ready skills do not have to be separate goals. Well-designed assessments can do both.

The challenge, therefore, is not stronger surveillance but better assessment design: assessments in which students’ understanding, reasoning, and capabilities remain visible, regardless of whether generative AI has contributed to the final artifact, and in which the way students are required to use, verify, and account for AI use itself becomes the means of developing the AI-ready skills that higher education must now cultivate in its graduates.

02 Three constructs of the framework

The AI-Resilient Assessment Design Framework frames assessment design around three interdependent constructs: (1) Authorization, (2) Verification, and (3) Accountability. Each corresponds to a distinct design question. Answered together, the three constructs support the design of assessments that measure what the student actually knows and can do.

Held together, these same three decisions serve a second purpose: they determine how often and how directly students practise using, evaluating, and taking responsibility for AI-assisted work — the same practice that builds their AI-readiness.

Authorization

Determines what GenAI is permitted to do in producing the assessed work. It specifies not only how much AI use is allowed but also what the tool may do, at which stage of the task, and under what conditions.

Verification

Generates evidence that the student possesses the competence the finished work implies. This evidence may come from visibility into the student’s process and interactions, dialogic or performative demonstrations of understanding, and transfer, multimodal, or triangulated evidence of retained capability.

Accountability

Establishes the responsibilities the student retains for AI-assisted work through transparency and provenance, validation and ethical responsibility, and ownership, reproducibility, and defensibility of the final work.

The central proposition

As the authorized level of generative AI contribution rises, the strength and diversity of verification and accountability must rise with it. This is the calibration principle, what distinguishes the framework from a checklist.

Verification Authorization Accountability HELD IN BALANCE AI-resilient assessment AI-ready skills
Figure 1 — The three constructs, held in balance, serving both AI-resilient assessment and AI-ready skills.

03 How to use the framework

The five authorization levels below show how the framework’s calibration principle works. As AI contributes more substantially to the work, stronger verification and accountability are needed to provide evidence of the student’s process, understanding, and responsibility.

Choose the level that best matches the intended competence, then use the corresponding verification and accountability measures. The final column shows the AI-ready skills each level tends to develop through the learning task.

The five authorization levels, with their verification bundle, accountability bundle, and potential AI-ready skills
Level What AI is authorized to do Verification bundle Accountability bundle Potential AI-ready skills
Level 1No AI Generative AI is excluded from the assessed task. Students are not allowed to use AI. Standard supervision or in-class conditions. No further evidence is required beyond the assessed performance. Ordinary academic-integrity expectations apply. Foundational judgement: demonstrating independent thinking.
Level 2Ideation & planning AI supports brainstorming, structuring, and organization. Analysis and drafting remain the student’s own work. A planning note, source map, or brief oral discussion of central decisions. A short AI-use statement identifying the tools used and the purposes served. Directing and framing: generating productive prompts, selecting and organizing AI-suggested ideas, and retaining ownership of the direction of the work.
Level 3Editing & revision AI supports language, style, and revision. Core disciplinary decisions remain the student’s. Version comparison or revision rationale, and a brief oral spot-check explaining key revisions and disciplinary decisions. A record of accepted and rejected AI edits; independent fact-checking and source verification. Critical appraisal: evaluating AI’s language and style suggestions, accepting or rejecting them with justification, and verifying facts and sources independently.
Level 4AI completion, human evaluation AI generates substantial elements. The student selects, evaluates, revises, and integrates the outputs. Demonstrate competence through an oral explanation, parallel assessment task (assessment twin), or transfer task, supported by evidence of key decisions made during the work. Disclose how AI was used, check the accuracy and ethical appropriateness of its output, and accept responsibility for the final work. Evaluation and integration: judging the quality of substantial AI output, revising and integrating it selectively, validating its accuracy, and taking responsibility for the final result.
Level 5Full AI AI creates most of the content. The student guides it, checks everything, and explains the result. A real-time walkthrough in which the student explains, adapts, and justifies the work while demonstrating retained disciplinary capability. For example, AI may generate a campaign package, but the student explains and defends the key creative and strategic decisions. A clear record of how AI was used, and the student defends the work. Orchestration and governance: directing AI systems, critiquing and adapting their output, and ensuring reproducibility, ethical integrity, and accountability throughout the process.

Note: The verification and accountability bundles are not intended to create a heavy documentation burden. They are flexible options, and educators should select only the elements that are proportionate to the task, the intended competence, and the role of AI. Stronger verification and accountability do not necessarily mean more paperwork; they may instead involve brief oral explanations, live demonstrations, or other embedded forms of evidence.

Reviewing and recalibrating

Review each level periodically. The following signals indicate when the design may need adjustment:

  • Level 1 — No AI: Check that supervision or in-class conditions do not create unnecessary accessibility barriers or affect what the assessment is intended to measure.
  • Level 2 — Ideation & planning: Increase or simplify the evidence depending on whether planning records clearly show students’ reasoning and decision-making.
  • Level 3 — Editing & revision: Review whether the editing boundary is sufficiently clear and equitable across language backgrounds.
  • Level 4 — AI completion, human evaluation: Use performance patterns to refine the parallel assessment task, reduce unnecessary documentation, and check whether the assessment format disadvantages some students.
  • Level 5 — Full AI: Recalibrate if the evidence no longer clearly shows what the student has learned, or if new AI tools make it possible to complete the task without demonstrating that learning.

04 The calibrated framework in action

Flow diagram. Intended competence leads to Authorization, which leads to the Calibration principle. The calibration principle splits into a Verification bundle and an Accountability bundle. The verification bundle leads to Construct-valid inference; the accountability bundle leads to AI-ready skills development, and the verification bundle also informs AI-ready skills development. Below, an Evidence-informed feedback and recalibration loop — within-cycle feedback, design review after implementation, and revision for the next cycle — feeds back up into Authorization and into AI-ready skills development.
Figure 2 — A Calibrated Framework for AI-Resilient Assessment Design and AI-Ready Skills Development in the Age of AI.

05 Applying the framework

Make three decisions explicit and keep them in balance. Because these decisions also shape how students practise with AI, each step designs for AI-readiness as well as validity, not a separate exercise layered on top.

01

Name the competence

Write one sentence: at the end of this assessment, the student should be able to demonstrate that they can […]. That statement is the construct everything follows from.

02

Determine the authorization level

Decide what role AI can play without displacing the student's reasoning. Choose a level 1–5, specify what the tool may do and when, and state which decisions stay the student's own.

03

Design the verification bundle

Identify the evidence that will show the competence, and match the bundle's strength to the authorization level. At higher levels, combine process, real-time performance, and transfer evidence.

04

Specify the accountability requirements

Determine what the student must disclose, validate, and defend. Scale it to how much the AI-performed part actually matters, not the volume of AI use.

Reviewing & recalibrating

Three questions guide each cycle

  1. 1

    Was the authorized use appropriate, and did students understand what was permitted and expected?

  2. 2

    Did the verification bundle generate evidence that could be interpreted, without redundancy, disproportion, or avoidable inequity?

  3. 3

    Were the accountability requirements proportionate to the consequential role of generative AI in the assessed performance?

06 Worked examples

The five examples below apply the framework across higher-education disciplines, arranged in ascending order of the authorized role of generative AI. In each case the same three constructs are applied, while the verification and accountability required vary with the role AI is permitted to play. The AI-ready skills listed are those that arise meaningfully from each task.

Level 1 · Example 1

Psychology concept identification task

Minimal AI · First-year undergraduate

Intended competence

Read a short introductory psychology text and accurately identify one foundational concept in their own words, demonstrating unaided disciplinary understanding at the start of the course.

Authorization

AI may not be used to summarise, explain, paraphrase, or interpret the assigned text. The reading and written response must be completed independently. AI may only be used to check spelling in the final submission.

Verification

A short written response identifying and explaining one core concept (such as classical conditioning or working memory) in a single paragraph, plus a brief in-class recall activity where the student restates the concept from memory.

Accountability

Standard academic-integrity expectations apply. The supervised task assures that no AI was used.

AI-ready skills

No specific AI skill. This task encourages independent reasoning, comprehension, and recall.

Level 2 · Example 2

Political science policy analysis task

Limited AI · First-year undergraduate

Intended competence

Identify a current policy issue, explain why it matters, and construct a short reasoned position using introductory political science concepts.

Authorization

AI may clarify unfamiliar terms and suggest possible angles or structures. It may not write any part of the analysis or supply the argument. Students work on an approved platform.

Verification

The completed analysis plus a short explanation of early decisions: what the student asked the AI to do, which suggested ideas they kept or rejected, and why. This shows the student directed the AI rather than relying on it.

Accountability

A simple statement identifying the AI tool used and its purpose. If the tool suggested any facts, these are checked against a trusted source and cited.

AI-ready skills

Directing and framing AI input; selecting and refining ideas responsibly.

Level 3 · Example 3

Nursing clinical reflection task

Editing and revision · Second-year undergraduate

Intended competence

Analyse a clinical scenario, apply relevant nursing concepts, and communicate the reasoning clearly.

Authorization

AI may improve clarity, style, and grammar. It may not generate clinical reasoning, identify risks, or propose interventions. All substantive decisions remain the student’s own.

Verification

The written reflection plus a short account of how the text changed in revision (which AI-suggested edits were accepted or rejected, and why), and a brief in-class oral check where the student explains one clinical judgement.

Accountability

The student discloses which sentences were AI-edited, confirms all clinical reasoning and examples are their own, and independently verifies any factual claims.

AI-ready skills

Critical appraisal of AI edits; responsible verification of factual content.

Level 4 · Example 4

Engineering design project

Substantial AI · Upper-level undergraduate

Intended competence

Define a design problem, evaluate alternative solutions, justify design decisions, and integrate technical constraints into a coherent proposal.

Authorization

AI may generate draft design options, produce preliminary calculations, and suggest optimisation paths. The student evaluates each suggestion, revises the outputs, and integrates only what meets the requirements. Sensitive or proprietary data may not be entered into any tool.

Verification

A short explanation of the design process: what the student asked the AI to do, how they evaluated the outputs, and what changes ensured feasibility and safety. If required, the student explains one design decision live and makes a small adjustment to show evaluative judgement.

Accountability

Clear provenance for all AI-assisted steps, validation of calculations and assumptions, and a short responsibility statement confirming the final design decisions are the student’s own.

AI-ready skills

Evaluating and integrating substantial AI output; validating technical content; taking responsibility for design choices.

Level 4 · Example 5

Digital humanities cultural analysis project

Substantial AI · Postgraduate coursework

Intended competence

Analyse a cultural dataset, interpret emerging patterns, and justify methodological choices in a way that reflects disciplinary reasoning.

Authorization

AI may generate code, produce visualisations, and suggest interpretive angles. The student evaluates, revises, and integrates only what aligns with the research question. Sensitive data may not be entered into any tool.

Verification

A short explanation of the analytical process (prompts used, outputs received, revisions made), a brief live explanation of one interpretive decision, and a small adaptation task done without AI to show the method transfers to a modified scenario.

Accountability

Provenance for AI-generated code and visualisations, verification of sources and methodological assumptions, and a responsibility statement confirming ownership of the final interpretation.

AI-ready skills

Evaluating AI-generated analytical work; verifying sources and methods; responsible integration.

Two things to hold onto

On equity

The calibration principle can add evidentiary burden on students who use AI most heavily. Detector bias, uneven tool access, and unequal digital preparation mean this can fall disproportionately on multilingual students, students with disabilities who use AI as access technology, and those in resource-constrained settings. Verification should measure disciplinary reasoning, not fluency, typing speed, or tool familiarity, and accountability should scale to how much the AI-performed part actually matters, not the volume of use.

What the framework is not

It is not a detection system and does not try to determine whether a student used AI. It is an assessment-design framework that keeps the intended evidence of learning visible, understanding, reasoning, judgement, and capability, whether or not AI contributed. Nor is AI-readiness a separate add-on: it is built through the same authorization, verification, and accountability decisions that protect validity.

Frequently asked questions

Why five levels?

Different assignments allow different amounts of AI support, but the student must always keep ownership of the thinking. As AI authorization increases, the risk of losing ownership increases too, so verification increases to match it. The five levels balance AI support with the checks needed to confirm that the work still belongs to the student.

Does having an instructor directly involved change how well AI-supported learning goes?

Yes. Instructor-mediated AI use, tools used by or alongside an instructor or tutor, tends to support learning more effectively than open-ended, unsupervised use. This is why the higher levels include clear boundaries and instructor-set expectations rather than unrestricted access.

Do I need to use all five levels?

No. Staying at Levels 1 through 3 for an entire course is a complete use of the framework, not a partial one. Many programmes are still cautious about AI use, so starting low is common, not a sign of falling behind.

How can I use this with a large class without spending hours on verification?

Large classes can use the framework without extra workload. Verification stays light at the lower levels, where short notes and random checks are usually enough. At the higher levels, students add brief process notes or short recordings, and instructors review these only when needed. In practice, most educators are already managing AI without structured systems, so the framework is designed to keep verification light rather than requiring review of every AI interaction. Deeper checks are used only for major projects at the highest levels, not for everyday work.

What if I suspect a student’s AI-use disclosure isn’t accurate?

Treat it as suspected, not proven. AI detectors are not reliable enough to serve as evidence alone. Ask for an alternate demonstration instead, such as a short live walkthrough.

What should students not put into an AI tool?

Personally identifiable information, theirs or anyone else’s, should never go into a public AI tool. Confidential or proprietary data should not either. This matters most at Levels 4 and 5, where real data or sources are more likely involved.

Attribution, reuse & contact

Citation Dilinika, J. M. S. (2026). The AI-Resilient Assessment Design Framework: A guide for educators.

CC BY 4.0 Open access

This work is licensed under a Creative Commons Attribution 4.0 International License. You are free to share it and to adapt it — including for commercial purposes — provided you give appropriate credit, link to the licence, and indicate whether changes were made. Institutions are welcome to adopt the framework, translate it, or rework it for local policy without seeking further permission.

Suggested attribution“A Calibrated Framework for AI-Resilient Assessment Design and AI-Ready Skills Development” by J. M. Shalani Dilinika, licensed under CC BY 4.0.

Questions & feedback

If you are adapting the framework for your institution, have a question about applying it to a particular assessment, or want to share how it worked in your context, get in touch. Feedback from practice shapes later editions.

shalanijayamanne@gmail.com

Further reading

  • Bearman, M., Tai, J., Dawson, P., Boud, D., & Ajjawi, R. (2024). Developing evaluative judgement for a time of generative artificial intelligence. Assessment & Evaluation in Higher Education, 49(6), 893–905.
  • Lodge, J. M., Howard, S., Bearman, M., Dawson, P., & Associates. (2023). Assessment reform for the age of artificial intelligence. Tertiary Education Quality and Standards Agency.
  • Perkins, M., Furze, L., Roe, J., & MacVaugh, J. (2024). The Artificial Intelligence Assessment Scale (AIAS): A framework for ethical integration of generative AI in educational assessment. Journal of University Teaching and Learning Practice, 21(6).