
How to tool up for assessment change without automating confusion.
Assessment tools in the AI era: a practical market map for universities
As universities navigate changing demands across departments, frameworks and platforms, contradictory AI-use guidance documents often emerge in parallel. The tool conversation runs ahead of the policy conversation, and the policy conversation runs ahead of the workflow it has to land in. When an institution then procures a new assessment tool, the tool lands on top of those contradictory documents and amplifies the confusion.
This is what automating confusion looks like in practice. Universities are not short of tools; they are working under pressures that rarely leave room for the patient analytical work of asking which problem each tool is meant to solve.
This piece is a public-source market map and a stance on how to read it. Product features, integrations and licensing vary by institution, configuration and contract. Always check details with vendors, procurement colleagues, internal technical teams and relevant policy owners.
In brief
Most assessment-technology decisions are workflow decisions in disguise. A market map separates the question into nine connected tool families, but the harder question is which family is solving which problem. This piece argues for one ordering: understand the assessment workflow first, treat the market map as evidence rather than as the answer, and choose the tool the workflow actually needs.
Generative AI has made older assessment questions harder to ignore, and has also created a genuinely new challenge: students can now generate plausible-looking work, and plausible-looking evidence of how that work was produced, faster than any current tool can reliably detect. That is the territory the market map sits in.
The QAA's Reconsidering assessment for the ChatGPT era names the same challenge from the UK regulator's vantage point. The structural argument here applies broadly; the local vocabulary will need translation for non-UK contexts.
Wasted friction and useful effort
Not all friction is equal.
Some friction is waste: duplicated admin, unclear handoffs, manual grade handling, confusing platform steps, inconsistent guidance and processes that rely on local memory.
Some effort is valuable: drafting, judging, moderating, interpreting evidence, giving feedback and helping students understand how to improve.
Assessment technology should reduce the first without accidentally removing the second. The alternative to unsupported automation is not unsupported human judgement, but accountable judgement: clear criteria, transparent evidence, moderation, escalation routes and human interpretation where decisions affect students.
Why the conversation runs ahead of the workflow
AI has not created every assessment problem universities are now facing. In many cases, it has made existing problems harder to ignore. Assessment design was already uneven. Feedback workflows were already time-intensive. Marking and moderation already varied between departments.
AI has added pressure because the stakes now feel higher. Institutions are thinking about evidence, fairness, workload, academic integrity, student communication, staff confidence, governance and acceptable use at the same time.
Two sector signals are worth reading together. Jisc's AI in marking and feedback pilot separates specialist tools designed for marking and feedback from general-purpose AI tools being used within controlled assessment workflows. EDUCAUSE's 2026 research on AI and work in higher education shows the same implementation gap from another angle: institutional AI strategies are becoming more common, but policy awareness, ROI measurement, tool governance and day-to-day staff use are still uneven.
The implementation gap is where assessment-technology decisions become complex. Universities are not simply choosing software. They are making decisions about evidence, trust, workload, integration, support and educational purpose.
The market is not one market
The assessment-technology market can look like a single category from the outside. In practice, it is a set of overlapping tool families that solve different problems. LMS-native tools support submission, marking, feedback and quiz workflows. Process-visibility tools give more context around how written work is produced. Similarity and originality tools provide document-level signals. AI-supported feedback tools sit closer to marking and review. Digital exam platforms focus on secure delivery. Portfolio and peer-review tools support different kinds of assessment design. STEM and programming tools often solve highly discipline-specific problems.
Grouping all of these together as assessment tools makes procurement conversations harder than they need to be. It also increases the risk that a university chooses a strong product for the wrong use case.
Market map at a glance
Table 1: Market map at a glance
If the main problem is… | Look at… | Examples | What to check |
|---|---|---|---|
Improving existing assignment, quiz or marking workflows | LMS-native assessment tools | Moodle Assignment, Moodle Quiz, Canvas Assignments, SpeedGrader, New Quizzes | Is the existing LMS being used well, or are unclear processes making the platform look like the problem? |
Understanding written-assessment process and authorship | Written-assessment and process-visibility tools | Cadmus | Do you need visibility into how work is produced, not only a final submission check? |
Checking similarity, originality or possible AI-written text | Similarity/originality tools | Turnitin, Inspera Originality | What kind of evidence is produced, and how will staff interpret it fairly? |
Supporting AI-assisted marking or feedback | AI-supported marking and feedback tools | Graide, KEATH, TeacherMatic, general-purpose AI tools used within controlled workflows | What does human oversight mean, and how will quality, consistency and workload be evaluated? |
Delivering secure digital exams at scale | End-to-end digital exam platforms | Inspera, WISEflow, Ans, ExamSoft, BetterExaminations, Questionmark, Cirrus | Are exam design, delivery, invigilation, accessibility, marking, grade return and support all in scope? |
Marking handwritten, scanned or problem-based work at scale | Online grading tools | Gradescope, Crowdmark | Who scans, allocates, marks, moderates, releases and resolves grades? |
Supporting peer, group, reflective or authentic assessment | Peer, portfolio and evidence tools | FeedbackFruits, PebblePad, Watermark | Does the tool strengthen assessment design rather than only monitor behaviour? |
Automating discipline-specific assessment | STEM, maths and programming tools | STACK, CodeRunner, Numbas, Möbius, CodeGrade | Does the tool fit the discipline, question type and staff capability? |
Strengthening marking, moderation or verification in Moodle | Moodle/Totara workflow extensions | Accipio Grade, Moodle Coursework/double-marking options | Can the institution improve Moodle workflows before procuring a separate platform? |
The table is a way of thinking, not a shortlist. Most UK universities I work with prioritise LMS configuration and Moodle marking and moderation extensions before adopting separate AI-supported feedback tools. The order in which the categories matter is institution-specific.
Start with what the LMS already does
Before looking outward, it is worth understanding what Moodle, Canvas or another institutional LMS can already support.
Moodle Assignment can support submission, grading, feedback, rubrics or marking guides, marking workflow and marker allocation, depending on configuration. Canvas offers native assignment, grading, rubric and moderated grading workflows through SpeedGrader and related features.
That does not mean the LMS can do everything. It often cannot. Some platform problems are actually process problems. A department may be using Moodle differently from another department. Marking states may not be consistently applied. Offline marking may rely on local workarounds. The gradebook may not match the assessment process people think they are running.
Sometimes a new platform is needed. More often, better configuration, clearer workflows, improved guidance or targeted development gets further.
Process visibility, similarity and what each one actually shows
Cadmus sits in the written-assessment and process-visibility part of the market. It provides a structured environment for written tasks and integrates with Moodle and Canvas through LTI 1.3, including launch, grade pass-back and membership synchronisation. Its Activity Reports give staff process-level information such as writing activity, work sessions and pasted content.
Process visibility can support better conversations and stronger evidence. It does not remove the need for judgement. Institutions still need to decide how the data will be interpreted, who will review it, what students will be told, and what will happen when the evidence is ambiguous. Reading process reports across 80 essays is not a small additional read; pilots should budget marker time for interpretation alongside the tool licence.
Turnitin sits in a different part of the landscape. It is best understood as a similarity, feedback and academic-integrity workflow tool around submitted documents, rather than a full writing-process environment. Its own guidance is clear that an AI-writing score should not be used as the sole basis for adverse action against a student.
A similarity report, an AI-writing indicator, a writing-process report, an exam log and a draft history are all different forms of evidence. Treating them as interchangeable creates risk.
AI feedback tools change where the workload sits
Tools such as Graide, KEATH and TeacherMatic sit in the AI-supported marking and feedback category. Jisc's current pilot is exploring these specialist tools alongside wider work with general-purpose AI tools such as ChatGPT, Claude, Gemini and Copilot.
This is the most operationally sensitive part of the assessment-technology landscape. Detection is one strand of the AI conversation; the more pressing question is whether AI can support clearer, more consistent or more timely feedback without weakening academic judgement.
A tool may generate feedback quickly. The real test is whether that feedback is accurate, fair, criteria-aligned, transparent to students and manageable for staff once setup, checking and governance are included.
Pilots in this category should ask: does the tool reduce workload, or move it into setup and review? Does it improve feedback quality, or produce more feedback of uneven quality? Does it support marker judgement, or create pressure to accept AI output too quickly?
Human judgement can be biased. Automation may be more consistent. That is precisely why human judgement needs structure, transparency and moderation. The capacity to judge the quality of one's own work and of others' work, including AI outputs, is what Bearman, Tai, Dawson, Boud and Ajjawi (2024) call evaluative judgement. It is a learnable academic capacity, not a technical one.
Digital exams and discipline-specific tooling
Some institutions are not primarily trying to improve coursework, but to manage formal digital exams, secure delivery, exam authoring, invigilation, marking, moderation, feedback release and reporting at scale. Inspera, WISEflow, Ans, ExamSoft, BetterExaminations, Questionmark and Cirrus enter the conversation here. Used well, an end-to-end platform creates more consistent institutional processes. Used without enough workflow design, it becomes a large technical project that still relies on manual workarounds.
For disciplines where the pressure is grading large volumes of handwritten or problem-based work consistently, Gradescope and Crowdmark are good examples. They focus on a specific grading pain point. The surrounding operating model still matters: who scans, allocates, marks, moderates, releases and resolves grades. In most UK departments, the answer to all six questions is the same teaching-track colleague. A tool buys time only if the operating model around it is staffed.
For STEM, maths and programming, STACK (open-source assessment for mathematics), Numbas (developed at Newcastle), CodeRunner and CodeGrade are discipline-specific assessment engines. Most UK universities I work with adopt these as departmental tools, not as whole-institution platforms.
For peer, portfolio and authentic assessment, FeedbackFruits supports peer feedback and group work; PebblePad supports portfolios, reflection, placement and longitudinal evidence; Watermark sits closer to outcomes evidence and accreditation. These tools matter because AI-era assessment conversations narrow too quickly when they focus only on misuse. Universities also need tools that support better assessment design.
For Moodle institutions, a category easy to miss: tools that extend Moodle rather than replace it. Accipio Grade supports advanced grading and quality assurance with internal and external verification. Moodle Coursework and double-marking plugins support allocation, double marking and grade agreement. If the institution is already committed to Moodle, the choice is rarely Moodle versus a new platform. More often it is better configuration, targeted plugins, bespoke development, or an external tool for a specific assessment type.
What “text tracking” really means
One of the most common AI-era questions is whether a tool can track the text. The phrase needs unpacking because it can refer to very different kinds of evidence.
Table 2: Text tracking evidence
Type of evidence | What it may show | What it does not automatically prove |
|---|---|---|
Writing-process data | Editing patterns, work sessions, pasted content, drafts or activity timelines | That misconduct has or has not occurred |
Similarity reports | Text matches with other sources | Whether the match is inappropriate or whether misconduct occurred |
AI-writing indicators | A statistical signal that text may have been AI-generated | Proof of AI use or misconduct |
Exam logs | Access, actions, timings or answer changes during an assessment | The full context behind a student’s behaviour |
Version history | How a file or document changed over time | Why those changes were made or who made every decision |
Declarations or process evidence | Student explanation of tools, drafting, sources or AI use | Independent verification of the whole process |
The practical issue is not whether data exists. It is whether the institution has a fair, transparent and workable process for interpreting that data.
Equity and the unintended cost of surveillance
The most under-discussed tool-selection question in 2026 is the equity dimension of detection and similarity tooling.
AI-writing indicators have well-documented higher false-positive rates for non-native English writing, neurodivergent writing patterns and writing produced with accommodations. An institution that pilots an AI-writing indicator without testing its false-positive rates against its own student demographic is buying a tool that may systematically disadvantage already-marginalised students. Vendor-claimed rates are not a substitute for institutional testing.
For students with English as an additional language, AI assistance is now functionally a translation tool. For neurodivergent students, AI scaffolding is closer to accommodation than to misuse. For international students, cultural variation in academic norms around assistance is significant and uneven across departments.
A tool that strengthens student agency and a tool that increases monitoring can look identical at procurement. The difference shows up in pilots, if the pilot is designed to surface it.
Integration is more than launching
Integrates with Moodle or integrates with Canvas can mean several different things: single sign-on, LTI launch, roster synchronisation, grade pass-back, group sync, deep linking, API access, student-record integration, or several of these in combination.
A technically available integration may still leave manual work in exactly the place the institution was trying to reduce it. One tool may launch neatly from the LMS but not return grades in the required format. Another may sync users but not groups. At small scale, those gaps may be manageable; at institutional scale, they become the work.
The sharper question is what data moves, when it moves, who triggers it, and what happens when it fails.
Trust extends into data protection and the regulator
Tool decisions are also data protection decisions. New processors mean new data flows, new consent questions and new retention regimes. UK universities should expect their data protection officer to be part of the pilot scoping conversation, not the procurement sign-off.
The regulatory shape matters too. The Office for Students' Condition B4 requires that assessment is valid and reliable, and that awards are credible. Workflows that support assessment quality also support the evidence trails that regulators and accreditors require. A senior assessment lead reading this piece should be able to put it in front of a Quality Committee.
Costs vary widely by institution size and contract type. Tools in the digital-exams category typically run six figures annually at institutional scale. Marking-extension and discipline-specific tools typically run smaller. A workflow diagnostic to scope what the institution actually needs is usually the cheaper investment to make first.
Questions before procurement, tests during the pilot
A small set of questions before procuring or piloting:
-
Which assessment types does the tool support well, and which does it not?
-
Does it strengthen student agency, or only increase monitoring?
-
Does it reduce inequity, or advantage students and staff who already have better AI access?
-
What evidence does it produce, and who will interpret it?
-
What manual work remains outside the platform?
-
What data protection, retention, consent and transparency questions need to be answered?
A pilot needs to test more than basic functionality and assess how effectively the tool works inside the institution's real assessment conditions. That means testing common assessment types, the marking and moderation workflow, student access and communication, staff setup time, feedback quality, exception handling, grade return, support burden, accessibility and accommodations, and whether the tool reduces or redistributes workload.
Pilots also need to be honest about what they cannot prove. A controlled use case may not reveal departmental variation. A positive staff experience in one context may not translate automatically to another. The capacity to give feedback that students can act on is itself a learnable academic capacity; Carless (2025) calls this academic feedback literacy. A tool that generates feedback quickly is not the same as one that supports feedback literacy.
The risk of automating confusion
An unclear moderation process remains unclear inside a new platform. A policy gap does not disappear because a tool has been configured. A weak integration may simply move manual work into a different team. An AI indicator may create more risk if staff are not supported to interpret it carefully.
This is why most assessment-technology conversations need to start with the workflow, not the demo. Universities with contradictory guidance documents looking to procure a tool, need to ask what the problem is that needs solving, before searching for a tool that fits.
Make the right work easier to do
There is no single best assessment platform for universities; there are better and worse fits for particular assessment types, institutional contexts and workflows.
A university reviewing written coursework needs a different solution from one reviewing secure exams. A STEM department needs a different tool from a portfolio-heavy professional programme. A Moodle institution with complex marking and moderation does not need the same approach as a university trying to replace a whole assessment-management ecosystem.
The strategy that underpins tool use matters more than the tool's feature set. In the AI era, the strategy is to reduce the friction that wastes time and attention, and to protect the intellectual and relational work that assessment still needs: agency, judgement, feedback, moderation, trust and learning.
That is how institutions avoid automating confusion, and make change workable.
About the author
Naomi Rowan is the founder of Gratitude Worldwide Ltd, a UK consultancy supporting universities on AI-era assessment, feedback, Moodle and platform workflow, and staff adoption. Recent engagements include work with the London School of Economics (assessment workflows and Moodle-based processes) and Nord Anglia Education.
The nine-category market map and the text-tracking evidence table are Gratitude Worldwide working tools, not sector canon. Gratitude Worldwide Ltd has no commercial relationship with any of the vendors named in this piece.
If you need help turning this market scan into requirements, pilot criteria, or a decision-ready shortlist
Further reading
-
Bearman, M., Tai, J., Dawson, P., Boud, D. & Ajjawi, R. (2024). Developing evaluative judgement for a time of generative artificial intelligence. Assessment & Evaluation in Higher Education, 49(6), 893-905. DOI: 10.1080/02602938.2024.2335321
-
Carless, D. (2025). Feedback Literacy Concepts and Practices: Toward Academic Feedback Literacy. Change: The Magazine of Higher Learning, 57(5). DOI: 10.1080/00091383.2025.2539038
-
Office for Students. Condition B4: Assessment and awards. OfS regulatory framework
-
Quality Assurance Agency for Higher Education (QAA) (2023). Reconsidering assessment for the ChatGPT era. PDF
By Naomi Rowan, Founder & Consultant, Gratitude Worldwide Ltd Published: April 2026 | Last updated: 19 May 2026