Improving exams: a practical guide
Background
Invigilated exams are perhaps the most widely used form of assessment in higher education, and their use is increasing in response to the threat to assessment security posed by generative AI. Although such examinations are favoured for their perceived security, they have many pedagogical drawbacks (French et al., 2024). Options for addressing these pedagogical drawbacks include replacing, re-weighting or redesigning examinations (Mulder & French, 2023).
When replacing examinations is not possible or feasible, substantial improvements can be made to examinations through thoughtful redesign. In this guide, we suggest practical, evidence-based strategies that can help educators avoid one or more common deficiencies in exam design and implementation. The guide is organised around four main themes: 1) ensuring validity and reliability; 2) enhancing authenticity and real-world relevance; 3) promoting fairness and equity; and 4) reducing unnecessary stress and anxiety.
Ensuring validity and reliability
Do our exams really measure what we intend, and do they do so consistently?
For an exam to be effective, it must be both valid and reliable. Validity is about whether an examination is actually measuring the knowledge and skills it is intended to measure (so that we can draw meaningful conclusions about student performance), while reliability refers to the consistency of the examination results (i.e. whether the same student would obtain similar results under comparable conditions, such as a different format of the same exam, or with different markers). These concepts are important because an exam that is inconsistent (unreliable) cannot support trustworthy conclusions, while one that does not measure the right construct (invalid) cannot meaningfully represent student learning, regardless of how consistent it is (Fig 1).
Figure 1. Validity and reliability illustrated via the metaphor of a target.

While there is no single tool that performs a comprehensive validity or reliability check for university exams, several user-friendly tools and practical frameworks are available that staff can use in complementary ways to evaluate and strengthen their assessments. These can be used to support two key stages in the exam workflow: validity checking before delivery; and analysis and refinement of the exam after delivery (Fig 2).
Figure 2. Steps in validity and reliability checking before and after exam delivery.

Before delivery
Click on the tabs below for validity and reliability checking strategies before exam delivery.
-
Creating an exam blueprint is an excellent way to ensure that the exam is a representative sample of the course learning outcomes and content domains (i.e. to ensure that it has high content validity). The blueprint helps to ensure that questions are intentionally sampled, rather than being selected ad hoc, or from a potentially biased subset.
Blueprints can be useful not only for designing new exams from scratch, but also for evaluating and updating existing exams, or creating multiple versions of the same exam that are of comparable difficulty.
Exam blueprints generally use a simple matrix format that allows the examiner to visualise how questions are distributed with respect to intended learning outcomes, levels of cognitive difficulty, and the relative weighting or importance of different areas of content. A hypothetical example is shown below (Fig 3).
Figure 3. Example exam blueprint

Ensuring that questions have varying levels of cognitive difficulty helps ensure that there is not over-emphasis on simple ‘one right answer’ questions, which focus on factual recall and encourage ‘surface learning’. Questions that require higher cognitive skill levels not only lead to deeper learning and improve knowledge retention (McConnell et al., 2015) but also improve validity by providing a more meaningful picture of a students’ learning. For example, questions that ask students to evaluate, synthesise, analyse and apply knowledge better distinguish understanding than those that emphasise remembering or understanding. In practice, this might mean using scenario-based short-answers or multi-step problems that require reasoning. An example can be found here.
The following online resources are useful for help with exam blueprinting:
- A practical guide to test blueprinting (Raymond & Grande, 2019) – explains blueprinting as part of validity evidence, using a 4-stage development process.
- General Test Construction (Canvas Course)– a freely accessible Canvas subject from Weber State University that helps educators plan, design and implement exams. A useful starting point for those new to exam design.
- Sample multiple choice questions that test higher order thinking – examples spanning economics, physics, psychology, ethics, and philosophy (Kimberley Green, Washington State University).
- Exam questions: types, characteristics and suggestions - a University of Waterloo guide to seven common exam question types, with advice on their strengths and limitations, and suggestions for effective design.
-
Exam blueprinting can strengthen content validity and ensure an appropriate distribution of cognitive demand across an exam. However, validity also depends on whether questions elicit the intended knowledge and capabilities without introducing ambiguity, unintended cues, construct-irrelevant difficulty, or avoidable barriers for students.
Academic colleagues are a valuable and sometimes under-used resource for evaluating these aspects, bringing additional expert judgement and insight into how students might interpret and respond to questions. Peer review can further strengthen validity by confirming that questions are clear, aligned with the intended learning outcomes, appropriate for the stage of learning, and fair.
The format of academic peer review can range from informal discussion with a colleague to structured moderation involving several academics or a formal committee. More formal processes may be used to review the exam as a whole, resolve differences in judgement, and to collectively agree on the expected standard and the quality of student responses.
After delivery
Click on the tabs below for validity and reliability checking strategies after exam delivery.
-
When assessments involve multiple markers, some variation in scoring is almost inevitable, even when a rubric or marking guide is provided. Differences can arise from marker experience, disciplinary or grading norms, and differing interpretations of criteria; less experienced markers may also apply marking guides more rigidly than intended. Clear rubrics and sufficiently detailed marking guides can reduce this variation and improve inter-rater reliability. In particular, marking guides should recognise multiple valid ways of demonstrating understanding rather than relying too closely on a single model answer. For example, a 10-mark question might identify more than 10 distinct elements for which marks could reasonably be awarded.
Moderation and calibration provide an additional safeguard. Where markers are responsible for different subsets of students, all markers can independently score a common sample of responses, compare their judgements, and discuss the reasons for any differences before substantial marking begins. An early check after the first 5-10% of exam scripts have been marked can identify emerging differences in interpretation or severity while they can still be corrected. If systematic differences persist, they can be addressed through further calibration, adjustment of individual marking practice, or where justified, moderation of scores to improve comparability across markers.
-
Statistical analysis of student performance on exam questions can be helpful for identifying problems with exam items (Rao et al., 2016). Such analyses are particularly useful for examinations that involve multiple choice questions (MCQs) and should be routinely conducted after each delivery of an exam.
Various statistical outputs can provide useful information about the reliability and internal functioning of an exam. Cronbach’s alpha provides an estimate of internal consistency, while item difficulty, discrimination indices and point-biserial correlations can help identify questions that may not be functioning as intended. For example, a question that is answered incorrectly by many students who otherwise perform well on the exam overall may may have low or negative discrimination, suggesting that it warrants closer scrutiny for ambiguity, misleading cues, incorrect scoring, or other problems.
These statistics should be treated as diagnostic indicators, rather than automatic grounds for removing an item. Reviewing flagged questions, and where appropriate, revising or replacing them can improve the quality, validity and reliability of subsequent examinations. The resources below provide accessible introductions to conducting item analyses and interpreting their outputs.
- Understanding Item Analysis (University of Washington) – an introductory resource that explains item difficulty, discrimination, response frequencies/distractors and reliability in accessible language.
- Canvas quiz and item analysis report – a practical guide to analyses that can be performed on quizzes delivered on Canvas.
- Outside of the LMS environment here are now a range of accessible tools that support item analysis, from free browser-based R/Shiny applications such as ShinyItemAnalysis to applications that use AI-assisted reporting to create items or identify and interpret poorly functioning items.
Enhancing authenticity and real-world relevance
Are closed-book exams reflective of real-world skills and contexts?
Traditional closed-book exams are often criticised as lacking authenticity, because they tend to present decontextualised problems under artificial conditions (isolated, time-pressured, without access to resources) that students rarely encounter outside academia. Students learn more and demonstrate higher-order skills when assessments mirror real-life tasks and challenges (Villarroel et al., 2019). Students are also more likely to be intrinsically motivated when they can see the real-world relevance and application of their assessment tasks (James & Casidy, 2018). Closed-book examinations can be redesigned in ways that are more authentic by incorporating realistic scenarios and applied problems that require students to exercise evaluative judgement, and to arrive at solutions through dialogue and debate. Click on the tabs below for more details.-
Consider framing exam questions around practical or professional scenarios that are relevant to the discipline. For instance, in a business subject, instead of a generic question about marketing theory, a student could be presented with a brief case of a company facing a sales decline and be invited to analyse the situation and propose a marketing strategy, drawing on principles they have learned in the subject. Scenario-based questions improve the authenticity of the exam and better assess the students’ ability to transfer knowledge to new problems. To maintain realism but ensure students are genuinely applying knowledge, such scenarios should be unfamiliar, but plausible.
-
Assessing quality and making judgements is part and parcel of professional life. While closed-book, time-limited conditions limit authenticity to some degree (students are unable to do lengthy research), exams can nevertheless encourage professional reasoning. Providing students with something to interpret (a dataset, a piece of text, an artwork, a legal case, a lab result, etc) provides them with a more authentic context to demonstrate understanding than an abstract question might. Having to analyse or critique provided material during an exam mimics real-life tasks (interpreting a lab result; debating a legal strategy; engaging in a public debate around a particular topic) and moves beyond pure recall. It requires students to draw on theoretical concepts they may have learned in class and apply them to a concrete artefact or problem, which is more authentic than recalling theory in isolation (Villarroel et al., 2019).
Authentic tasks of this nature often have no single right answer. For instance, students in an architecture subject might be provided with a design challenge that involves working with a set of constraints (space, cost, etc). An exam question might include a worked solution that includes some flaws, which students are then invited to identify and improve on. Such tasks require students to analyse, exercise judgement, make decisions and consider compromises. This more closely mirrors workplace practice than rote solutions because it requires higher-order thinking and may culminate in a range of possible outcomes.
-
In the workplace, it is common practice for individuals to first prepare work on their own and then collaborate with others. Two-stage exams can enhance the authenticity of closed-book assessment by introducing the dialogic context of a workplace into a task that is typically purely an individual exercise. In a two-stage exam, students first complete the exam individually. Then, in a second stage, they re-engage with the same or related questions collaboratively, discussing their reasoning, and debating before settling on a consensus answer that is submitted on behalf of the group.
This two-stage structure not only improves conceptual understanding and reduces exam-related anxiety (Gilley & Clarkston, 2014; Leight et al., 2012) but also provides opportunities for dialogue, argumentation, negotiation and peer feedback. In this way, two-stage exams foster the kinds of communication, critical dialogue, negotiation and peer evaluation that are central to authentic professional practice, while retaining the security and comparability that is valued in summative testing. Collaborative testing furthermore lowers students’ actual (Russo & Warren, 1999) and perceived (LoGiudice et al., 2015) anxiety, and provides opportunities for developing collective knowledge, which can support fair assessment, especially for marginalised student groups (Darabi Bazvand & Rasooli, 2022). This case study is instructive.
-
Involving external stakeholders (e.g. industry, employers, practitioners) in assessment design or review can help align academic assessment with professional expectations (Chiang, 2021). Industry partners can help develop prompts that reflect authentic workplace problems, identify competencies and knowledge most relevant to contemporary practice, and review test items for realism, relevance and suitability, thereby strengthening the exam’s external validity.
Involving students in assessment design can also help them to feel more invested, motivated and engaged (Smith et al., 2025), while modelling professional practice in which graduates actively contribute ideas rather than simply respond to predetermined tasks. Co-design may involve students developing learning objectives, designing rubrics or marking criteria, or writing pools of exam questions for peers. It can also strengthen relationships between staff and students, foster collaboration and promotes a more student-centric and inclusive pedagogy (Ahn & Class, 2011).
Promoting fairness and equity
Do our exams provide an equal opportunity for all to succeed?
Traditional exams can inadvertently favour or disadvantage certain groups of students, raising concerns about fairness and equity (French et al., 2024). Timed, high-pressure exams may disadvantage students from underrepresented or marginalized groups. Gender gaps have also been documented (Ballen et al., 2017) and students for whom English is an additional language often underperform on essay exams relative to native speakers, due to language barriers or lack of familiarity with the testing culture. Socio-economic factors, race/ethnicity, and disability can also intersect with exam performance differences (Richardson, 2015; Tai et al., 2022), and it is widely argued that exams disadvantage Indigenous students (Klenowski, 2009; Preston & Claypool, 2021; Trumbull & Nelson-Barber, 2019). Our goal as educators is to design exams that minimise barriers, so that grades reflect learning, not a student’s background or test-taking savvy. A series of strategies can be implemented at each of the phases of exam design, delivery and grading to help reduce inequities and promote inclusivity and fairness. Click on the tabs below for more details.-
Exam questions that privilege students with higher levels of education history and acculturated learning can have adverse impacts on those from marginalised backgrounds. Tests focussing on prior and acquired knowledge tend to result in greater differences in performance between groups than those that focus on working memory, goal-directed outcomes, reasoning and novel problem solving (Burgoyne et al., 2021). Therefore, designing questions that ask students to solve problems and apply reasoning could help mitigate potential differences in performance for some equity groups. Similarly, assumed cultural knowledge can confuse or bias certain student groups. Avoid unnecessary language complexity, colloquialisms, idioms, or contexts that require cultural knowledge unrelated to the learning outcomes. If technical terminology is necessary, ensure that it is defined or has been taught. The aim is to check for unintended difficulty that does not stem from the subject matter itself.
-
While approved Academic Adjustment Plans (for example, extra time, separate quiet spaces or pods, assistive technologies) can reduce barriers for individual students, such accommodations have limits. They are typically reactive and depend on students recognising that they need support and asking for it, placing the burden of addressing inequity on the student. Many students may not seek adjustments or realise they are available. A more proactive approach is to design examinations to be as inclusive as possible from the outset. Universal Design for Learning (UDL) provides one framework for doing this, by anticipating learner diversity rather than responding to it only after barriers arise (Thompson et al., 2024). Applied across an assessment program, UDL encourages multiple opportunities and modes through which students can demonstrate learning. Within examinations specifically, inclusive designs can include clear and accessible instructions, avoidance of unnecessary linguistic or procedural complexity, and the use of varied questions and formats where these are compatible with the intended learning outcomes.
Traditional exams rely heavily on reading and written responses, which can create additional barriers for some students, including those from language backgrounds other than English and, in some contexts, Indigenous students (International Test Commission (ITC), 2019; Preston & Claypool, 2021). Where reading or writing proficiency is not itself the construct being assessed, unnecessary dependence on these skills can introduce ‘construct-irrelevant difficulty’. Allowing greater variety in question and response formats may permit students to demonstrate knowledge and reasoning in different ways without lowering academic standards. Although the range of formats possible in a timed examination is necessarily constrained, multimodal elements can still be incorporated – for example, by asking students to analyse an image, diagram, artwork, audio excerpt or video. Where practicable and consistent with the learning outcomes, alternative response modes may also provide students with more than one way to demonstrate what they know and can do (e.g. verbal, logical, visual, interpersonal).
-
Exams should assess students’ understanding of content rather than test-taking familiarity or the capacity to perform under time-pressure. Rigid time limits can exacerbate disparities and disadvantage certain groups. For instance, students who process English more slowly, neurodivergent students, those who have anxiety, and those who experience disability may know the material but take longer to provide an answer and require more flexible timing. While not always feasible, it is also good practice to allow short breaks in longer exams (this also mirrors authentic workplace practice). Compassion with respect to timing acknowledges human limits and can be of particular help to students prone to anxiety or panic under pressure.
Exam literacy training can enhance the fairness of exams for marginalised groups by increasing transparency and giving students greater confidence in their understanding of the assessment process. Providing students with resources, support and sufficient opportunities to undertake practice exams is important for supporting equity, and for reducing pre-exam stress and anxiety, as discussed in the following section.
-
Rubric criteria should be shared with students in advance of an exam so that expectations and standards are explicit (Jonsson & Svingby, 2007). Where practical, anonymous or blind marking should also be used, for example by requiring student ID numbers rather than names on examination booklets or by enabling anonymous grading in digital assessment systems. Removing identifying information can reduce the risk of conscious or unconscious biases influencing the grading. Anonymous marking can also reduce ‘halo effects’, where favourable prior impressions of a student affect judgements about their current work (Malouff et al., 2014). However, anonymity involves trade-offs: it can constrain personalised feedback and relationship-building and may make it more difficult for markers to recognise unusual changes in a student’s writing or performance that might otherwise prompt concerns about authorship or contract cheating.
Reducing stress and anxiety
How do we minimise the effect of stress and anxiety on exam performance
Excessive test anxiety can significantly reduce student motivation, negatively impact student wellbeing and impair student performance (Von Der Embse et al., 2018; Wolf & Smith, 1995). Students also differ in how they experience and respond to stress and anxiety because of a range of personal, cultural and biological factors (Zhang et al., 2011). As a result, highly stressful examination conditions may affect some students more than others, creating additional concerns about fairness and equity. Applying UDL and other inclusive design principles can help reduce avoidable sources of anxiety by making expectations, procedures and assessment conditions more predictable and accessible, while also reducing the need for individual accommodations.Click on the tabs below for a range of evidence-based strategies that can make examinations less stressful without compromising academic standards. These include measures taken before the exam to familiarise students with its format and expectations, as well as exam-day practices that promote a calm, predictable and equitable testing environment.
-
Students need explicit guidance about what an exam will require and repeated opportunities to prepare for it. Clearly communicate the exam format, scope, question types and marking expectations well in advance, and provide opportunities for students to ask questions. A short set of FAQs can help ensure that all students receive the same information. Reviewing sample questions and model responses, explaining how marks would be allocated, and providing practice questions or a mock exam can reduce uncertainty and improve performance (Chen et al., 2017).
Frequent, low-stakes quizzes and practice tests can further build confidence, reduce test anxiety and support test-enhanced learning (Yang et al., 2023). Benefits are particularly strong when successive relearning is built into the curriculum, with students repeatedly retrieving important material across spaced practice opportunities (Rawson et al., 2013; Roediger & Karpicke, 2006). Formative testing also helps students identify gaps in understanding, receive feedback and adjust their study strategies before the exam. Where possible, practice should resemble the kinds of questions, reasoning and time pressures students will encounter in the final assessment.
Students can also be encouraged to use evidence-based study strategies such as retrieval practice and peer teaching, and to support their preparation by getting adequate sleep (Smith, 2001) and regular exercise (Engle-Friedman et al., 2003).
-
Some degree of exam anxiety is normal, and acknowledging this without portraying the exam as threatening can help create a more supportive atmosphere. On the day of the exam, small actions such as a welcoming greeting, calm instructions and a positive but professional demeanour can help students settle into the assessment. The physical environment should also minimise unnecessary sources of stress, with comfortable temperature and lighting, minimal noise and distractions, and clear procedures for asking questions or seeking assistance.
Where regulations permit, simple measures such as allowing water bottles or unobtrusive stress-relief items may also make the environment more comfortable. Brief, evidence-based anxiety-reduction activities can be considered where appropriate; for example, a short expressive-writing exercise immediately before a high-stakes test has been shown to improve performance among highly test-anxious students (Ramirez & Beilock, 2011).
-
If appropriate, allowing students to bring permitted materials into an otherwise closed-book exam can ease student nerves. A permitted ‘cheat sheet’ might include key formulas, definitions, diagrams, worked examples, reference tables or other concise materials that support application and reasoning. Students report feeling less anxious when such aids are available (Gharib et al., 2012), and there is no evidence that such allowances inflate scores or undermine learning. Indeed, they are likely to have the positive effect of shifting emphasis in preparation from pure memorisation to application. Even the process of creating a cheat sheet can serve as a valuable learning exercise.
Conclusion
The renewed emphasis on secure assessment creates an important opportunity to reconsider the role of examinations, but it also carries a risk: that security becomes the dominant design principle at the expense of validity, fairness and learning. Exams can provide strong assurance that submitted work reflects a student’s own knowledge and capabilities, but security alone does not make an assessment educationally sound. The approaches outlined in this guide are intended to help assure that, where exams are used, they are designed to be as valid, reliable, authentic, inclusive and equitable as possible, while minimising avoidable sources of anxiety and disadvantage.
At the same time, the push toward secure assessment should not be conflated solely with a return to examinations. Well-designed exams can make an important contribution to a subject’s secure assessment requirement, alongside other approaches such interactive oral assessments, supervised practical tasks or other forms of authenticated performance. The move toward greater security also should not displace open assessment. Open tasks remain critical where students need to research, collaborate, create, use contemporary tools, engage with AI appropriately, and work on complex problems over extended periods.
The aim should therefore be to achieve a balanced assessment regime in which secure and open assessments serve different but complementary purposes. Secure tasks can provide confidence that students have attained essential knowledge and capabilities, while open tasks can support richer learning and assess outcomes that are difficult to capture under constrained conditions. A diverse mix of assessments also allows students to demonstrate a broader range of skills and provides a more defensible and holistic picture of achievement than any single assessment can offer.
Assessment timing and weighting also matter. A secure final examination should not become a single, high-stakes point of potential failure. Formative and lower-stakes tasks distributed throughout a subject can provide students with early feedback, support learning, familiarise them with expectations and reduce the pressure associated with the final assessment.
The current push towards secure assessment should therefore be understood not as a return to exams by default, but as an impetus to use them in a more considered way. The challenge is not simply how to make an assessment secure, but how to design an assessment system that provides credible assurance of learning while remaining valid, inclusive, authentic and educationally worthwhile.
Case studies
References
-
Ahn, R., & Class, M. V. (2011). Student-Centered Pedagogy: Co-Construction of Knowledge through Student-Generated Midterm Exams. The International Journal of Teaching and Learning in Higher Education, 23, 269–281.
Ballen, C. J., Salehi, S., & Cotner, S. (2017). Exams disadvantage women in introductory biology. PLoS ONE, 12(10). https://doi.org/10.1371/journal.pone.0186419
Brown, G. T. L. (2010). The Validity of Examination Essays in Higher Education: Issues and Responses. Higher Education Quarterly, 64(3), 276–291. https://doi.org/10.1111/j.1468-2273.2010.00460.x
Burgoyne, A. P., Mashburn, C. A., & Engle, R. W. (2021). Reducing adverse impact in high-stakes testing. Intelligence, 87, 101561. https://doi.org/10.1016/j.intell.2021.101561
Callaghan, K., Kestin, G., Klales, A., McCarty, L., & Deslauriers, L. (2025). Active learning through flexible collaborative exams: Improving assessments across disciplines. Active Learning in Higher Education, 14697874251344293. https://doi.org/10.1177/14697874251344293
Chiang, W. S. (2021). Involvement of External Stakeholders in Designing Pedagogy for Experiential Learning at University Level: A Case Study. EDUCATUM Journal of Social Sciences, 7(2), 45–56. https://doi.org/10.37134/ejoss.vol7.2.5.2021
Darabi Bazvand, A., & Rasooli, A. (2022). Students’ experiences of fairness in summative assessment: A study in a higher education context. Studies in Educational Evaluation, 72, 101118. https://doi.org/10.1016/j.stueduc.2021.101118
French, S., Dickerson, A., & Mulder, R. A. (2024). A review of the benefits and drawbacks of high-stakes final examinations in higher education. Higher Education, 88(3), 893–918. https://doi.org/10.1007/s10734-023-01148-z
Gharib, Afshin, Phillips, William, & Mathew, Noelle. (2012). Cheat Sheet or Open-Book? A Comparison of the Effects of Exam Types on Performance, Retention, and Anxiety. Journal of Psychology Research, 2(8). https://doi.org/10.17265/2159-5542/2012.08.004
Gilley, B., & Clarkston, B. (2014). Collaborative Testing: Evidence of Learning in a Controlled In-Class Study of Undergraduate Students. Journal of College Science Teaching, 43(3), 83–91. https://doi.org/10.2505/4/jcst14_043_03_83
International Test Commission (ITC). (2019). ITC Guidelines for the Large-Scale Assessment of Linguistically and Culturally Diverse Populations. International Journal of Testing, 19(4), 301–336. https://doi.org/10.1080/15305058.2019.1631024
James, L. T., & Casidy, R. (2018). Authentic assessment in business education: Its effects on student satisfaction and promoting behaviour. Studies in Higher Education, 43(3), 401–415. https://doi.org/10.1080/03075079.2016.1165659
Jonsson, A., & Svingby, G. (2007). The use of scoring rubrics: Reliability, validity and educational consequences. Educational Research Review, 2(2), 130–144. https://doi.org/10.1016/j.edurev.2007.05.002
Klenowski, V. (2009). Australian Indigenous students: Addressing equity issues in assessment. Teaching Education, 20(1), 77–93.
Leight, H., Saunders, C., Calkins, R., & Withers, M. (2012). Collaborative Testing Improves Performance but Not Content Retention in a Large-Enrollment Introductory Biology Class. CBE—Life Sciences Education, 11(4), 392–401. https://doi.org/10.1187/cbe.12-04-0048
LoGiudice, A. B., Pachai, A. A., & Kim, J. A. (2015). Testing together: When do students learn more through collaborative tests? Scholarship of Teaching and Learning in Psychology, 1(4), 377–389. https://doi.org/10.1037/stl0000041
Malouff, J. M., Stein, S. J., Bothma, L. N., Coulter, K., & Emmerton, A. J. (2014). Preventing halo bias in grading the work of university students. Cogent Psychology, 1(1), 988937. https://doi.org/10.1080/23311908.2014.988937
McConnell, M. M., St-Onge, C., & Young, M. E. (2015). The benefits of testing for learning on later performance. ADVANCES IN HEALTH SCIENCES EDUCATION, 20(2), 305–320. https://doi.org/10.1007/s10459-014-9529-1
Mulder, R. A., & French, S. (2023). Reconsidering the role of high-stakes examinations in higher education. 1595021 Bytes. https://doi.org/10.26188/21951287
Preston, J. P., & Claypool, T. R. (2021). Analyzing Assessment Practices for Indigenous Students. Frontiers in Education, 6. https://doi.org/10.3389/feduc.2021.679972
Ramirez, G., & Beilock, S. L. (2011). Writing About Testing Worries Boosts Exam Performance in the Classroom. Science, 331(6014), 211–213. https://doi.org/10.1126/science.1199427
Rao, C., Kishan Prasad, H., Sajitha, K., Permi, H., & Shetty, J. (2016). Item analysis of multiple choice questions: Assessing an assessment tool in medical students. International Journal of Educational and Psychological Researches, 2(4), 201. https://doi.org/10.4103/2395-2296.189670
Rawson, K. A., Dunlosky, J., & Sciartelli, S. M. (2013). The Power of Successive Relearning: Improving Performance on Course Exams and Long-Term Retention. Educational Psychology Review, 25(4), 523–548. https://doi.org/10.1007/s10648-013-9240-4
Raymond, M. R., & Grande, J. P. (2019). A practical guide to test blueprinting. Medical Teacher, 41(8), 854–861. https://doi.org/10.1080/0142159X.2019.1595556
Richardson, J. T. E. (2015). The under-attainment of ethnic minority students in UK higher education: What we know and what we don’t know. Journal of Further and Higher Education, 39(2), 278–291. https://doi.org/10.1080/0309877X.2013.858680
Roediger, H. L., & Karpicke, J. D. (2006). Test-Enhanced Learning: Taking Memory Tests Improves Long-Term Retention. Psychological Science, 17(3), 249–255. https://doi.org/10.1111/j.1467-9280.2006.01693.x
Russo, A., & Warren, S. H. (1999). Collaborative Test Taking. College Teaching, 47(1), 18–20. https://doi.org/10.1080/87567559909596072
Smith, A., McConnell, L., Iyer, P., Allman-Farinelli, M., & Chen, J. (2025). Co-designing assessment tasks with students in tertiary education: A scoping review of the literature. Assessment & Evaluation in Higher Education, 50(2), 199–218. https://doi.org/10.1080/02602938.2024.2376648
Tai, J., Ajjawi, R., Bearman, M., Boud, D., Dawson, P., & Jorre de St Jorre, T. (2022). Assessment for inclusion: Rethinking contemporary strategies in assessment design. Higher Education Research & Development, 1–15. https://doi.org/10.1080/07294360.2022.2057451
Trumbull, E., & Nelson-Barber, S. (2019). The Ongoing Quest for Culturally-Responsive Assessment for Indigenous Students in the U.S. Frontiers in Education, 4. https://doi.org/10.3389/feduc.2019.00040
Villarroel, V., Boud, D., Bloxham, S., Bruna, D., & Bruna, C. (2019). Using principles of authentic assessment to redesign written examinations and tests. Innovations in Education and Teaching International, 1–12. https://doi.org/10.1080/14703297.2018.1564882
Von Der Embse, N., Jester, D., Roy, D., & Post, J. (2018). Test anxiety effects, predictors, and correlates: A 30-year meta-analytic review. Journal of Affective Disorders, 227, 483–493. https://doi.org/10.1016/j.jad.2017.11.048
Wolf, L. F., & Smith, J. K. (1995). The Consequence of Consequence: Motivation, Anxiety, and Test Performance. Applied Measurement in Education, 8(3), 227–242. https://doi.org/10.1207/s15324818ame0803_3
Yang, C., Li, J., Zhao, W., Luo, L., & Shanks, D. R. (2023). Do Practice Tests (Quizzes) Reduce or Provoke Test Anxiety? A Meta-Analytic Review. Educational Psychology Review, 35(3), 87. https://doi.org/10.1007/s10648-023-09801-w
Professor Raoul Mulder, Centre for the Study of Higher Education, and Dr Sarah French, Faculty of Education
Last updated: September 2026