Chat Gpt vs. Notebook Lm: Which Ai Tool Better Generates Mcq Listening Test Items?
Keywords:
AI-assisted language assessment, EFL listening assessment, multiple-choice items, stem clarity, distractor plausibilityAbstract
Abstract
This study compares the quality of multiple-choice listening test items generated by ChatGPT and NotebookLM for EFL classroom-based assessment, focusing on two key dimensions: stem clarity and distractor plausibility. Using a small-scale comparative evaluative design, the study analyzed thirty AI-generated listening items, consisting of fifteen items produced by ChatGPT and fifteen items produced by NotebookLM. The items were generated from three listening transcripts representing different listening contexts and objectives through an identical zero-shot prompting procedure. Each item was evaluated using an analytic rubric covering stem clarity and distractor plausibility, supported by qualitative content analysis of the stems, options, answer keys, and alignment with the source transcripts. The findings show that both tools generated highly clear and focused stems, although ChatGPT performed slightly better in producing concise, direct, and learner-friendly question stems. However, a more notable difference was found in distractor plausibility. NotebookLM produced more plausible, competitive, and transcript-grounded distractors, particularly for speaker intention, main idea, and simple inference items, whereas ChatGPT often produced formally consistent distractors that were sometimes weakly connected to the listening input, easily eliminated through contextual guessing, or affected by answer-key patterns. These findings indicate that ChatGPT and NotebookLM offer complementary strengths in AI-assisted listening item development. ChatGPT may be more useful for drafting clear stems, while NotebookLM may provide stronger support for constructing source-based distractors. Nevertheless, neither tool can be used as a standalone solution for developing valid listening assessments. Human expert review remains essential to ensure construct validity, improve distractor function, remove technical clues, and align AI-generated items with sound principles of EFL listening assessment.