LLM-Assisted assessment in GIS education
An empirical evaluation of opportunities and risks
Bibliographic Data
| ID | 22415136 |
|---|---|
| Authors | Kamyar Hasanzadeh (0000-0002-0705-7662, University of Helsinki, corresponding author), Anna Saarinen (University of Helsinki) |
| Year | 2026 |
| Pages | 1-13 |
| Publication date | 2026-07-04 |
| Peer Reviewed | Yes |
| Open Access | Yes |
| Type | ARTICLE |
| Venue | Journal of Geography in Higher Education (JOURNAL) |
| Journal identifiers | ISSN: 0309-8265 • E-ISSN: 1466-1845 |
| Publisher | Informa UK Limited (PUBLISHER • GB) |
| DOI | 10.1080/03098265.2026.2699176 |
| OpenAlex | W7167353369 |
| Language | EN |
| References cited | 31 |
Assessment in Geographic Information Systems (GIS) education is complex, as student work combines technical implementation, spatial reasoning, interpretation, and visual communication. While large language models (LLMs) are increasingly discussed in higher education, empirical evidence on their role in assessment remains limited. This study addresses a sensitive issue many instructors have considered yet remains cautiously discussed: whether LLMs can responsibly support grading. Using a comparative design, we analyzed manual and LLM-based grading across three master-level courses: programming-based spatial analysis, interpretative GIS, and cartography. Identical course-specific rubrics were provided to an instructor and the LLM. Outcomes were compared in terms of grades, grading time, and feedback length, alongside qualitative analysis of feedback structure. Results show high agreement in relative student rankings, although alignment varied across course types. Agreement in absolute grades was closer in interpretative courses, while larger differences emerged in programming tasks, where the LLM graded more strictly. Feedback was generally more detailed, and model execution time per submission was substantially lower than manual grading time. The findings support a constrained, rubric-driven role for LLMs as structured assessment support within instructor-led workflows. Responsible implementation requires transparency, instructor oversight, and reproducible evaluation procedures. Further research is needed before broader adoption in education.
Data collection · Empirical research · Evaluation methods · Geographic information system · Risk assessment · Educational Assessment and Pedagogy · Geography Education and Pedagogy · Mathematics Education and Programs
A Conversation on Artificial Intelligence, Chatbots, and Plagiarism in Higher Education
Developing evaluative judgement
A scoping review on how generative artificial intelligence transforms assessment in higher education
Assessment and Classroom Learning
Formative assessment and the design of instructional systems
ChatGPT for good? On opportunities and challenges of large language models for education
ChatGeoAI
Generative AI in Undergraduate Education
Assessing ChatGPT for GIS education and assignment creation
Assessment and Evaluation of GIScience Curriculum using the Geographic Information Science and Technology Body of Knowledge
Advancements in artificial intelligence and the reframing of student assessment in geography education
Referencing, Internalizing, or Activating Knowledge
Mark my words
| Citation velocity | historical |
|---|---|
| Highly cited | No |