HybridDetect: Detecting Partial AI Contamination in Student Essays via Multi-Class Classification
Daniel Dan, Minoru Nakayama
DOI: http://dx.doi.org/10.15439/2026F3512
Citation: Daniel Dan, Minoru Nakayama (2026). HybridDetect: Detecting Partial AI Contamination in Student Essays via Multi-Class Classification. In M. Bolanowski, M. Ganzha, M. Grzegorowski, L. Maciaszek, M. Paprzycki, A. Paszkiewicz, D. Ślęzak (eds), Proceedings of the 21st Conference on Computer Science and Intelligence Systems (FedCSIS). ACSIS, Vol. 47, pages 507–512.
Abstract. The proliferation of large language models (LLMs) in higher education has created an urgent need for tools that detect AI-generated content. While existing detectors such as GPTZero and Turnitin operate as binary classifiers distinguishing fully human from fully AI-generated text, they fail to address the increasingly common scenario where students partially incorporate AI-generated content into their writing. We introduce HybridDetect, a multi-class classification framework that detects varying levels of AI contamination in student essays across seven granularity levels: 0\%, 10\%, 25\%, 50\%, 75\%, 90\%, and 100\% AI content. Our approach combines a fine-tuned RoBERTa model for semantic analysis with an XGBoost classifier trained on 36 handcrafted linguistic features, integrated through weighted soft voting. We construct a synthetic dataset of 50,001 hybrid essays by blending sentences from approximately 487,000 human-written and 487,000 AI-generated essays at controlled contamination levels. On a held-out test set of 7,501 essays, the ensemble achieves 79.9\% accuracy and 0.798 macro F1-score on the 7-class task, with 0.876 F1-score on binary classification (≥50\% AI). Results reveal that semantic features captured by RoBERTa are essential for fine-grained detection (0.798 macro F1), while surface-level linguistic features alone achieve only 0.394 macro F1, indicating that AI contamination manifests primarily in semantic patterns rather than stylistic markers. These findings have direct implications for academic integrity policies, suggesting that institutions should adopt graduated response frameworks rather than binary pass/fail judgments.
References
- E. Tian, “GPTZero: An AI text detector,” https://gptzero.me, 2023, accessed: 2026-03-01.
- E. Mitchell, Y. Lee, A. Khazatsky, C. D. Manning, and C. Finn, “DetectGPT: Zero-shot machine-generated text detection using probability curvature,” in Proceedings of the 40th International Conference on Machine Learning (ICML), 2023. https://dx.doi.org/10.48550/arXiv.2301.11305 pp. 24 950–24 962.
- D. R. E. Cotton, P. A. Cotton, and J. R. Shipway, “Chatting and cheating: Ensuring academic integrity in the era of ChatGPT,” Innovations in Education and Teaching International, vol. 61, no. 2, pp. 228–239, 2024. https://dx.doi.org/10.1080/14703297.2023.2190148
- J. Kirchenbauer, J. Geiping, Y. Wen, J. Katz, I. Miers, and T. Goldstein, “A watermark for large language models,” in Proceedings of the 40th International Conference on Machine Learning (ICML), 2023. https://dx.doi.org/10.48550/arXiv.2301.10226 pp. 17 061–17 084.
- W. Liang, M. Yuksekgonul, Y. Mao, E. Wu, and J. Zou, “GPT detectors are biased against non-native English writers,” Patterns, vol. 4, no. 7, p. 100779, 2023. https://dx.doi.org/10.1016/j.patter.2023.100779
- V. S. Sadasivan, A. Kumar, S. Balasubramanian, W. Wang, and S. Feizi, “Can AI-generated text be reliably detected?” arXiv preprint https://arxiv.org/abs/2303.11156, 2023. https://dx.doi.org/10.48550/arXiv.2303.11156
- D. Weber-Wulff, A. Anohina-Naumeca, S. Bjelobaba, T. Foltýnek, J. Guerrero-Dib, O. Popoola, P. Šigut, and L. Waddington, “Testing of detection tools for AI-generated text,” International Journal for Educational Integrity, vol. 19, no. 26, 2023. https://dx.doi.org/10.1007/s40979-02300146-z
- P. Gryka, K. Gradoń, M. Kozłowski, M. Kutyła, and A. Janicki, “Impact of spelling and editing correctness on detection of LLM-generated emails,” in Proceedings of the 19th Conference on Computer Science and Intelligence Systems (FedCSIS), ser. Annals of Computer Science and Information Systems, vol. 39, 2024. https://dx.doi.org/10.15439/2024F8906 pp. 603–608.
- Turnitin, “AI writing detection,” https://www.turnitin.com/solutions/ ai-writing, 2023, accessed: 2026-04-01.
- Winston AI, “AI content detection for education,” https://gowinston.ai, 2024, accessed: 2026-04-01.
- Copyleaks, “AI content detector,” https://copyleaks.com/ ai-content-detector, 2024, accessed: 2026-04-01.
- L. Dugan, D. Ippolito, A. Kirubarajan, S. Shi, and C. Callison-Burch, “Real or fake text?: Investigating human ability to detect boundaries between human-written and machine-generated text,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 11, 2023. https://dx.doi.org/10.1609/aaai.v37i11.26501 pp. 12 763–12 771.
- J. W. Cutler, L. Dugan, S. Havaldar, and A. Stein, “Automatic detection of hybrid human-machine text boundaries,” University of Pennsylvania, CIS 520 Course Project, 2021, available: https://www.seas.upenn.edu/ ~steinad/papers/cis520_project.pdf.
- A. Uchendu, T. Le, and D. Lee, “Attribution and obfuscation of neural text authorship: A data mining perspective,” ACM SIGKDD Explorations Newsletter, vol. 25, no. 1, pp. 1–18, 2023. https://dx.doi.org/10.1145/3606274.3606276
- D. Baidoo-Anu and L. O. Ansah, “Education in the era of generative artificial intelligence (AI): Understanding the potential benefits of ChatGPT in promoting teaching and learning,” Journal of AI, vol. 7, no. 1, pp. 52–62, 2023. https://dx.doi.org/10.61969/jai.1337500
- A. Smolansky, A. Cram, C. Raduescu, S. Zeivots, E. Huber, and R. F. Kizilcec, “Educator and student perspectives on the impact of generative AI on assessments in higher education,” in Proceedings of the Tenth ACM Conference on Learning @ Scale (L@S), 2023. https://dx.doi.org/10.1145/3573051.3596191 pp. 378–382.
- M. Perkins, L. Furze, J. Roe, and J. MacVaugh, “The artificial intelligence assessment scale (AIAS): A framework for ethical integration of generative AI in educational assessment,” Journal of University Teaching and Learning Practice, vol. 21, no. 6, 2024. https://dx.doi.org/10.53761/q3azde36
- D. Dan, A. Wróblewska, B. Grabek, M. Taczała, and M. Nakayama, “Applications and challenges of artificial intelligence in educational course design and delivery,” in Proceedings of the 20th Conference on Computer Science and Intelligence Systems (FedCSIS), ser. Annals of Computer Science and Information Systems, vol. 43, 2025. https://dx.doi.org/10.15439/2025F8073 pp. 675–680.
- M. Ejaz, “AI vs human text dataset,” Kaggle, 2024, https://www.kaggle. com/datasets/muqaddasejaz/ai-vs-human-text-dataset.
- B. Guo, X. Zhang, Z. Wang, M. Jiang, J. Nie, Y. Ding, J. Yue, and Y. Wu, “How close is ChatGPT to human experts? comparison corpus, evaluation, and detection,” arXiv preprint arXiv:2301.07597, 2023. https://dx.doi.org/10.48550/arXiv.2301.07597
- Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov, “RoBERTa: A robustly optimized BERT pretraining approach,” arXiv preprint arXiv:1907.11692, 2019. https://dx.doi.org/10.48550/arXiv.1907.11692
- T. Chen and C. Guestrin, “XGBoost: A scalable tree boosting system,” in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2016. https://dx.doi.org/10.1145/2939672.2939785 pp. 785–794.
- T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. L. Scao, S. Gugger, M. Drame, Q. Lhoest, and A. M. Rush, “Transformers: State-of-the-art natural language processing,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 2020. https://dx.doi.org/10.18653/v1/2020.emnlp-demos.6 pp. 38–45.