Logo PTI Logo FedCSIS

Proceedings of the 21st Conference on Computer Science and Intelligence Systems (FedCSIS)

Annals of Computer Science and Information Systems, Volume 47

A Real-World Dataset of LLM Jailbreak Prompts from a Capture-the-Flag Challenge

,

DOI: http://dx.doi.org/10.15439/2026F5600

Citation: Jan Czajkowski,

Full text

Abstract. We introduce a real-world dataset for studying prompt injection and jailbreak behavior in Large Language Models, collected during Capture the Flag (CTF) events in which participants attempted to extract a hidden secret string protected by defensive prompts. Unlike benchmark data generated primarily through automated methods, the corpus captures human-authored offensive and defensive prompts produced in a realistic educational and competitive setting. The raw dataset contains roughly 8,000 interaction records; through filtering for successful, non-trivial, and analytically useful cases, we derive a curated subset of 115 examples for detailed study. We analyze both defensive and offensive prompts as interacting linguistic artifacts and show that defensive strategies are relatively concentrated, while offensive prompts display substantially greater structural diversity. Clustering identifies three major families of defensive prompts and eight clusters of offensive prompts, reflecting an asymmetry between reusable protection templates and varied attack strategies. The dataset supports qualitative and quantitative research on prompt injection, including the study of secrecy-preserving prompt design, attack typology, and the security--usability trade-off in LLM-based systems.

References

  1. J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al. GPT-4 technical report. arXiv preprint https://arxiv.org/abs/2303.08774, 2023.
  2. M. Andriushchenko et al. A benchmark for measuring harmfulness of LLM agents. arXiv preprint arXiv:2410.09024, 2024.
  3. T. Chothia and C. Novakovic. An offline capture the {Flag-Style} virtual machine and an assessment of its value for cybersecurity education. In 2015 USENIX Summit on Gaming, Games, and Gamification in Security Education (3GSE 15), 2015.
  4. Gray Swan. Adversarial attacks on aligned language models. https: //www.grayswan.ai/research/adversarial-attacks-on-aligned-language-m odels, 2024. Research overview discussing GCG-style jailbreak attacks. Accessed: 2026-04-15.
  5. Lakera. Who is Gandalf? the AI challenge that tests your prompt injection skills. https://www.lakera.ai/blog/who-is-gandalf, 2023. Accessed: 2026-04-15.
  6. K. Leune and S. J. Petrilli Jr. Using capture-the-flag to enhance the effectiveness of cybersecurity education. In Proceedings of the 18th annual conference on information technology education, pages 47–52, 2017.
  7. Y. Liu, Y. Jia, R. Geng, J. Jia, and N. Z. Gong. Formalizing and benchmarking prompt injection attacks and defenses. In 33rd USENIX Security Symposium (USENIX Security 24), pages 1831–1847. USENIX Association, 2024.
  8. OpenAI. text-embedding-3-small model. https://developers.openai. com/api/docs/models/text-embedding-3-small, 2026. Official model documentation. Accessed: 2026-04-15.
  9. OpenAI. Vector embeddings. https://developers.openai.com/api/do cs/guides/embeddings, 2026. Official embeddings guide. Accessed: 2026-04-15.
  10. N. Pfister, V. Volhejn, M. Knott, S. Arias, J. Bazińska, M. Bichurin, A. Commike, J. Darling, P. Dienes, M. Fiedler, D. Haber, M. Kraft, M. Lancini, M. Mathys, D. Pascual-Ortiz, J. Podolak, A. Romero-López, K. Shiarlis, A. Signer, Z. Terek, A. Theocharis, D. Timbrell, S. Trautwein, S. Watts, Y.-H. Wu, and M. Rojas-Carulla. Gandalf the Red: Adaptive security for LLMs. arXiv preprint arXiv:2501.07927, 2025.
  11. J. Yi, Y. Xie, B. Zhu, E. Kiciman, G. Sun, X. Xie, and F. Wu. Benchmarking and defending against indirect prompt injection attacks on large language models. arXiv preprint arXiv:2312.14197, 2023.
  12. A. Zou, Z. Wang, N. Carlini, M. Nasr, J. Z. Kolter, and M. Fredrikson. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043, 2023.