Logo PTI Logo FedCSIS

Proceedings of the 21st Conference on Computer Science and Intelligence Systems (FedCSIS)

Annals of Computer Science and Information Systems, Volume 47

Generating Quality Word-Association Puzzles

, , ,

DOI: http://dx.doi.org/10.15439/2026F7130

Citation: Ashish Khadka, , ,

Full text

Abstract. Word-association puzzles require players to identify a hidden relationship or a connection between a group of keywords. Generating such puzzles is possible with proprietary/frontier LLM models, probably due to their reasoning capacities, but we don't have many insights on how quality puzzles could be generated with limited resources. Open-source instruction-tuned language models show limited capability as they only generate poor quality and repetitive puzzles. In this paper, we explore how an open-source language model (i.e., Llama 3.1-8B-Instruct, Mistral-7B-Instruct-v0.3, and Qwen2.5-7B-Instruct) could be leveraged to generate high-quality and diverse puzzles. We demonstrate a method based on Supervised Fine-Tuning (SFT) from a teacher model (i.e., Claude) for generating quality word association puzzles. We validate the puzzles and assess their diversity using OpenAI's GPT-5.2 Thinking as a judge. Our findings show a low-cost and efficient method for distilling a lateral thinking skill from a large reasoning model.

References

  1. A. Khadka, M. Nassar, and S. Khare, “Are language models good at lateral thinking?” in 2025 3rd International Conference on Foundation and Large Language Models (FLLM). IEEE, 2025, pp. 565–572.
  2. I. Arcuschin, J. Janiak, R. Krzyzanowski, S. Rajamanoharan, N. Nanda, and A. Conmy, “Chain-of-thought reasoning in the wild is not always faithful,” arXiv preprint https://arxiv.org/abs/2503.08679, 2025.
  3. E. De Bono, Lateral Thinking: Creativity Step by Step, ser. Harper colophon books. Harper & Row, 1970. [Online]. Available: https://books.google.com/books?id= H-ROAAAAMAAJ
  4. LinkedIn Games, “Pinpoint,” 2024, https://www.linkedin. com/games/pinpoint/. Accessed: 2025-01-21.
  5. Anthropic, “Claude opus 4.6,” 2025, https://www. anthropic.com/news/claude-opus-4-6. Accessed: 202603-27.
  6. G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” arXiv preprint arXiv:1503.02531, 2015.
  7. M. F. Maleki and R. Zhao, “Procedural content generation in games: A survey with insights on emerging llm integration,” in Proceedings of the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment, vol. 20, no. 1, 2024, pp. 167–178.
  8. T. Merino, S. Earle, R. Sudhakaran, S. Sudhakaran, and J. Togelius, “Making new connections: Llms as puzzle generators for the new york times’ connections word game,” in Proceedings of the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment, vol. 20, 2024, pp. 87–96.
  9. S. Sudhakaran, M. González-Duque, M. Freiberger, C. Glanois, E. Najarro, and S. Risi, “Mariogpt: Openended text2level generation through large language models,” Advances in Neural Information Processing Systems, vol. 36, pp. 54 213–54 227, 2023.
  10. X. Xu, M. Li, C. Tao, T. Shen, R. Cheng, J. Li, C. Xu, D. Tao, and T. Zhou, “A survey on knowledge distillation of large language models,” arXiv preprint arXiv:2402.13116, 2024.
  11. A. Gudibande, E. Wallace, C. Snell, X. Geng, H. Liu, P. Abbeel, S. Levine, and D. Song, “The false promise of imitating proprietary llms,” arXiv preprint arXiv:2305.15717, 2023.
  12. T. Hagendorff, S. Fabi, and M. Kosinski, “Human-like intuitive behavior and reasoning biases emerged in large language models but disappeared in chatgpt,” Nature Computational Science, vol. 3, no. 10, pp. 833–838, 2023.
  13. L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing et al., “Judging llm-as-a-judge with mt-bench and chatbot arena,” Advances in neural information processing systems, vol. 36, pp. 46 595–46 623, 2023.
  14. L. Li, L. Sleem, Y. Xu, Y. Song, A. Jia, J. Francois, and R. State, “The necessity of setting temperature in llm-asa-judge,” arXiv preprint arXiv:2603.28304, 2026.
  15. A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan et al., “The llama 3 herd of models,” arXiv preprint arXiv:2407.21783, 2024.
  16. A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M.A. Lachaux, P. Stock, T. Le Scao, T. Lavril, T. Wang, T. Lacroix, and W. El Sayed, “Mistral 7b,” arXiv preprint arXiv:2310.06825, 2023.
  17. B. Hui, J. Yang, Z. Cui, J. Yang, D. Liu, L. Zhang, T. Liu, J. Zhang, B. Yu, K. Lu et al., “Qwen2.5-coder technical report,” arXiv preprint arXiv:2409.12186, 2024.
  18. A. Singh, A. Fry, A. Perelman, A. Tart, A. Ganesh, A. ElKishky, A. McLaughlin, A. Low, A. Ostrow, A. Ananthram et al., “Openai gpt-5 system card,” arXiv preprint arXiv:2601.03267, 2025.
  19. N. Aspert, V. Miz, B. Ricaud, and P. Vandergheynst, “A graph-structured dataset for wikipedia research,” in Companion Proceedings of The 2019 World Wide Web Conference, 2019, pp. 1188–1193.
  20. L. von Werra, Y. Belkada, L. Tunstall, E. Beeching, T. Thrush, N. Lambert, S. Huang, K. Rasul, and Q. Gallouédec, “TRL: Transformers Reinforcement Learning,” 2020, https://github.com/huggingface/trl.
  21. E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” in International Conference on Learning Representations, 2022. [Online]. Available: https://openreview.net/forum?id=nZeVKeeFYf9
  22. L. Chen, G. Varoquaux, and F. M. Suchanek, “Learning high-quality and general-purpose phrase representations,” arXiv preprint arXiv:2401.10407, 2024.
  23. I. Shumailov, Z. Shumaylov, Y. Zhao, N. Papernot, R. Anderson, and Y. Gal, “Ai models collapse when trained on recursively generated data,” Nature, vol. 631, no. 8022, pp. 755–759, 2024.