Plot-and-Ask: Multimodal Local LLMs in Smart Environments for Visual IoT Analytics
Aygün Varol, Asif Shaikh, Naser Hossein Motlagh, Mirka Leino, Johanna Virkki
DOI: http://dx.doi.org/10.15439/2026F4180
Citation: Aygün Varol, Asif Shaikh, Naser Hossein Motlagh, Mirka Leino, Johanna Virkki (2026). Plot-and-Ask: Multimodal Local LLMs in Smart Environments for Visual IoT Analytics. In M. Bolanowski, M. Ganzha, M. Grzegorowski, L. Maciaszek, M. Paprzycki, A. Paszkiewicz, D. Ślęzak (eds), Proceedings of the 21st Conference on Computer Science and Intelligence Systems (FedCSIS). ACSIS, Vol. 47, pages 609–614.
Abstract. Smart environments generate continuous heterogeneous sensor data that require real-time processing to support occupant services and decision-making. Existing approaches either compromise data sovereignty through cloud dependencies or rely on rigid rule-based systems that lack natural-language interaction. This paper presents a local multimodal analytics architecture for smart indoor environments that combines edge-resident multimodal large language models with a knowledge graph policy layer. The proposed plot-and-ask workflow couples automatic time-series visualization with vision-enabled language interpretation, enabling non-expert users to explore indoor sensing data without programming expertise. The system supports natural-language-to-SQL querying, plot generation with knowledge-graph-derived comfort bands, and multimodal plot interpretation while keeping all processing on local hardware. We evaluate the framework in a two-week indoor air quality deployment across office, hallway, and kitchen locations using four local models: Gemma 3 4B, Ministral 3 3B, Qwen 3 4B, and Granite 3 2B. All four models achieved 100\% accuracy for structured query synthesis and plot generation. For plot interpretation, Ministral 3 3B and Gemma 3 4B provided the best balance between accuracy and latency, reaching 100\% and 96.67\% accuracy with 5.84 s and 5.34 s end-to-end latency, respectively. A knowledge graph grounding ablation showed the largest gain for Qwen 3 4B, with a 25.0 percentage-point improvement and latency reduced from 27.98 s to 15.21 s. A supplementary indoor-scene safety check also shows that the same models can support direct image understanding. Together, these results show that plot-based summarization and lightweight knowledge graph grounding improve local multimodal IoT analytics under edge constraints. Overall, hybrid knowledge-graph and multimodal-language-model architectures can support privacy-preserving visual IoT analytics at the edge while enabling interactive exploration by non-expert users.
References
- N. H. Motlagh, M. A. Zaidan, L. Lovén, P. L. Fung, T. Hänninen, R. Morabito, P. Nurmi, and S. Tarkoma, “Digital twins for smart spaces—beyond IoT analytics,” IEEE Internet of Things Journal, vol. 11, no. 1, pp. 573–583, 2023.
- A. Varol, N. H. Motlagh, M. Leino, S. Tarkoma, and J. Virkki, “Creation of AI-driven smart spaces for enhanced indoor environments—a survey,” Internet of Things, p. 101876, 2026.
- M. Tawfik, A. H. Abdelhaliem, and I. S. Fathi, “Transforming IoT security through large language models: A comprehensive systematic review and future directions,” Statistics, Optimization & Information Computing, vol. 14, no. 2, pp. 1018–1044, 2025.
- A. Sangwan, P. Sen, A. Singh, S. Zaman, Y. Zhou, and M. Medley, “From cloud to edge: Enabling offline IoT with small-scale language models,” IEEE Internet of Things Magazine, 2025.
- A. Varol, N. H. Motlagh, M. Leino, and J. Virkki, “Performance of large language models across edge and cloud platforms in smart spaces,” in 2025 10th International Conference on Smart and Sustainable Technologies (SpliTech). IEEE, 2025, pp. 1–6.
- K. Mohsenzadegan, V. Tavakkoli, W. V. Kambale, and K. Kyamakya, “A hybrid AI framework integrating ontology learning, knowledge graphs, and large language models for improved data model translation in smart manufacturing and transportation,” in 2024 Sensor Data Fusion: Trends, Solutions, Applications (SDF). IEEE, 2024, pp. 1–6.
- K. Mohsenzadegan, V. Tavakkoli, W. V. Kambale, and K. Kyamakya, “Towards seamless data translation based on data models: A hybrid AI framework for smart transportation and manufacturing,” in 2024 19th International Workshop on Semantic and Social Media Adaptation & Personalization (SMAP). IEEE, 2024, pp. 68–73.
- C. Zhou and J. Yang, “HoloLLM: Multisensory foundation model for language-grounded human sensing and reasoning,” arXiv preprint https://arxiv.org/abs/2505.17645, 2025.
- P. P. Ray and M. P. Pradhan, “LLMEdge: A novel framework for localized LLM inferencing at resource-constrained edge,” in 2024 International Conference on IoT Based Control Networks and Intelligent Systems (ICICNIS). IEEE, 2024, pp. 1–8.
- M. Zhang, X. Shen, J. Cao, Z. Cui, and S. Jiang, “EdgeShard: Efficient LLM inference via collaborative edge computing,” IEEE Internet of Things Journal, vol. 12, no. 10, pp. 13 119–13 131, 2024.
- A. Varol, K. Kołodziej, Łukasz Sobczak, M. Romaszewski, P. Głomb, N. H. Motlagh, M. Leino, and J. Virkki, “Enabling cloud-level accuracy in edge ai through iot data preprocessing,” 2026. [Online]. Available: https://arxiv.org/abs/2606.22496
- G. M. Yilma, J. A. Ayala-Romero, A. Garcia-Saavedra, and X. CostaPerez, “TelecomRAG: Taming telecom standards with retrievalaugmented generation and LLMs,” ACM SIGCOMM Computer Communication Review, vol. 54, no. 3, pp. 18–23, 2025.
- M. Klapperstueck, T. Czauderna, C. Goncu, J. Glowacki, T. Dwyer, F. Schreiber, and K. Marriott, “Contextuwall: Multi-site collaboration using display walls,” Journal of Visual Languages & Computing, vol. 46, pp. 35–42, 2018.
- M. K. Konkel, B. Ullmer, O. Shaer, and A. Mazalek, “Envisioning tangibles and display-rich interfaces for co-located and distributed genomics collaborations,” in Proceedings of the 8th ACM International Symposium on Pervasive Displays, 2019, pp. 1–8.
- S. K. Badam, E. Fisher, and N. Elmqvist, “Munin: A peer-to-peer middleware for ubiquitous analytics and visualization spaces,” IEEE Transactions on Visualization and Computer Graphics, vol. 21, no. 2, pp. 215–228, 2014.
- Gemma Team, “Gemma 3,” Technical report, Google DeepMind, 2025, https://goo.gle/Gemma3Report.
- A. H. Liu, K. Khandelwal et al., “Ministral 3,” 2026. [Online]. Available: https://arxiv.org/abs/2601.08584
- S. Bai, Y. Cai et al., “Qwen3-VL technical report,” arXiv preprint arXiv:2511.21631, 2025.
- IBM Granite Team, “Granite 3.0 language models,” October 2024. [Online]. Available: https://github.com/ibm-granite/granite-3. 0-language-models/
- C. Tian, X. Qin, K. Tam, L. Li, Z. Wang, Y. Zhao, M. Zhang, and C. Xu, “CLONE: Customizing LLMs for efficient latency-aware inference at the edge,” arXiv preprint arXiv:2506.02847, 2025.
- S. S. Kim, Q. V. Liao, M. Vorvoreanu, S. Ballard, and J. W. Vaughan, “” i’m not sure, but...”: Examining the impact of large language models’ uncertainty expression on user reliance and trust,” in Proceedings of the 2024 ACM conference on fairness, accountability, and transparency, Rio de Janeiro, Brazil, 2024, pp. 822–835.
- A. Quattoni and A. Torralba, “Recognizing indoor scenes,” in 2009 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2009, pp. 413–420.
- A. K. Putra, “Indoor fire smoke dataset,” Mar. 2025. [Online]. Available: https://doi.org/10.5281/zenodo.15826133
- A. Shaikh, A. Varol, and J. Virkki, “From prompts to motors: Man-inthe-middle attacks on llm-enabled vacuum robots,” IEEE Access, 2025.