Logo PTI Logo FedCSIS

Communication Papers of the 21st Conference on Computer Science and Intelligence Systems (FedCSIS)

Annals of Computer Science and Information Systems, Volume 49

Parsing Hierarchical Document Structure with Graph Neural Networks

, , , , ,

DOI: http://dx.doi.org/10.15439/2026F3376

Citation: Thomas Reiser, , , , ,

Full text

Abstract. Document hierarchy parsing is used to identify hierarchical relationships between layout elements in documents, such as headings, paragraphs, and list items. This information makes documents that are not born-digital easier to navigate for human readers and can support automated processing of these documents. Similar to existing approaches, we insert a sequence of nodes that represent layout elements into a layout tree to incrementally reconstruct the original document hierarchy. Different to existing research, our approach uses a GraphSAGE-based feature extraction network and a deep neural network based scorer to predict parent relationships between layout element pairs using text and layout features. Our data set contains 50 historical German documents and we apply k-fold cross validation with five splits to validate the model's ability to generalize from a small data set to unseen data. The model was able to obtain an average F1-score of over 91\\% on the selected data set.

References

  1. H. Xing, C. Cheng, F. Gao, Z. Shao, Z. Yu, J. Bu, Q. Zheng, and C. Yao, “DocHieNet: A large and diverse dataset for document hierarchy parsing,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y.-N. Chen, Eds. Miami, Florida, USA: Association for Computational Linguistics, Nov. 2024. https://dx.doi.org/10.18653/v1/2024.emnlp-main.65 pp. 1129–1142. [Online]. Available: https://aclanthology.org/2024.emnlp-main.65/
  2. Z. Zhao, H. Kang, B. Wang, and C. He, “Doclayoutyolo: Enhancing document layout analysis through diverse synthetic data and global-to-local adaptive perception,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12628
  3. R.-Y. Cao, Y.-X. Cao, G.-B. Zhou, and P. Luo, “Extracting variable-depth logical document hierarchy from long documents: Method, evaluation, and application,” Journal of Computer Science and Technology, vol. 37, no. 3, pp. 699–718, 2022. https://dx.doi.org/10.1007/s11390-021-1076-7. [Online]. Available: https://doi.org/10.1007/s11390-021-1076-7
  4. J. Wang, K. Hu, Z. Zhong, L. Sun, and Q. Huo, “Detect-order-construct: A tree construction based approach for hierarchical document structure analysis,” Pattern Recognition, vol. 156, p. 110836, 2024. https://dx.doi.org/10.1016/j.patcog.2024.110836. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0031320324005879
  5. W. L. Hamilton, R. Ying, and J. Leskovec, “Inductive representation learning on large graphs,” 2018. [Online]. Available: https://arxiv.org/abs/1706.02216
  6. X. Zhong, J. Tang, and A. Jimeno Yepes, “Publaynet: Largest dataset ever for document layout analysis,” in 2019 International Conference on Document Analysis and Recognition (ICDAR), 2019. https://dx.doi.org/10.1109/ICDAR.2019.00166 pp. 1015–1022.
  7. B. Pfitzmann, C. Auer, M. Dolfi, A. S. Nassar, and P. Staar, “Doclaynet: A large human-annotated dataset for document-layout segmentation,” in Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, ser. KDD ’22. New York, NY, USA: Association for Computing Machinery, 2022. https://dx.doi.org/10.1145/3534678.3539043. ISBN 9781450393850 p. 3743–3751. [Online]. Available: https://doi.org/10.1145/3534678.3539043
  8. M. Li, Y. Xu, L. Cui, S. Huang, F. Wei, Z. Li, and M. Zhou, “DocBank: A benchmark dataset for document layout analysis,” in Proceedings of the 28th International Conference on Computational Linguistics, D. Scott, N. Bel, and C. Zong, Eds. Barcelona, Spain (Online): International Committee on Computational Linguistics, Dec. 2020. https://dx.doi.org/10.18653/v1/2020.coling-main.82 pp. 949–960. [Online]. Available: https://aclanthology.org/2020.coling-main.82/
  9. J. Li, A. Sun, J. Han, and C. Li, “A survey on deep learning for named entity recognition,” IEEE Transactions on Knowledge and Data Engineering, vol. 34, no. 1, pp. 50–70, 2022. https://dx.doi.org/10.1109/TKDE.2020.2981314
  10. Y. Wang, H. Tong, Z. Zhu, and Y. Li, “Nested named entity recognition: A survey,” ACM Trans. Knowl. Discov. Data, vol. 16, no. 6, Jul. 2022. https://dx.doi.org/10.1145/3522593. [Online].
  11. R. Avyodri, S. Lukas, and H. Tjahyadi, “Optical character recognition (ocr) for text recognition and its post-processing method: A literature review,” in 2022 1st International Conference on Technology Innovation and Its Applications (ICTIIA), 2022. https://dx.doi.org/10.1109/IC-TIIA54654.2022.9935961 pp. 1–6.
  12. G. Jaume, H. K. Ekenel, and J.-P. Thiran, “Funsd: A dataset for form understanding in noisy scanned documents,” 2019. [Online]. Available: https://arxiv.org/abs/1905.13538
  13. A. Abdallah, D. Eberharter, Z. Pfister, and A. Jatowt, “Transformers and language models in form understanding: A comprehensive review of scanned document analysis,” 2024. [Online]. Available: https://arxiv.org/abs/2403.04080
  14. Z. Wang, M. Zhan, X. Liu, and D. Liang, “DocStruct: A multimodal method to extract hierarchy structure in document for general form understanding,” in Findings of the Association for Computational Linguistics: EMNLP 2020, T. Cohn, Y. He, and Y. Liu, Eds. Online: Association for Computational Linguistics, Nov. 2020. https://dx.doi.org/10.18653/v1/2020.findings-emnlp.80 pp. 898–908. [Online]. Available: https://aclanthology.org/2020.findings-emnlp.80/
  15. S. Wei and N. Xu, “Paragraph2graph: A gnn-based framework for layout paragraph analysis,” 2023. [Online]. Available: https://arxiv.org/abs/2304.11810
  16. A. Conneau, K. Khandelwal, N. Goyal, V. Chaudhary, G. Wenzek, F. Guzmán, E. Grave, M. Ott, L. Zettlemoyer, and V. Stoyanov, “Unsupervised cross-lingual representation learning at scale,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), 2020. https://dx.doi.org/10.18653/v1/2020.acl-main.747 pp. 8440–8451.
  17. R. Smith, “An overview of the tesseract OCR engine,” in International Conference on Document Analysis and Recognition (ICDAR), vol. 2, 2007. https://dx.doi.org/10.1109/ICDAR.2007.4376991 pp. 629–633.
  18. Y. Huang, T. Lv, L. Cui, Y. Lu, and F. Wei, “Layoutlmv3: Pre-training for document ai with unified text and image masking,” 2022. [Online]. Available: https://arxiv.org/abs/2204.08387
  19. Text Encoding Initiative Consortium, “TEI P5: Guidelines for Electronic Text Encoding and Interchange,” https://tei-c.org/guidelines/, 2023, version P5.