
10 F. Holm et al.
6. Devlin, J., Chang, M.W., Lee, K., Toutanova, K.: Bert: Pre-training of deep bidi-
rectional transformers for language understanding. In: Proceedings of the 2019
conference of the North American chapter of the association for computational
linguistics: human language technologies, volume 1 (long and short papers). pp.
4171–4186 (2019)
7. Funke, I., Rivoir, D., Speidel, S.: Metrics matter in surgical phase recognition
(2023), https://arxiv.org/abs/2305.13961
8. Gastager, D., Ghazaei, G., Patsch, C.: Watch and learn: Leveraging expert knowl-
edge and language for surgical video understanding. ArXiv abs/2503.11392
(2025), https://api.semanticscholar.org/CorpusID:277044026
9. Grammatikopoulou, M., Flouty, E., Kadkhodamohammadi, A., Quellec, G., Chow,
A., Nehme, J., Luengo, I., Stoyanov, D.: Cadis: Cataract dataset for image seg-
mentation. arXiv preprint arXiv:1906.11586 (2019)
10. Holm, F., Ghazaei, G., Czempiel, T., Özsoy, E., Saur, S., Navab, N.: Dynamic
scene graph representation for surgical video. In: Proceedings of the IEEE/CVF
international conference on computer vision. pp. 81–87 (2023)
11. Ji, J., Krishna, R., Fei-Fei, L., Niebles, J.C.: Action genome: Actions as composi-
tions of spatio-temporal scene graphs. In: Proceedings of the IEEE/CVF conference
on computer vision and pattern recognition. pp. 10236–10247 (2020)
12. Köksal, Ç., Ghazaei, G., Holm, F., Farshad, A., Navab, N.: Sangria: Surgical
video scene graph optimization for surgical workflow prediction. arXiv preprint
arXiv:2407.20214 (2024)
13. Liu, Z., Ning, J., Cao, Y., Wei, Y., Zhang, Z., Lin, S., Hu, H.: Video swin trans-
former. In: Proceedings of the IEEE/CVF conference on computer vision and pat-
tern recognition. pp. 3202–3211 (2022)
14. Luo, Z., Durante, Z., Li, L., Xie, W., Liu, R., Jin, E., Huang, Z., Li, L.Y., Wu,
J., Niebles, J.C., Adeli, E., Fei-Fei, L.: MOMA-LRG: Language-refined graphs for
multi-object multi-actor activity parsing. In: Thirty-sixth Conference on Neural
Information Processing Systems Datasets and Benchmarks Track (2022), https:
//openreview.net/forum?id=eJhc_CPXQIT
15. Murali, A., Alapatt, D., Mascagni, P., Vardazaryan, A., Garcia, A., Okamoto, N.,
Mutter, D., Padoy, N.: Encoding surgical videos as latent spatiotemporal graphs
for object and anatomy-driven reasoning. In: International Conference on Medi-
cal Image Computing and Computer-Assisted Intervention. pp. 647–657. Springer
(2023)
16. Rodin, I., Furnari, A., Min, K., Tripathi, S., Farinella, G.M.: Action scene graphs
for long-form understanding of egocentric videos. In: Proceedings of the IEEE/CVF
Conference on Computer Vision and Pattern Recognition. pp. 18622–18632 (2024)
17. Wang, G., Li, Z., Chen, Q., Liu, Y.: Oed: towards one-stage end-to-end dynamic
scene graph generation. In: Proceedings of the IEEE/CVF Conference on Com-
puter Vision and Pattern Recognition. pp. 27938–27947 (2024)
18. Wang, J., Wen, Z., Li, X., Guo, Z., Yang, J., Liu, Z.: Pair then relation: Pair-net
for panoptic scene graph generation. IEEE Transactions on Pattern Analysis and
Machine Intelligence (2024)
19. Yang, J., Ang, Y.Z., Guo, Z., Zhou, K., Zhang, W., Liu, Z.: Panoptic scene graph
generation. In: European Conference on Computer Vision. pp. 178–196. Springer
(2022)
20. Yang, J., Peng, W., Li, X., Guo, Z., Chen, L., Li, B., Ma, Z., Zhou, K., Zhang, W.,
Loy, C.C., et al.: Panoptic video scene graph generation. In: Proceedings of the
IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 18675–
18685 (2023)