The CASTLE 2024 Dataset : Advancing the Art of Multimodal Understanding

dc.contributor.authorRossetto, Luca
dc.contributor.authorBailer, Werner
dc.contributor.authorDang-Nguyen, Duc Tien
dc.contributor.authorHealy, Graham
dc.contributor.authorJónsson, Björn Pór
dc.contributor.authorKongmeesub, Onanong
dc.contributor.authorLe, Hoang Bao
dc.contributor.authorRudinac, Stevan
dc.contributor.authorSchöffmann, Klaus
dc.contributor.authorSpiess, Florian
dc.contributor.authorTran, Allie
dc.contributor.authorTran, Minh Triet
dc.contributor.authorTran, Quang Linh
dc.contributor.authorGurrin, Cathal
dc.contributor.departmentDepartment of Computer Science
dc.date.accessioned2026-09-24T14:11:01Z
dc.date.available2026-09-24T14:11:01Z
dc.date.issued2025-10-27
dc.descriptionPublisher Copyright: © 2025 Copyright held by the owner/author(s).en
dc.description.abstractEgocentric video has seen increased interest in recent years, as it is used in a range of areas. However, most existing datasets are limited to a single perspective. In this paper, we present the CASTLE 2024 dataset, a multimodal collection containing ego- and exo-centric (i.e., first- and third-person perspective) video and audio from 15 time-aligned sources, as well as other sensor streams and auxiliary data. The dataset was recorded by volunteer participants over four days in a common location and includes the point of view of 10 participants, with an additional 5 fixed cameras providing an exocentric perspective. The entire dataset contains over 600 hours of UHD video recorded at 50 frames per second. In contrast to other datasets, CASTLE 2024 does not contain any partial censoring, such as blurred faces or distorted audio. The dataset is available via https://castle-dataset.github.io/.en
dc.description.versionPeer revieweden
dc.format.extent7
dc.format.extent3328930
dc.format.extent12629-12635
dc.format.extent
dc.identifier.citationRossetto, L, Bailer, W, Dang-Nguyen, D T, Healy, G, Jónsson, B P, Kongmeesub, O, Le, H B, Rudinac, S, Schöffmann, K, Spiess, F, Tran, A, Tran, M T, Tran, Q L & Gurrin, C 2025, The CASTLE 2024 Dataset : Advancing the Art of Multimodal Understanding. in MM 2025 - Proceedings of the 33rd ACM International Conference on Multimedia, Co-Located with MM 2025. MM 2025 - Proceedings of the 33rd ACM International Conference on Multimedia, Co-Located with MM 2025, Association for Computing Machinery, Inc, pp. 12629-12635, 33rd ACM International Conference on Multimedia, MM 2025, Dublin, Ireland, 27/10/25. https://doi.org/10.1145/3746027.3758199en
dc.identifier.citationconferenceen
dc.identifier.doi10.1145/3746027.3758199
dc.identifier.isbn9798400720352
dc.identifier.other251045021
dc.identifier.otherb9447912-b7d5-4e82-a677-617452633f90
dc.identifier.other105024064173
dc.identifier.urihttps://hdl.handle.net/20.500.11815/8361
dc.language.isoen
dc.publisherAssociation for Computing Machinery, Inc
dc.relation.ispartofseriesMM 2025 - Proceedings of the 33rd ACM International Conference on Multimedia, Co-Located with MM 2025; ()en
dc.relation.ispartofseriesMM 2025 - Proceedings of the 33rd ACM International Conference on Multimedia, Co-Located with MM 2025; ()en
dc.relation.urlhttps://www.scopus.com/pages/publications/105024064173en
dc.rightsinfo:eu-repo/semantics/openAccessen
dc.subjectdataseten
dc.subjectegocentric visionen
dc.subjectlifeloggingen
dc.subjectmulti-perspective videoen
dc.subjectmultimodal understandingen
dc.subjectHuman-Computer Interactionen
dc.subjectSoftwareen
dc.subjectArtificial Intelligenceen
dc.subjectComputer Graphics and Computer-Aided Designen
dc.titleThe CASTLE 2024 Dataset : Advancing the Art of Multimodal Understandingen
dc.type/dk/atira/pure/researchoutput/researchoutputtypes/contributiontobookanthology/conferenceen

Skrár

Original bundle

Niðurstöður 1 - 1 af 1
Nafn:
3746027.3758199.pdf
Stærð:
3.17 MB
Snið:
Adobe Portable Document Format