Skip to content
KAISTARRCARRC

active

XR Memory

Development of Technology to Reenact Industrial Spatial Knowledge from Incomplete Data

Overview

AI-powered reconstruction and visualization of incomplete multimodal experiences with preserved spatiotemporal context.

xrmemory_eng.png

Final Research Goal This research aims to develop an intelligent XR platform that uses AI to integrate and analyze diverse forms of incomplete multimodal data, including video, audio, sensor data, location information, and EEG. The platform reconstructs these data into highly immersive 3D content while preserving their spatiotemporal context, and supports the authoring and visualization of such content.

Focus Areas - Multimodal memory data structuring and synchronization across heterogeneous data sources - Memory Scene Graph–based context modeling for semantic understanding, user intent inference, and memory retrieval - High-fidelity 3D spatial reconstruction from incomplete data - User cognitive state analysis and adaptive visualization for proactive safety guidance - AI agent–based memory content authoring and validation in real-world industrial environments.

Application The developed technologies can be applied to a wide range of areas, including industrial safety, expert knowledge transfer, education and training, accident reconstruction, and on-site task support. For example, the work processes and decision-making of skilled workers can be structured as Memory Scene Graphs and delivered to new workers as AR/XR-based guidance. Incomplete video and sensor data from industrial accidents can also be reconstructed as 3D environments to support accident analysis. In addition, the technologies can enable personalized job training and adaptive learning by adjusting the amount and level of guidance according to users' expertise and cognitive states.

Expected Contribution Preservation and Reuse of Experiential Knowledge: XRMemory reconstructs fragmented multimodal data into structured memory content that preserves spatial, temporal, and semantic context, enabling expert experience and tacit knowledge to be retained and reused - Improved Industrial Safety and Operational Efficiency: By jointly analyzing user states and task contexts, the system can identify potential hazards and provide timely, context-aware guidance to support safer and more effective task performance - Personalized and Adaptive Training: XRMemory tailors knowledge and guidance to users’ expertise, task progress, and cognitive states, supporting more effective skill transfer and learning - Accessible XR Content Authoring: AI agents and no-code authoring tools enable users without specialized development expertise to create, modify, and reuse XR content for specific environments and tasks.

Motivation

Problem

Much of the knowledge produced in real-world work and learning environments is not captured as complete, structured information. A skilled worker may recognize a subtle change in sound, notice an unusual vibration, shift attention toward a particular component, and make a decision based on years of accumulated experience. Yet conventional recordings capture only fragments of this process—such as video, audio, sensor readings, or system logs—without preserving how these signals relate to one another across space and time.

As a result, what is recorded often represents only what happened, while the surrounding context—where it happened, what objects and people were involved, what preceded the event, and what the user was attending to or trying to accomplish—is easily lost. Reusing these fragmented records for training, knowledge transfer, safety analysis, or immersive reconstruction therefore requires more than simply storing additional data. The data must be organized into a representation that preserves the relationships among people, objects, actions, space, time, and user state.

Existing limitations

Current multimodal and XR systems typically process individual data streams independently or combine them only at the level of synchronized playback. Video, audio, motion, physiological signals, and spatial information may be recorded together, but missing segments, sensor noise, tracking failures, and differences in temporal resolution make it difficult to reconstruct a coherent experience.

Digital twins and conventional 3D representations provide spatial models of environments and equipment, but they generally do not capture the semantic and cognitive context surrounding human activity. They represent where objects are and how they are configured, but not necessarily why a worker attended to a particular component, how an action relates to an earlier event, or which part of an experience is important for another user to understand.

Content creation is another barrier. Reconstructing incomplete experiences as interactive XR content typically requires extensive manual editing and specialized development expertise. This limits the ability of workers, trainers, and domain experts to directly preserve and reuse their own knowledge.

Our approach

XRMemory approaches these challenges by treating an experience as a structured combination of multimodal signals, spatial context, human behavior, and semantic relationships rather than as a collection of independent recordings.

First, heterogeneous data—including video, audio, IMU, location, EEG, and other sensor signals—are temporally and spatially aligned to establish a common representation of an event. AI-based processing is then used to identify relevant objects, actions, people, temporal events, and relationships, organizing them into a Memory Scene Graph (MSG). The MSG provides a structured representation through which fragmented observations can be connected, missing contextual information can be inferred, and relevant memory content can be retrieved according to the user or task.

In parallel, incomplete visual and spatial observations are reconstructed into high-fidelity 3D environments using AI-based spatial reconstruction techniques. Cognitive and behavioral signals are incorporated to represent not only what occurred in the environment but also aspects of how users perceived and responded to it, such as attention, cognitive load, or hazard awareness.

Finally, these representations are connected to an intelligent XR authoring and visualization platform. Natural-language interaction and AI agents allow domain experts and non-expert users to create, edit, and adapt memory-based XR content without conventional programming workflows. The resulting content can then be presented according to the user’s context and needs, enabling past experiences and expert knowledge to be reconstructed as reusable, adaptive XR experiences.