Skip to content
KAISTARRCARRC

Papers & Patents

Publications & Patents

194 outputs

202615

What Are You Really Asking For? A Comparative 5W1H Analysis of Learner Questioning in CPR Training with IVAs in Screen-based and Augmented Reality Environments
Hyerim Park* · Jinseok Hong · Heejeong Ko · Woontack Woo

Proceedings of the 2026 CHI Conference on Human Factors in Computing SystemsTop-tier

Question-asking is one of the key indicators of cognitive engagement. However, understanding how the distinct psychological affordances of presentation media shape learners’ spoken inquiries with embodied Intelligent Virtual Agents (IVAs) remains limited. To systematically examine this process, we propose a 5W1H-based framework for analyzing learner questions. Using this framework, we conducted a user study comparing an Augmented Reality-based IVA (AR-IVA) deployed in the physical environment with a screen-based IVA (Video-IVA) during cardiopulmonary resuscitation (CPR) instruction. Results showed that the AR-IVA elicited higher spatial and social presence and promoted more frequent and longer questions focused on clarification and understanding. In contrast, the Video-IVA encouraged questions regarding procedural refinement. Presence acted as a selective filter, shaping the timing and …

A Unified Hand and Gesture Tracking via Offloading Framework for Object-mediated Interaction in Wearable AR
Woojin Cho* · Taewook Ha · Taejun Son · Woontack Woo†

IEEE Conference on Virtual Reality and 3D User Interfaces (VR) 2026Top-tier

We propose a novel object-mediated hand interaction system that enables real-time operation with everyday objects on wearable augmented reality (AR) devices. Despite recent advances, both commercial and academic hand interaction techniques remain constrained, typically requiring external hardware or depending exclusively on bare-hand gestures. Motivated by these constraints, we developed an offloading framework that integrates a high-fidelity transformer-based 3D hand reconstruction model with a dynamic gesture recognition network powered by gated recurrent units (GRU). This architecture ensures stable and accurate gesture recognition even during interaction with physical objects. To evaluate its quantitative performance, we collected a custom dataset based on a predefined gesture set, achieving 93.0% accuracy in 5-fold cross-validation. The complete system implemented on Microsoft HoloLens 2 operates at a real-time framerate, and we further analyze the latency of each step in our framework. Through this interaction paradigm, users can experience immersive and intuitive AR in everyday environments with minimal disruption to natural action behavior. Our projects are available at https://github.com/kaist-uvrlab/UnifiedHOInteraction.

HMD-only Controllable 3D Gaussian Avatars in VR: Face and Full-body Demonstration
Seokhwan Yang · Hail Song · Woontack Woo

IEEE Conference on Virtual Reality and 3D User Interfaces Abstracts and Workshops (VRW) 2026

We present a VR research demo for practical telepresence with photorealistic 3D Gaussian Splatting (3DGS) avatars under HMD-only sensing. The demo consists of two parts that provide complementary experiences in a single session: a face part, where HMD-provided blendshape signals drive an expressive 3DGS head avatar rendered stereoscopically on the client, and a full-body part, where a user-specific 3DGS full-body avatar is created from a single image and driven using HMD signals. During the demo, participants first capture a single image for full-body avatar creation, experience facial control using a pre-built third-person head avatar while the full-body avatar is generated, and then control their personalized full-body avatar in VR. By enabling both facial and full-body controllability and contrasting two rendering-responsibility designs, client-side direct 3D rendering and server-side stereoscopic image rendering/streaming-the demo offers practical insights for designing responsive, realistic VR telepresence systems under HMD-only sensing.

Real-time Superquadric Representation of Cuboids and Cylinders
using Instant Surface Normal Map for Dynamic Object Integration in AR
Yohan Lim · Yoonseok Shin · Woontack Woo

IEEE Conference on Virtual Reality and 3D User Interfaces Abstracts and Workshops (VRW) 2026

We present a real-time AR system that enables interaction with physical cuboid- and cylinder-shaped objects by representing them as compact superquadrics. Our approach uses Instant Surface Normal (ISN) maps-color images encoding per-pixel geometry-to detect and segment objects based on shape rather than appearance. Superquadric parameters are extracted sequentially, allowing efficient and updatable 3D modeling in real time. Experimental results show accurate detection, consistent segmentation, and reconstruction quality comparable to existing methods, while achieving over 5× faster inference. This system allows physical objects to be seamlessly integrated as interactive elements in dynamic AR environments.

VR Zen Garden: Designing a Virtual Environment for Stress Relief
Hail Song · Jinseok Hong · Seonji Kim · Kyung Taek Oh · Woontack Woo · Sungyoung Kim

IEEE Conference on Virtual Reality and 3D User Interfaces Abstracts and Workshops (VRW) 2026

We propose VR Zen Garden, a virtual reality (VR) application designed to support stress relief and meditation through interactive spatial soundscapes. The system enables users to create personalized auditory environments by combining two key interactions: Object Placement, which spatializes sound-emitting objects, and Sand Pattern Raking, which modulates sound properties through visual patterns. These inputs are processed by the Sound Spatializer and the Metaphoric Sound Mixer, allowing users to intuitively shape immersive 3D soundscapes in real time. VR Zen Garden offers an engaging and embodied approach to mindfulness, highlighting the potential of sound-driven design in immersive wellness applications.

Event-Based Referred Vibrotactile Feedback for Bare-Hand XR Interaction
Juyoung Lee* · Hyunseo Seo · Hyunjin Lee · Minju Baeck · Hui-Shyong Yeo · Woontack Woo†

IEEE Transactions on Visualization and Computer Graphics (Proc. IEEE VR 2026)Top 10%

We propose event-based vibrotactile feedback as an effective approach for enhancing bare-hand XR interaction, delivered through commodity smartwatches that users already wear. While controllers and haptic gloves provide rich tactile feedback, they introduce additional hardware that users must carry and wear throughout the day. Importantly, they interpose hardware between the user and the virtual world, occupying the hands, constraining finger motion, and undermining truly unobstructed bare-hand interaction. Our proposed approach delivers short discrete pulses at key manipulation moments such as object contact and state changes, providing tactile confirmation at the wrist without requiring additional hardware or disrupting finger movement. Two user studies with 26 participants demonstrate that this referred feedback significantly enhances user experience at either wrist placement, with 94% finding it helpful and particular benefits during complex tasks like knob rotation where haptic cues reduced visual attention demands. These findings establish design guidelines for integrating accessible haptics into everyday XR through personal devices, supporting broader adoption of natural hand interaction.

ForceCtrl: Hand-Raycasting with User-Defined Pinch Force for Control-Display Gain Application
Seo Young Oh · Junghoon Seo · Juyoung Lee · Boram Yoon · Sang Ho Yoon · Woontack Woo

IEEE Transactions on Visualization and Computer GraphicsTop 10%

We present ForceCtrl, a novel 3D hand raycasting technique that enhances pointing precision based on control-display (CD) gain controlled with user-defined pinch force. We introduce a target-agnostic approach for refining raycasting precision, overcoming limitations in human motor accuracy. User-defined pinch force, detected with surface electromyography (sEMG), enables users to easily activate or deactivate CD gain during interaction. We propose three CD gain strategies and compare them through target selection and placement tasks. Our system reduces selection errors, placement jitters, and user workload, especially for distant targets in high-difficulty tasks. These results highlight the effectiveness of applying CD gain to hand raycasting and demonstrate the potential of user-defined pinch force as a robust input modality for precise hand interaction in AR/VR.

Int3DNet: Scene-Motion Cross Attention Network for 3D Intention Prediction in Mixed Reality
Taewook Ha · Woojin Cho · Dooyoung Kim · Woontack Woo

IEEE Transactions on Visualization and Computer Graphics (Proc. IEEE VR 2026)Top 10%

We propose Int3DNet, a scene-aware network that predicts 3D intention areas directly from scene geometry and head-hand motion cues, enabling robust human intention prediction without explicit object-level perception. In Mixed Reality (MR), intention prediction is critical as it enables the system to anticipate user actions and respond proactively, reducing interaction delays and ensuring seamless user experiences. Our method employs a cross attention fusion of sparse motion cues and scene point clouds, offering a novel approach that directly interprets the user's spatial intention within the scene. We evaluated Int3DNet on MoGaze and CIRCLE datasets, which are public datasets for full-body human-scene interactions, showing consistent performance across time horizons of up to 1500 ms and outperforming the baselines, even in diverse and unseen scenes. Moreover, we demonstrate the usability of proposed method through a demonstration of efficient visual question answering (VQA) based on intention areas. Int3DNet provides reliable 3D intention areas derived from head-hand motion and scene geometry, thus enabling seamless interaction between humans and MR systems through proactive processing of intention areas.

OFERA: Blendshape-driven 3D Gaussian Control for Occluded Facial Expression to Realistic Avatars in VR
Seokhwan Yang · Boram Yoon · Seoyoung Kang · Hail Song · Woontack Woo

IEEE Transactions on Visualization and Computer Graphics (Proc. IEEE VR 2026)Top 10%

We propose OFERA, a novel framework for real-time expression control of photorealistic Gaussian head avatars for VR headset users. Existing approaches attempt to recover occluded facial expressions using additional sensors or internal cameras, but sensor-based methods increase device weight and discomfort, while camera-based methods raise privacy concerns and suffer from limited access to raw data. To overcome these limitations, we leverage the blendshape signals provided by commercial VR headsets as expression inputs. Our framework consists of three key components: (1) Blendshape Distribution Alignment (BDA), which applies linear regression to align the headset-provided blendshape distribution to a canonical input space; (2) an Expression Parameter Mapper (EPM) that maps the aligned blendshape signals into an expression parameter space for controlling Gaussian head avatars; and (3) a Mapper-integrated Avatar (MiA) that incorporates EPM into the avatar learning process to ensure distributional consistency. Furthermore, OFERA establishes an end-to-end pipeline that senses and maps expressions, updates Gaussian avatars, and renders them in real-time within VR environments. We show that EPM outperforms existing mapping methods on quantitative metrics, and we demonstrate through a user study that the full OFERA framework enhances expression fidelity while preserving avatar realism. By enabling real-time and photorealistic avatar expression control, OFERA significantly improves telepresence in VR communication. A project page is available at https://ysshwan147.github.io/projects/ofera/.

SceneLinker: Compositional 3D Scene Generation via Semantic Scene Graph from RGB Sequences
Seok-Young Kim · Dooyoung Kim · Woojin Cho · Hail Song · Suji Kang · Woontack Woo

IEEE Transactions on Visualization and Computer Graphics (Proc. IEEE VR 2026)Top 10%

We introduce SceneLinker, a novel framework that generates compositional 3D scenes via semantic scene graph from RGB sequences. To adaptively experience Mixed Reality (MR) content based on each user's space, it is essential to generate a 3D scene that reflects the real-world layout by compactly capturing the semantic cues of the surroundings. Prior works struggled to fully capture the contextual relationship between objects or mainly focused on synthesizing diverse shapes, making it challenging to generate 3D scenes aligned with object arrangements. We address these challenges by designing a graph network with cross-check feature attention for scene graph prediction and constructing a graph-variational autoencoder (graph-VAE), which consists of a joint shape and layout block for 3D scene generation. Experiments on the 3RScan/3DSSG and SG-FRONT datasets demonstrate that our approach outperforms state-of-the-art methods in both quantitative and qualitative evaluations, even in complex indoor environments and under challenging scene graph constraints. Our work enables users to generate consistent 3D spaces from their physical environments via scene graphs, allowing them to create spatial MR content. Project page is https://scenelinker2026.github.io.

Spatial Affordance-Aware Affine Transformation Between Heterogeneous Spaces for Mixed Reality Remote Collaboration
Seonji Kim · Dooyoung Kim · Selin Choi · Woontack Woo

IEEE Transactions on Visualization and Computer Graphics (Proc. IEEE VR 2026)Top 10%

We propose a spatial affordance-aware affine transformation method between heterogeneous spaces for continuous multi-object matching in shared Mixed Reality (MR) spaces. While previous redirection and spatial mapping approaches utilize physical objects and walkable areas, a critical gap remains in enabling continuous mapping between dissimilar physical environments that supports both precise object alignment and seamless locomotion in a shared space. Our method structurally segments heterogeneous spaces into interaction zones and constructs affine patches based on object adjacency and facing configuration, enabling continuous correspondence. We evaluate our method using a dataset of paired dissimilar spaces and demonstrate that, unlike conventional grid-based methods, our approach achieves broader spatial alignment and richer object matching. The results show that our method can serve as an effective mapping framework for shared environments requiring semantic continuity and structural coherence across diverse real-world spaces.

Streamlined Facial Data Collection based on Utterance and Emotional Data for Human-to-Avatar Reconstruction
Seoyoung Kang · Seokhwan Yang · Hail Song · Boram Yoon · Jinwook Kim · Kangsoo Kim · Woontack Woo

IEEE Transactions on Visualization and Computer Graphics (Proc. IEEE VR 2026)Top 10%

This study explores a streamlined facial data collection method for conversational contexts, addressing the limitations of existing approaches that often require extensive datasets and prioritize technical metrics over user perception and experience. We systematically investigate which facial expression data are essential for reconstructing photorealistic avatars and how they can be captured efficiently. Our research employs a two-phase methodology to identify efficient facial data collection strategies and evaluate their effectiveness. In the first phase, we conduct facial data acquisition and evaluate reconstruction performance using utterance data and emotional data. In the second phase, we carry out a comprehensive user evaluation comparing three progressive conditions: utterance only, utterance and emotional data, and a control condition involving extensive data. Findings from 24 participants engaged in simulated face-to-face conversations reveal that targeted utterance and emotional data achieve comparable levels of perceived realism, naturalness, and telepresence, while reducing training time and data usage when compared to the extensive data collection approach. These results demonstrate that targeted data inputs can enable efficient avatar face reconstruction, offering practical guidelines for real-time applications such as AR/VR telepresence and highlighting the trade-off between data quantity and perceived quality.