
We investigate how input modality (gloves vs. controller) and its visualisation (hand vs. controller model) influence the execution of everyday tasks in VR. In a virtual kitchen, participants perform typical tasks (e.g. table setting, dishwashing, cutting, cleaning, pouring, showing/evaluating).
In addition to usability/workload, we record high-resolution motion sequences (trajectories) and interactions and analyse them using semantic segmentation (e.g. grasping → acting → putting down). This allows us to assess not only "how fast", but also "how" movements were performed and how consistent/natural they are.
Three input setups are compared:

M: Motion capture gloves with hand visualisation
H: Controller, hand is virtually reconstructed
C: Controller with visible controller model