Multimodal Contrastive Learning for Zero-Shot Instruction-Following Robot with Synthetic Data
Résumé
Robot trajectory prediction is heavily dependent on large-scale real-world demonstrations, which limit scalability, increase data acquisition costs, and eventually prevent zero-shot generalization. To address this limitation, this paper introduces Zero-Shot Task Learning (ZSTL), a multimodal framework that uses structurally aligned synthetic data and contrastive learning to enable instruction-based trajectory generation without reliance on real-world demonstrations. ZSTL jointly encodes natural-language instructions, depth observations, LiDAR-derived spatial representations, and action trajectories within a joined embedding space, allowing cross-modal alignment and conditional behavior synthesis. The proposed architecture preserves modality structure prior to fusion by representing depth inputs as spatial tokens and LiDAR observations as temporal tokens. Together with a text token, these form a 101-token multimodal context attended over by a Transformer decoder to predict full 50-step trajectories with Gaussian uncertainty estimates. The system integrates a pretrained Bidirectional Encoder Representations from Transformers (BERT) language encoder, a ResNet-18 depth backbone, a one-dimensional convolutional LiDAR sequence encoder, and a two-layer Transformer decoder comprising approximately 125M parameters. Training was conducted entirely on a procedurally generated synthetic dataset of 5,000 samples for 50 epochs. The results demonstrate stable convergence, with the trajectory-negative log likelihood decreasing from 3.465 to -0.695 on validation data and the combined loss reaching -0.540 at epoch 18 under cosine annealed learning. The contrastive objective (InfoNCE, τ = 0.07) stabilized near 1.55, indicating consistent cross-modal alignment. The trajectory evaluation yielded an average final position error of 9.97 cm, a collision-free execution rate of 65.9%, and a task success rate of 59.2%, showing that structured synthetic supervision can support physically meaningful motion generation.
Citer ce document
Accès au document
Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter
Voir l'article sur le site de la revueStatistiques
Consultations : 1
Téléchargements : 0