| | |
Last updated on August 7, 2026. This conference program is tentative and subject to change
Technical Program for Thursday August 6, 2026
| |
| TH0900-1 Invited Sessions, Ballroom I |
Add to My Program |
| Session 4.1: Application-Driven Robot Learning and Control I |
|
| |
| Chair: Liang, Wenyu | Institute for Infocomm Research, A*STAR |
| Co-Chair: Li, Dongyu | Beihang University |
| Organizer: Liang, Wenyu | Institute for Infocomm Research, A*STAR |
| Organizer: Li, Dongyu | Beihang University |
| |
| 09:00-10:00, Paper TH0900-1.1 | Add to My Program |
| SEMG-Only Interaction Stiffness Estimation for Contact-Rich Human-Robot Collaboration (I) |
|
| Xu, Chan | University of Chinese Academy of Sciences |
| Chen, Silu | Ningbo Institute of Materials Technology and Engineering, CAS |
| Wang, Dehao | Zhejiang University of Technology; Ningbo Institute of Materials Technology and Engineering, Chinese Academy of Sciences |
| Chen, Xiyu | The Ningbo Institute of Materials Technology and Engineering |
| Zhang, Chi | Ningbo Institute of Material Technology and Engineering, CAS |
| Yang, Guilin | Ningbo Institute of Material Technology and Engineering, Chinese Academy of Sciences |
| Fang, Zaojun | Ningbo Institute of Materials Technology & Engineering, CAS |
Keywords: Cybernetics Automation and Control, Robotics and Automation Applications, Methodologies for Robotics and Automation
Abstract: In contact-rich human–robot collaboration, estimating the operator's interaction stiffness is essential for compliant and stable task execution. Although interaction stiffness can be inferred solely from upper-limb surface electromyography (sEMG), identifying informative features remains challenging due to the low signal-to-noise ratio of the recorded signals. In this paper, a robust feature selection criterion is developed by removing noise entropy from both relevance and redundancy evaluation. The noise censoring threshold is first estimated from the expectation of the smallest extreme value distribution, enabling noise-free mutual information to be obtained without specifying an additional confidence level. A noise-free similarity metric is then constructed through the symmetric application of noise-free mutual information. Robust feature selection is finally achieved by maximizing relevance and minimizing redundancy under these noise-free measures. Experiments on a human–robot collaborative wiping task demonstrate that only ten sEMG features are sufficient for accurate stiffness reconstruction, achieving a reconstruction error rate of 2.10% and improving performance by 37.37% over three state-of-the-art feature selection methods.
|
| |
| 09:00-10:00, Paper TH0900-1.2 | Add to My Program |
| TopoFlatten: Action Decision Via Topology-Guided Confidence Scoring for Robotic Garment Unfolding |
|
| Yang, Qingyun | Jilin University |
Keywords: Embodied AI, Deep Learning, Robot Vision
Abstract: 机器人服装展开在无序初始配置中依然具有挑战性,因为折叠、重叠和自遮挡使得可靠的动作选择变得困难。局部几何线索常常模糊不清,导致决策不稳定。我们提出了TopoFlatten,一种基于拓扑的理解框架,利用结构表征进行决策。给定RGB-D观测,服装被编码为基于骨架的图,捕捉其拓扑结构。然后构建并利用学习得的置信度模型构建并评估一组拓扑条件候选动作,估计其预期的改进效果。选择并执行置信度最高的动作,形成闭环。通过结合拓扑表示与置信度驱动的评估,该方法实现候选动作的结构化比较和更可靠的展开决策。实验结果显示,在不同服装配置中,耐用性和效率均有所提升。
|
| |
| 09:00-10:00, Paper TH0900-1.3 | Add to My Program |
| Enhancing Spatial Understanding in Vision-Language Models Via Curriculum Learning (I) |
|
| Zhong, Zhe | Zhejiang University |
| Ren, Qinyuan | Zhejiang University |
Keywords: Robot Vision, Deep Learning, Embodied AI
Abstract: The development of Embodied AI urgently necessitates high-fidelity environment modeling enriched with spatial context. However, existing 3D semantic scene understanding methods predominantly focus on isolated instance-level labels or high-dimensional semantic vector embeddings, lacking effective modeling of explicit spatial-semantic relationships between objects within a scene. To address this issue, we propose a method that guides Vision-Language Model(VLMs) to learn scene spatial layout relationships via Supervised Fine-Tuning (SFT). First, based on the InteriorGS dataset, we construct a Visual Question Answering (VQA) dataset spanning 100 scenes, comprising RGB images and grounding images with 2D bounding box prompts. Within this dataset, we systematically annotate the spatial-semantic relationship graphs among visible instances. Second, we utilize the parameter-efficient fine-tuning strategy of Low-Rank Adaptation (LoRA) to enhance the spatial relationship reasoning capabilities of the baseline model, Qwen2.5-VL-7B-Instruct. Furthermore, we design an easy-to-hard, three-stage curriculum learning scheme: progressing from single-image single-instance relationship reasoning, advancing to single-image multi-instance relationship understanding, and ultimately achieving global spatial layout perception across continuous frames. Comprehensive evaluations demonstrate that our method enables the model to effectively comprehend and output spatial-semantic relationship triplets in a predefined format, significantly outperforming the baseline model in inter-instance spatial-semantic reasoning. Our research validates the feasibility of endowing existing VLMs with preliminary spatial intelligence via SFT, laying the foundation for constructing next-generation semantic scene representations enriched with spatial relational information.
|
| |
| 09:00-10:00, Paper TH0900-1.4 | Add to My Program |
| Cooperative Grasping Strategy for Dual-Drive Gantry Robots with Coaxial Dual-Gripper (I) |
|
| Lu, Jiaxiang | Ningbo Institute of Materials Technology and Engineering Chinese Academy of Sciences |
| Chen, Silu | Ningbo Institute of Materials Technology and Engineering, CAS |
| Zhang, Chi | Ningbo Institute of Material Technology and Engineering, CAS |
| Yang, Guilin | Ningbo Institute of Material Technology and Engineering, Chinese Academy of Sciences |
Keywords: Robotics and Automation Applications, Planning and Control, Methodologies for Robotics and Automation
Abstract: Aiming at the problems of spatial interference and low operating efficiency caused by rigid X-axis coupling and independent Y-axis movement of the coaxial dual grippers in dual-drive gantry robots, this paper establishes a kinematic constraint model and a multi-scenario grasping time-consuming model, and proposes a Partial Local Optimization (PLO) strategy based on distance ranking. This strategy constructs a local candidate set through distance ranking, performs traversal search for the optimal workpiece pairing under interference-free constraints, and ensures real-time computing performance and global optimization capability. Experimental results show that the PLO strategy improves efficiency by more than 20% compared with the single-gripper serial strategy and by approximately 7.56% compared with the nearest workpiece pair optimization strategy, which can meet the requirements of efficient cooperative operation in industrial high-speed sorting scenarios.
|
| |
| 09:00-10:00, Paper TH0900-1.5 | Add to My Program |
| Reinforcement Learning-Based Prescribed Performance Tracking Control for Robotic Manipulators |
|
| Li, Yan | Changchun University of Technology |
| Liu, Xin | Changchun University of Technology |
| Zhang, Yan | Changchun University of Technology |
| Fan, Jianlei | Changchun University of Technology |
| Lu, Zengpeng | Changchun University of Technology |
Keywords: Dynamics
Abstract: 本文提出了固定时间轨迹跟踪 面对动态的机器人操作手控制策略 不确定性和输入饱和,结合规定 与自适应强化学习的绩效控制。A 规定的性能功能被严格设计为 在预定义的瞬态和范围内约束跟踪误差 系统运行过程中的稳态边界。一个 自适应 具有动态参数调整的演员-批评者框架为 开发中,演员网络产生控制 政策 而 CRRITIC 网络则在 中评估动作-值函数 真实 时间基于系统状态。一种反饱和 薪酬 由变换误差信号推导出的定律被构造出来 到 减轻执行器饱和效应。此外,还有 自适应 非奇异终端滑动模式控制器是 并入 确保跟踪误差的固定时间收敛 独立候选人 初始条件。基于李雅普诺夫的稳定性分析 严格地建立了有界性和收敛性 性质 封闭式系统的
|
| |
| TH0900-2 Regular Sessions, Ballroom II |
Add to My Program |
| Session 4.2: Decision and Control |
|
| |
| Chair: Sun, Shuo | Singapore-MIT Alliance for Research and Technology (SMART) |
| Co-Chair: Bao, Yinyin | Lanzhou University of Technology |
| |
| 09:00-10:00, Paper TH0900-2.1 | Add to My Program |
| An Outcome-Oriented Multimodal Decision Support Framework for Three-Level Emergency Department Triage |
|
| Yao, Siyang | Northeastern University |
| Gao, Xing | Tianjin University Tianjin Hospital |
| Tao, Yushun | University of Chinese Academy of Sciences |
| Xiong, Jing | Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences |
Keywords: Decision Support Systems, Natural Language Processing, Neural Networks
Abstract: Emergency department (ED) triage is a safety-critical decision process, yet conventional triage practice may become inconsistent under crowding and high workload, especially for rare high-risk patients. This paper presents TriadNet, an outcome-oriented multimodal decision support framework for three-level ED triage. Instead of reproducing the conventional five-level triage scale, TriadNet formulates ED triage as an actionable prediction task with three categories: Emergency, Admission, and Home. The framework uses only intake-stage information available at triage, including short clinical text and structured variables, and integrates semantic and physiological risk cues through a dual-branch multimodal architecture. A lightweight fusion classifier is further combined with a reinforcement-adaptive loss controller to address severe class imbalance and improve recognition of underrepresented high-risk cases. Experiments on the MIMIC-IV-ED dataset show that TriadNet achieves a macro-F1 of 0.840, a macro PR-AUC of 0.857, and an Emergency recall of 0.88, outperforming conventional structured-data baselines and single-modality variants. Error analysis further shows that most residual Emergency errors are shifted to Admission rather than Home, indicating a safer misclassification pattern for deployment-oriented triage support. These results suggest that TriadNet provides an effective and practical framework for early-stage ED risk stratification without requiring post-triage laboratory tests or downstream clinical information.
|
| |
| 09:00-10:00, Paper TH0900-2.2 | Add to My Program |
| A Cooperative Game-Based Approach for Coupled Fault Identification of Unmanned Aerial Vehicles |
|
| Huang, Ping | School of Computer Science and Engineering, Changchun University of Technology |
| Cheng, Chao | Changchun University of Technology |
| Meng, Fantuo | School of Computer Science and Engineering, Changchun University of Technology |
| Li, Hualiang | School of Computer Science and Engineering, Changchun University of Technology |
Keywords: Decision Support Systems, Systems Modeling & Control
Abstract: This paper proposes a method for identifying coupled faults and assessing the health of UAVs based on Shapley values within a cooperative game framework: First, a UAV system dynamics model is constructed, and a state observer is used to extract residual features from multi-source heterogeneous data; second, key UAV components are defined as alliance members in a cooperative game, and a characteristic function is constructed based on system residuals; Building on this, Shapley value theory is introduced to calculate the marginal contribution rate of each component to system failures, thereby achieving multi-source coupled fault isolation and component attribution; finally, the responsibility weights are mapped to a dynamic system health index. The simulation results demonstrate that this method can effectively mitigate interference from closed-loop control mechanisms and reduce fault detection delay by 82.3%.
|
| |
| 09:00-10:00, Paper TH0900-2.3 | Add to My Program |
| Highway Ramp Merging Behavior of Autonomous Vehicles Using Human-Inspired Structured LSTM with Reward Machine Reinforcement Learning |
|
| Lv, Yiyan | Northeast Normal University |
| Yan, Chufei | Northeast Normal University |
| Cui, Zhihao | Tongji University |
| Wang, Yulei | Tongji University |
Keywords: Intelligent Transportation Systems, Transportation Systems, Deep Reinforcement Learning
Abstract: For autonomous vehicles (AVs), safe and efficient highway on-ramp merging remains a challenging decision-making problem. To jointly improve safety, merging success, and driving efficiency, this paper proposes a novel human-inspired structured Long Short-Term Memory (HIS-LSTM) with reward machine reinforcement learning (RMRL). To ensure highway driving safety, an RMRL framework for on-ramp merging is first developed by incorporating the safe distance derived from the Vienna Convention into both the safety constraints of the Deep Q-Network (DQN) and the reward function. For the ramp merging task, HIS-LSTM is designed to emulate the cognitive process of human drivers. Specifically, it first monitors the motion patterns of surrounding vehicles (SVs), then analyzes their potential driving intentions, and subsequently fuses this behavioral understanding with the features of the ego vehicle (EV) to generate reference speed. Simulation experiments are conducted in normal and complex highway on-ramp merging scenarios built in highway-env, with comparisons against eight baseline methods. Experimental results show that the proposed method achieves a superior trade-off among safety, merging success, and efficiency, while exhibiting strong robustness across different traffic complexities.
|
| |
| 09:00-10:00, Paper TH0900-2.4 | Add to My Program |
| Temperature Anomaly Detection and Early Warning Method for Wind Turbines Based on Multivariate Time-Series Representation Learning |
|
| Wang, Xiaodong | China Huadian Corporation Ltd |
| Wu, Baoming | China Huadian Corporation Ltd |
| Shao, Xiangwen | China Huadian Corporation Ltd |
| Fu, Yintao | China Huadian Corporation Ltd |
| Wang, Shuanhong | China Huadian Corporation Ltd |
Keywords: Energy Efficiency, Deep Learning, Smart Sensor Networks
Abstract: Untimely temperature monitoring in wind turbines often causes unplanned downtime and economic losses. To address this, this paper proposes an anomaly detection and early warning method using multivariate SCADA time-series data. First, an Attention-LSTM network is utilized to efficiently extract multivariate coupling relationships and long-range thermal inertia dependencies. Second, a Variational AutoEncoder (VAE) imposes distribution constraints on the latent space, significantly enhancing the model's robustness against transient sensor noise and operating fluctuations. Furthermore, a semi-supervised auxiliary discriminator (AD) leverages limited historical anomaly data to refine diagnostic boundaries and reduce false alarm rates. Finally, transfer validation experiments on real-world wind turbine datasets demonstrate that this approach outperforms traditional baselines in anti-noise performance and reliably captures thermodynamic degradation trends, such as gearbox oil overheating, achieving effective early fault warning.
|
| |
| 09:00-10:00, Paper TH0900-2.5 | Add to My Program |
| Cascade Control and Virtual Impedance Design for Hardware-In-The-Loop Simulation of Scaled Tracked Vehicles |
|
| Liang, Jingshi | Jilin University |
| Zhuang, Ye | Jilin University |
| Xu, Bin | Jilin University |
Keywords: Robotics and Automation in Unstructured Environment, Force/Impedance Control, Methodologies for Robotics and Automation
Abstract: Abstract—Hardware-in-the-Loop (HIL) simulation is crucial for tracked vehicle control verification but is often constrained by the high costs of multi-motor testbeds and dynamic inertia mismatch. To overcome these hardware limitations, this paper proposes a novel asymmetric mixed-reality HIL architecture. By coupling a physical drive-load motor pair (representing the inner track) with a high-fidelity virtual motor digital twin (representing the outer track), complex 3-DOF differential steering dynamics are validated using only a single physical testbed. A virtual impedance control strategy is designed to emulate the large transient inertia without noise-amplifying derivatives. Experimental results demonstrate that the proposed framework provides stable trajectory tracking and accurately replicates the skid-steering torque distribution.
|
| |
| TH1130-1 Invited Sessions, Ballroom I |
Add to My Program |
| Session 5.1: Application-Driven Robot Learning and Control II |
|
| |
| Chair: Gou, Yu | Changchun Institute of Optics, Fine Mechanics and Physics, Chinese Academy of Sciences |
| Co-Chair: Liang, Wenyu | Institute for Infocomm Research, A*STAR |
| Organizer: Liang, Wenyu | Institute for Infocomm Research, A*STAR |
| Organizer: Li, Dongyu | Beihang University |
| |
| 11:30-12:30, Paper TH1130-1.1 | Add to My Program |
| LLM-In-The-Loop Variable Impedance Control: Towards Safe Generalized and Personalized Robotic Interactions (I) |
|
| Xue, Junyuan | National University of Singapore |
| Liang, Wenyu | Institute for Infocomm Research, A*STAR |
| Xu, Yilan | National University of Singapore |
| Wu, Yan | A*STAR Institute for Infocomm Research |
| Lee, Tong Heng | National University of Singapore |
Keywords: Embodied AI, Planning and Control, Force/Impedance Control
Abstract: While impedance control provides compliance and flexibility in physical robotic interactions, it requires tedious parameter tuning for diverse tasks and user preferences. To overcome this limitation, a novel impedance control framework integrating a large language model is proposed. Leveraging few-shot learning and natural language comprehension, the LLM maps implicit user inputs to explicit controller parameters, enabling safe and personalized interactions across various tasks. Furthermore, an evaluator closes the control loop by supervising and refining the LLM's outputs. Physical experiments demonstrate that this approach significantly improves interaction safety, task generalization, and user personalization.
|
| |
| 11:30-12:30, Paper TH1130-1.2 | Add to My Program |
| Tactile-Reactive Surface Following Control on Rapidly Changing or Discontinuous Geometries (I) |
|
| Xu, Yilan | National University of Singapore |
| Xue, Junyuan | National University of Singapore |
| Li, Dongyu | Beihang University |
| Liang, Wenyu | Institute for Infocomm Research, A*STAR |
| Lee, Tong Heng | National University of Singapore |
Keywords: Force/Impedance Control, Haptics, Robotics and Automation in Unstructured Environment
Abstract: Existing force/torque (F/T) surface-following controllers assume smooth surface continuity and fail at geometric discontinuities, producing excessive force spikes or losing contact at sharp edges and rapid curvature transitions. We introduce a spatial leading-region tactile event detection paradigm that monitors the taxel-level force distribution, rather than aggregate F/T signals, to detect impending transitions before force spikes occur, and couples it with two event-triggered reactive pipelines: Discrete Re-orientation for discontinuous edges and Adaptive Velocity Regulation for continuous high-curvature turns. Under fixed pipeline-level parameters, the proposed framework achieves 100% success on all six main geometries, whereas the baseline remains unreliable on sharper transitions. A R{=}5,mm operating-envelope study reveals curvature-dependent velocity-scaling limits, motivating future adaptive parameter selection. These results support tactile spatial-event triggering as a practical path to safe, vision-free surface traversal on previously unknown geometries.
|
| |
| 11:30-12:30, Paper TH1130-1.3 | Add to My Program |
| Multimodal Emotion Recognition Using EEG and ECG: Exploring Cross-Modal Correlations and Feature Fusion Strategies (I) |
|
| Ding, Chenrui | Xi'an Jiaotong-Liverpool University |
| Wang, Xinrui | Xi'an Jiaotong-Liverpool University |
| Su, Chuanzhi | Xi'an Jiaotong-Liverpool University |
| Huang, Mengjie | Xi'an Jiaotong-Liverpool University |
| Yang, Rui | Xi'an Jiaotong-Liverpool University |
Keywords: Human/Machine Systems, Deep Learning, Human-Robot Interfaces
Abstract: Emotion recognition using physiological signals has attracted increasing attention in affective computing due to its objectivity and robustness, with electroencephalography (EEG) and electrocardiography (ECG) providing complementary information reflecting central and autonomic nervous system activities. This paper presents a multimodal emotion recognition framework based on EEG and ECG signals using the DREAMER dataset, where a unified processing pipeline is designed to ensure comparability between modalities through preprocessing, baseline correction, and feature extraction. Specifically, EEG features are extracted using power spectral density, while ECG features are derived from heart rate variability analysis. In addition, Pearson correlation analysis is conducted to investigate intra-modal, cross-modal, and feature-label relationships, and a support vector machine (SVM) classifier is employed for binary classification of valence, arousal, and dominance. Experimental results show that EEG and ECG exhibit low cross-modal correlation, and fused features do not consistently outperform single-modality features across tasks. These findings indicate that simple feature-level fusion may be insufficient to fully exploit multimodal information, suggesting that simple feature-level fusion may be insufficient, potentially due to feature heterogeneity.
|
| |
| 11:30-12:30, Paper TH1130-1.4 | Add to My Program |
| Diffusion-Based Reinforcement Learning for Multi-Objective Adaptive Control in Autonomous Car-Following (I) |
|
| Ke, Weiling | Tongji University |
| Li, Shanghao | Beijing Institute of Technology |
| Tian, Daiying | Agency for Science, Technology and Research |
| Fang, Hao | Beijing Institute of Technology |
Keywords: Intelligent Transportation Systems, Robotics and Automation Applications, Deep Reinforcement Learning
Abstract: This paper investigates adaptive multi-objective control for autonomous car-following using model predictive control (MPC). Conventional MPC with predefined cost-function parameters has limited adaptability in balancing safety and comfort under varying traffic conditions. To address this limitation, an intent-driven adaptive MPC framework is proposed. A long short-term memory network is first employed to infer the driving intent of the preceding vehicle. Then, a modified twin delayed deep deterministic policy gradient (TD3) algorithm is developed, in which the conventional multilayer perceptron (MLP) actor is replaced by a diffusion policy actor. The diffusion actor generates a continuous action that adaptively configures the MPC cost weights. Simulation results on real-world NGSIM trajectories show that the proposed method achieves ride comfort comparable to that of the conventional MPC baseline. In emergency braking, the proposed method improves safety regulation over the MPC baseline. Compared with the MLP-based TD3 baseline, the proposed method produces smoother control responses, demonstrating the effectiveness of the diffusion-based actor for adaptive MPC.
|
| |
| 11:30-12:30, Paper TH1130-1.5 | Add to My Program |
| Elasto-Geometric Calibration of Serial Robots Using Joint Torque Sensing (I) |
|
| Zhou, Yaohua | Eastern Institute of Technology, Ningbo |
| Li, Xiaocong | Eastern Institute of Technology, Ningbo |
Keywords: Kinematics, Methodologies for Robotics and Automation, Modeling
Abstract: The loaded pose accuracy of lightweight collaborative robots is often degraded by both geometric errors and joint elastic deflections. Conventional kinematic calibration mainly compensates for geometry-induced deviations and is therefore insufficient to describe load-dependent pose errors. This paper proposes a joint torque sensing-based elasto-geometric calibration method for serial robots. A local product-of-exponentials based loaded kinematic model is first established, in which local geometric errors and torque-induced joint deflections are integrated into a unified formulation. The measured joint torques are then used to estimate elastic joint deflections through a joint compliance model. Based on the proposed model, a linearized calibration equation is derived, and the geometric and compliance parameters are simultaneously identified using damped least squares. Simulation results on a Franka Research 3 robot demonstrate the effectiveness of the proposed method.
|
| |
| TH1130-2 Invited Sessions, Ballroom II |
Add to My Program |
| Session 5.2: Embodied Perception, Navigation and Control |
|
| |
| Chair: Xu, Jiajun | Nanjing University of Aeronautics and Astronautics |
| Organizer: Yue, Yufeng | Beijing Institute of Technology |
| Organizer: Wang, Yuanzhe | Shandong University |
| Organizer: Song, Wenjie | Beijing Institute of Technology |
| Organizer: Wang, Yutang | Changchun Institute of Optics, Fine Mechanics and Physics, CAS |
| |
| 11:30-12:30, Paper TH1130-2.1 | Add to My Program |
| Terrain-Aware Heuristic for Hybrid A*: Low-Cost Path Planning in Off-Road Environments (I) |
|
| Wang, Kai | Beijing Institute of Technology |
| Wang, Meiling | Beijing Institute of Technology |
| Mao, Zihao | Beijing Institute of Technology |
| Song, Wenjie | Beijing Institute of Technology |
Keywords: Transportation Systems, Decision Support Systems
Abstract: Path planning for unmanned ground vehicles (UGVs) in off-road environments is hindered by unstructured terrains and the trade-off between kinematic feasibility and computational efficiency. To address this issue, a terrain-aware heuristic-based hybrid A* path planning method is proposed. First, a multi-layer terrain traversability cost map is constructed by fusing obstacle, slope and self-supervised learning-based bumpiness cost layers, providing comprehensive terrain information. Then, a terrain-aware A* algorithm generates an initial reference path, which is further embedded into an adaptively fused heuristic function for the hybrid A* algorithm. This design retains vehicle kinodynamic constraints while reducing computational complexity by leveraging terrain guidance. Real-world experiments on a UGV platform demonstrate that the proposed method outperforms standard HA*, T-HA* and SC-HA* in balancing terrain adaptability, path feasibility and planning efficiency, with the shortest computational time and fewer expanded nodes while achieving low terrain cost.
|
| |
| 11:30-12:30, Paper TH1130-2.2 | Add to My Program |
| Research on Simulated Human-In-The-Loop Reinforcement Learning Methods for Robotic Grasping Oriented to Sample Efficiency Improvement (I) |
|
| Wei, Zihe | Beijing Institute of Technology |
| Zhou, Tianxing | Beijing Institute of Technology |
| Zhou, Zichen | Beijing Institute of Technology |
Keywords: Embodied AI, Deep Reinforcement Learning, Planning and Control
Abstract: Robotic grasping based on reinforcement learning often suffers from low sample efficiency and unstable exploration. To address this issue, this paper proposes a learning framework for robotic grasping that combines simulated Human-in-the-Loop intervention, category-aware PCA-based reward shaping, adaptive reward fusion, and mixed demonstration and intervention replay. This method improves policy learning by integrating human corrective guidance with dense reward signals. The main contribution is not any single module in isolation, but a task-level integration in which intervention data, classifier features, confidence-gated shaping rewards, and replay composition are coupled in one closed training loop and evaluated under controlled component ablations. Experiments on three simulated robotic manipulation tasks show that the proposed framework achieves higher success rates and faster convergence than representative baseline methods. Ablation studies further confirm that all components contribute positively and the complete framework achieves the best overall performance.
|
| |
| 11:30-12:30, Paper TH1130-2.3 | Add to My Program |
| A Multi-Parameter Acquisition and Reproduction Framework for Human Manipulation Based on 6-DOF Bilateral Teleoperation (I) |
|
| Wang, Zhao | University of Chinese Academy of Sciences |
| Gao, Yang | Changchun Institute of Optics, Fine Mechanics and Physics, Chinese Academy of Sciences |
| Deng, Xu | Changchun Institute of Optics, Fine Mechanics and Physics (CIOMP), Chinese Academy of Sciences |
Keywords: Tele-robotics, Haptics, Force/Impedance Control
Abstract: Human manual manipulation skills are challenging to quantify due to their high-dimensional motion-force coupling and inter-operator variability. This paper presents a multi-parameter acquisition and reproduction framework for human manipulation based on a 6-DOF bilateral teleoperation robotic system. The proposed platform enables synchronous measurement of end-effector position, orientation, force, and torque in full six degrees of freedom. A comprehensive system is developed, encompassing kinematic modeling, Newton–Euler dynamic formulation, admittance-based bilateral control, gravity compensation, and passivity-based stability regulation is developed. To transform raw manipulation data into reproducible parametric representations, a polynomial-based trajectory fitting method integrated with adaptive smoothing and cycle extension is proposed. A force-position hybrid control scheme is implemented for motion reproduction. System transparency, interaction stability, gravity compensation accuracy, and reproduction smoothness are verified experimentally. Expert manipulation capture, tissue-simulated experiments, and preliminary human trials demonstrate that the proposed framework achieves reliable high-fidelity motion reproduction while guaranteeing safe interaction. The presented approach provides a systematic solution for quantitative modeling and robotic replication of human manipulation skills.
|
| |
| 11:30-12:30, Paper TH1130-2.4 | Add to My Program |
| HydrotheraBots: Integrating Robotics, Artificial Intelligence, and Extended Reality for Hybrid Land–Water Rehabilitation |
|
| Cabibihan, John-John | Qatar University |
| Sha, Mizaj Shabil | Qatar University |
Keywords: Medical Robots and Systems, Personal and Service Robotics, Virtual Reality
Abstract: Physical rehabilitation is essential for restoring mobility and functional independence following neurological and musculoskeletal disorders. Although hydrotherapy has exhibited significant benefits in terms of buoyancy-assisted movement and reduced joint loading, it's still underutilised due to therapist dependency, restricted facility availability, and high operational costs. This paper presents HydrotheraBots, a novel rehabilitation framework integrating robotic rollator systems, artificial intelligence (AI), multimodal sensing, and extended reality (XR) technologies for rehabilitation across both terrestrial and aquatic environments. The suggested architecture includes TerraBot-1, a land-based robotic rollator, and HydraBot-1, an underwater robotic rollator developed for hydrotherapy-assisted gait training. The system uses inertial measuring units (IMUs), electromyography (EMG), electroencephalography (EEG), force sensors, and stereoscopic vision to create a closed-loop rehabilitation ecology. AI-driven analytics continuously evaluate patient performance and adapt rehabilitation protocols in real time. The proposed framework solves the constraints of current rehabilitation technologies by allowing for a seamless transition between land and water therapy while also offering objective assessment and individualised intervention. The paper outlines the system architecture, design approach, and anticipated contributions to next-generation rehabilitation robotics.
|
| |
| 11:30-12:30, Paper TH1130-2.5 | Add to My Program |
| DR-TSLM: Deictic Reference Resolution Via Temporal-Spatial Large Language Model for Human-Robot Interaction |
|
| Tan, Runjia | Nanyang Technological University |
| Lou, Shanhe | A*STAR, Advanced Remanufacturing and Technology Centre |
| Yu, Lan | Guangzhou Cloudbutterfly Technology Co., Ltd |
| Tian, Xuesong | Guangzhou Cloudbutterfly Technology Co., Ltd |
| Lv, Chen | Nanyang Technological University |
Keywords: Embodied AI, Robotics and Automation Applications, Human-Robot Interfaces
Abstract: Human-robot interaction (HRI) is crucial for en-abling seamless collaboration between robots and humans in diverse applications, yet traditional methods often struggle with ambiguities in natural language instructions, particularly when non-verbal cues like co-speech referential gestures are involved. This paper introduces DR-TSLM, a novel multimodal system that leverages a fine-tuned large language model to ground human intentions spatial-temporally by integrating audio and video streams. DR-TSLM extracts word-level timestamps from speech to identify relevant video frames (temporal alignment) and maps ambiguous pronouns to specific visual regions via segmentation masks (spatial grounding), effectively resolving referential ambiguities without multi-round dialogues. To train and evaluate DR-TSLM, we designed a training methodology with pretraining and Chain-of-Thought (CoT) training phases, supported by three datasets: a 1,200-instance deictic words detection dataset for temporal alignment, an enhanced 100DOH dataset with 1,000 annotated scenes for spatial grounding, and a customized dataset with 100 speech-gesture pairs from 10 participants for CoT training. Experiments demonstrate improved performance over existing methods, advancing HRI by integrating linguistic and gestural modalities for enhanced efficiency and intuitiveness for non-expert users.
|
| |
| TH1430-1 Invited Sessions, Ballroom I |
Add to My Program |
| Session 6.1: Trustworthy Intelligence and Mobility Systems |
|
| |
| Chair: He, Xiangkun | University of Electronic Science and Technology of China |
| Co-Chair: Henglai, Wei | Beihang University |
| Organizer: He, Xiangkun | University of Electronic Science and Technology of China |
| Organizer: Henglai, Wei | Beihang University |
| Organizer: Hang, Peng | Tongji University |
| Organizer: Sun, Bohua | Jilin University |
| |
| 14:30-15:30, Paper TH1430-1.1 | Add to My Program |
| Improved YOLOv8-Based Ground Small Object Perception for UAV (I) |
|
| Yan, Yongjun | Nanjing University of Science and Technology |
| Yinyin, Cao | Nanjing University of Science and Technology |
| Peng, Zhang | Nanjing University of Science and Technology |
| Xuanjiang, Liu | Nanjing University of Science and Technology |
| Wang, Hongliang | Nanjing University of Science and Technology |
| Pi, Dawei | Nanjing University of Science and Technology |
Keywords: Image Processing, Transportation Systems
Abstract: To address the challenges of ground small-object detection in low-altitude UAV nadir-view scenes, such as small target scale, weak texture features, occlusion, and complex background interference, this paper proposes an improved YOLOv8-based method for UAV ground small-object detection. The network is lightweightly optimized by introducing CBAM and ECA into the backbone to enhance the representation of key regions and fine-grained features of small objects, adopting a PANet+ cross-scale feature fusion structure in the neck to improve multi-scale information interaction, and designing an improved decoupled detection head with an auxiliary objectness branch to improve localization accuracy, category discrimination, and small-object recall. Experiments on the VisDrone2019 dataset show that the proposed model achieves better detection performance and robustness in complex urban scenes. Compared with the baseline YOLOv8, Precision increases from 93.2% to 96.1%, mAP@0.5 improves from 55.2% to 64.5%, and Loss decreases from 0.48 to 0.08. The proposed method effectively improves the detection accuracy of UAV ground small objects while maintaining high real-time performance, providing support for UAV-based ground target perception and intelligent monitoring in complex scenes.
|
| |
| 14:30-15:30, Paper TH1430-1.2 | Add to My Program |
| Physics-Constrained Deep Residual Learning for Dynamics Modeling of Autonomous Racing Cars (I) |
|
| Zhao, Bolin | Nanyang Technological University |
| Shan, Zitong | Nanyang Technological University |
| He, Xianqi | Nanyang Technological University |
| Lou, Baichuan | Nanyang Technological University |
| Lv, Chen | Nanyang Technological University |
Keywords: Dynamics, Motion Control, Systems Modeling & Control
Abstract: Vehicle dynamics modeling at high speeds remains challenging for autonomous racing due to strongly nonlinear tire behavior and transient discontinuities induced by gear-shift events. These effects can degrade prediction accuracy and compromise physical consistency, motivating learning-based models that preserve the structure of vehicle dynamics while capturing unmodeled phenomena. This paper proposes a physics-constrained deep residual learning framework for high-speed dynamics modeling of autonomous racing cars. The method combines a physics-based backbone with a residual neural network that learns unmodeled dynamics while maintaining structural coherence. To handle gear-shift-induced wheel-speed discontinuities without violating vehicle-speed continuity, we introduce a shift-aware soft gating mechanism that selectively modulates wheel-speed residuals conditioned on the shift state. To further enforce physical plausibility, a constrained parameter system keeps learned dynamics-related parameters within realistic bounds during neural network training. In addition, the model incorporates a dynamic cornering-stiffness module with a two-stage degradation effect to capture tire nonlinearity under high lateral accelerations. Experiments on high-speed racing datasets show that the proposed model achieves high-fidelity predictions for both vehicle velocity and wheel speeds, and consistently outperforms competitive baselines, while remaining lightweight with only 220K parameters.
|
| |
| 14:30-15:30, Paper TH1430-1.3 | Add to My Program |
| Cross-Domain Cockpit Fusion in Autonomous Vehicles: Intelligent Decision-Making and Comfort Optimization Via Large Language Models (I) |
|
| Lan, Yuhao | Tongji University |
| Yan, Zhoudong | Tongji University |
| Hang, Peng | Tongji University |
Keywords: Large Language Models, Robotics and Automation Applications, Planning and Control
Abstract: With the rapid development of Autonomous Driving Systems (ADSs), passengers increasingly engage in non-driving tasks, e.g., using mobile phones, leading to higher motion sickness rates. To address this limitation, this paper proposes Comfort-LLM, a human-centered autonomous driving framework that leverages Large Language Models (LLMs) as the core reasoning engine. The proposed framework objectively evaluates comfort using the Motion Sickness Index (MSI) and integrates gaze detection to identify whether passengers are viewing the road or using mobile phones, thereby improving the in-cabin passenger experience. The LLM-based decision module dynamically generates adaptive control parameters, including target velocity, PID gains, and sigmoid trajectory planning coefficients. Experimental results in CARLA demonstrate that Comfort-LLM substantially improves passenger comfort: the MSI is reduced by up to 28.58%, confirming that controlling the chassis domain in ADS can effectively enhance the in-cabin experience.
|
| |
| 14:30-15:30, Paper TH1430-1.4 | Add to My Program |
| CL-DPL: Continual Learning Method Based on Dynamic Prompts for Driver Distraction Detection (I) |
|
| Yang, Lie | Nanchang University |
| Jin, Nan | Nanchang University |
| Li, Jing | Nanchang University |
| Yang, Yan | Nanchang University |
| Chen, Xuanyu | Nanchang University |
| Zhang, Tingfang | Nanchang University |
Keywords: Transportation Systems
Abstract: Driver distraction detection is critical to road traffic safety, and vision-based methods have become mainstream owing to their non-intrusive nature. However, real-world driving scenarios are dynamic, necessitating models that continuously adapt to new tasks. Traditional deep learning methods suffer from catastrophic forgetting when learning sequentially. Existing prompt learning approaches alleviate this issue to some extent but rely on a fixed-capacity prompt pool and retrieval mechanism. This limits their ability to capture subtle variations in driving scenes, resulting in coarse-grained recognition. To address these limitations, this paper proposes a continual learning framework based on dynamic prompts learning (CL-DPL), which uses a lightweight convolutional network to generate input-dependent prompts in real time, which guide a frozen Vision Transformer (ViT) to efficiently adapt to new tasks. Binary masks isolate task-specific parameters to prevent inter-task interference, while an Additive Angular Margin Penalty (AAMP) enhances feature discriminability. Experiments on CIFAR-100, State-Farm, and SAM-DD show that CL-DPL outperforms existing methods such as L2P and EWC in both higher average accuracy and lower forgetting rate. Ablation studies confirm the effectiveness of each core component. Moreover, the proposed method incurs extremely low parameter overhead, enabling forgetting-resistant continual learning with strong potential for vehicle-edge deployment.
|
| |
| 14:30-15:30, Paper TH1430-1.5 | Add to My Program |
| Serial Cycle Consistent Video Joint Embedding Predictive Architecture with Phase Wise Optimization |
|
| Yan, Fei | Jilin University |
| Miao, Chengyu | Jilin University |
| Yao, Zhiyi | Changchun University of Science and Technology |
| Han, Wei | Jilin University |
| Liu, Yu | Jilin University |
| Guo, Hetian | Southern University of Science and Technology |
| Huang, Tianlv | Jilin University |
| Wang, Tianshuo | High School Attached to Northeast Normal University |
| Fan, Zipei | Jilin University |
| Song, Xuan | Jilin University |
Keywords: Embodied AI, Deep Learning, Robotics and Automation Applications
Abstract: Joint embedding predictive architectures offer a compelling alternative to pixel space reconstruction for visual self-supervised learning, especially in video. Existing video joint embedding predictive architecture methods primarily emphasize one-way masked prediction, where the model predicts masked target representations from visible context. However, the predicted latent tokens are not explicitly constrained to preserve enough information to recover a more complete semantic target. This allows shortcut solutions in which predicted tokens are sufficient for local matching while discarding global structure. In this work, we propose a serial cycle-consistent extension to video joint embedding predictive architecture. The method first performs standard forward masked prediction in latent space and then applies a backward predictor to reconstruct full target features from the predicted masked tokens. This yields a cycle-consistency regularizer that encourages structurally recoverable latent representations. To stabilize optimization, we adopt a two-stage training procedure. Phase 1 warms up the backward branch under a frozen encoder and forward predictor, and Phase 2 jointly optimizes both objectives with an exponential moving average teacher. We further evaluate the learned representations with frozen probe protocols on downstream video classification and action anticipation tasks, using only the pretrained encoder as the transferred backbone. Our experiments show that the proposed architecture converges stably without representation collapse and learns robust spatiotemporal representations that better preserve global semantic structure. More broadly, by improving the predictive usefulness and recoverability of latent video representations, the proposed framework highlights the potential of JEPA-style world models for action understanding, anticipation, and future embodied robotic perception and decision making.
|
| |
| TH1430-2 Invited Sessions, Ballroom II |
Add to My Program |
| Session 6.2: Emerging Paradigms in Robotics and Autonomous Systems |
|
| |
| Chair: Nie, Zifei | Jilin University |
| Organizer: Hu, Yunfeng | Jilin University |
| Organizer: Sun, Zhongbo | Changchun University of Technology |
| Organizer: Jin, Long | Lanzhou University |
| Organizer: Sun, Ning | Nankai University |
| |
| 14:30-15:30, Paper TH1430-2.1 | Add to My Program |
| DPRF-CLIP: Dual-Path Residual Fusion for Zero-Shot Industrial Anomaly Detection (I) |
|
| Liu, Shuaishi | Changchun University of Technology |
| Ma, Tianjiao | Changchun University of Technology |
Keywords: Deep Learning, Image Processing
Abstract: 零射点异常检测(ZSAD)旨在实现 从未见过的类别中,精心识别异常样本 无需依赖目标类训练数据。在 工业质量 检查场景,收集训练样本 目标缺陷 由于生产问题,类别往往不切实际 约束和 数据稀缺性,以及ZSAD方法可以有效解决 查尔 在有限数据下可靠异常检测的长度 条件。 近年来,视觉语言模型表现出强劲 通用化 极大地提升了TION和固有的零发能力 促进他们的 在零射工业异常检测中的广泛应用 任务 具有竞争力且可靠的检测性能。 然而, 它们存在关键的实际局限:不足 注意 f 图像细节偏离颗粒,且适应性较差 具体情况 工业异常检测任务的要求。前往 地址 基于这些限制,我们提出了基于CLIP的ZSAD DPRF-CLIP 框架。它使用预训练的ResNet网络进行提取 好吧 颗粒
|
| |
| 14:30-15:30, Paper TH1430-2.2 | Add to My Program |
| Multimodal Sentiment Analysis Via Modal Experts and Missing-Modal Prompt Generation (I) |
|
| Liu, Shuaishi | Changchun University of Technology |
| Wu, Xuanyu | Changchun University of Technology |
Keywords: Deep Learning, Natural Language Processing
Abstract: Multimodal sentiment analysis aims to extract affective signals from multiple modalities and perform feature analysis, modality fusion and sentiment prediction. However, in practical application scenarios, input modality information is highly prone to be missing due to various objective factors, which further causes the model to produce biased or even completely erroneous sentiment judgment results. To address this critical issue, a novel model named Modal Experts and Missing Prompt Generation (MEMPG) is proposed. The model employs a brand-new dual-residual modality expert architecture to integrate the knowledge of hybrid modality experts. This architecture can not only retain the original information but also preserve the contextual information learned by the attention mechanism, thus exhibiting better robustness. The model adopts a two-stage training strategy: In the first stage, each modality expert branch is independently pre-trained to extract low-level basic features of a single modality and strengthen the feature extraction capability of modality experts. In the second stage, all input modality features are fed into all modality expert branches for processing, and adaptive weights are assigned to modality features via an adaptive gating mechanism to obtain fused modality features with stronger representation ability. Subsequently, the enhanced features are input into the missing modality prompt generation module, which incorporates prompt learning to guide the reconstruction of missing modalities. Extensive comparative experiments are conducted on two standard multimodal sentiment datasets, namely MOSI and MOSEI, under various common modality missing scenarios. Experimental results demonstrate that, compared with current mainstream baseline models, the proposed MEMPG model achieves superior sentiment prediction performance.
|
| |
| 14:30-15:30, Paper TH1430-2.3 | Add to My Program |
| Variable Impedance Control Method for Upper Limb Rehabilitation Robot Based on Fuzzy Network (I) |
|
| Pang, Zaixiang | Changchun University of Technology |
| Liu, Guangjun | Changchun University of Technology |
| Wang, Yaxuan | Changchun University of Technology |
| Li, Shuang | Changchun University of Technology |
Keywords: Force/Impedance Control, Medical Robots and Systems, Dynamics
Abstract: Rehabilitation robots can reproduce many movements in traditional rehabilitation training through control methods designed for different types of rehabilitation robots, so as to achieve the goal of training the motor ability of the affected area of patients. In this process, encouraging the enthusiasm of patients for training is crucial to the effectiveness of rehabilitation treatment. However, due to the differences in patients' motor abilities, the design of personalized control in robot - assisted treatment remains challenging. To solve this problem, the controller designed in this paper conducts dynamic modeling based on the Lie group theory configuration during the design process, which solves the problems of singularity and non - globality that usually exist in local coordinates. The concept of patient ability factor is introduced, and through comprehensive consideration of multiple aspects, the controller can adaptively switch control modes. In order to be closer to the real training situation, a muscle fatigue model is introduced to simulate the real - time situation of patients' muscle fatigue and recovery. In terms of impedance adjustment, a fuzzy neural network is incorporated to ensure personalization and active participation. In terms of data learning and updating, an adaptive and iterative learning mechanism is adopted to enhance convergence. In terms of safety, multiple safety constraints are integrated to ensure the safety of human - machine interaction. The stability of the designed control system is verified using Lyapunov theory. Finally, through MATLAB simulation experiments, the trajectory tracking results are obtained, and the errors are all within the ideal range, which verifies the feasibility of the control system.
|
| |
| 14:30-15:30, Paper TH1430-2.4 | Add to My Program |
| A Zeroing Neural Dynamics-Based Predictive Control Algorithm for Manipulator Trajectory Tracking and Obstacle Avoidance (I) |
|
| Zhang, Zhishuo | Jilin University |
| Tang, Shijun | Jilin University |
| Chong, Zhang | Jilin University |
| Hu, Yunfeng | Jilin University |
Keywords: Planning and Control
Abstract: To address the problem of robotic manipulator trajectory tracking in complex environments subject to obstacle constraints and multilevel joint constraints, this paper proposes a trajectory tracking and obstacle avoidance control method based on model predictive control (MPC) and zeroing neural dynamics (ZND). First, a unified optimization model is established by simultaneously considering trajectory tracking error, control input smoothness, obstacle avoidance requirements, and multilevel joint constraints, so as to achieve safe motion control of the manipulator in complex environments. Then, for the resulting time-varying optimization problem, a corresponding ZND-based online solver is designed to dynamically update the optimization variables in real time, thereby improving the computational efficiency of MPC and satisfying the real-time requirement of the control system. The proposed method enables the end-effector to accurately track the desired trajectory while effectively avoiding collisions with obstacles and satisfying multilevel joint constraints. Simulation results demonstrate that the proposed method achieves favorable performance in terms of tracking accuracy, obstacle avoidance safety, and online computational efficiency.
|
| |
| TH1600-1 Invited Sessions, Ballroom I |
Add to My Program |
Session 7.1: Perception, Control and Learning for Intelligent Autonomous
Systems |
|
| |
| Chair: Gong, Xun | Jilin University |
| Organizer: Hu, Yunfeng | Jilin University |
| Organizer: Li, Yongfu | Chongqing University of Posts and Telecommunications |
| Organizer: Ren, Bingtao | Beihang University |
| Organizer: Gong, Xun | Jilin University |
| |
| 16:00-17:00, Paper TH1600-1.1 | Add to My Program |
| Temporal Cross-Modal Alignment Attack against Perception-Oriented Vision-Language Models for Autonomous Driving (I) |
|
| Chen, Xianglong | Jilin University |
| Hu, Yunfeng | Jilin University |
| Ma, Bin | College of Communication Engineering, Jilin University |
| Zou, Bosong | China Software Testing Center |
Keywords: Cybernetics Automation and Control, Computational Intelligence, Systems Modeling & Control
Abstract: Vision-language models (VLMs) have shown strong performance in autonomous driving (AD) tasks, supporting scene understanding and safety-related multimodal reasoning. However, robustness under adversarial perturbations remains critical, and the alignment vulnerability between visual evidence and task semantics under sequential observations is insufficiently explored. This paper proposes TCMA, a Temporal Cross-Modal Alignment Attack combining a task-oriented objective, an alignment disruption loss, and a lightweight temporal propagation mechanism to attack perception-oriented VLMs in AD. Specifically, TCMA constructs a semantic anchor from the task prompt to suppress correct visual-text alignment, and warm-starts each frame's attack from the previous perturbation while enforcing temporal consistency. On BDD100K with Dolphins, TCMA achieves 50.0% overall center-frame targeted ASR across traffic-light, pedestrian, and rider tasks. Transfer evaluation on Qwen2.5-VL further reaches 73.3% center-frame and 76.7% vote-level overall targeted ASR, demonstrating strong cross-model generalizability.
|
| |
| 16:00-17:00, Paper TH1600-1.2 | Add to My Program |
| SmartHerd: An IoT-Enabled Deep Representation Learning Framework for UWB Trajectory Clustering of Dairy Cows |
|
| Yao, Lei | Jilin University |
| Kong, Fanrong | Jilin University |
| Liu, Jin | Jilin University |
| Hong, Weinan | Jilin University |
| Fan, Zipei | Jilin University |
| Lei, Lin | Jilin University |
| Li, Xinwei | Jilin University |
Keywords: Neural Networks, Computational Intelligence, Internet of Things
Abstract: The automated acquisition and processing of spatiotemporal data via wearable Internet of Things (IoT) devices is crucial for Precision Livestock Farming (PLF). While Ultra-Wideband (UWB) sensors provide high-fidelity trajectory coordinate streams, bridging the semantic gap between this non-linear raw data and discrete behavioral states remains a significant computational challenge. This paper presents SmartHerd, an end-to-end IoT computational framework for unsupervised trajectory clustering. We empirically compare two distinct analytical paradigms: a baseline pipeline relying on handcrafted kinematic features, and a deep representation learning approach. The proposed deep learning architecture leverages a Long Short-Term Memory (LSTM) network to capture long-range temporal dependencies, integrated with a Variational Autoencoder (VAE) to map these dynamics into a probabilistically regularized latent manifold. Topological analysis of the resulting embeddings indicates that the LSTM-VAE framework produces more cohesive and better-separated clusters than the feature-based baseline, with clusters semantically associated with key behaviors including resting, feeding, and walking. These findings suggest that deep representation learning offers a promising computational basis for decoding stochastic IoT trajectory data and supporting automated welfare monitoring.
|
| |
| 16:00-17:00, Paper TH1600-1.3 | Add to My Program |
| HuMam: Humanoid Motion Control Via End-To-End Deep Reinforcement Learning with Mamba (I) |
|
| Wang, Yinuo | Woven by Toyota, US. Inc |
| Tao, Xiaowen | Trinity College Dublin |
| Hu, Yunfeng | Jilin University |
| Liang, Yang | Wuhan University of Technology |
| Song, Ziyu | Jilin University |
| Ding, Haitao | Jilin University |
| Zhou, Jinzhao | University of Technology Sydney |
Keywords: Legged Robots, Motion Control, Deep Learning
Abstract: End-to-end reinforcement learning for humanoid locomotion remains challenging due to training instability, inefficient multimodal feature fusion, and high actuation cost. We present HuMam, a state-centric end-to-end RL framework that employs a single-layer Mamba encoder to fuse robot-centric proprioceptive states with oriented footstep targets and a continuous phase clock. The policy outputs joint position targets executed by a low-gain proportional–derivative controller, and is optimized via proximal policy optimization with a six-term reward balancing contact quality, swing smoothness, foot placement accuracy, posture, body height, and upper-body stability. Experiments on the JVRC-1 humanoid in mc-mujoco across forward, backward, lateral, curved walking, and standing tasks demonstrate that HuMam consistently outperforms a strong feedforward baseline: it improves peak return by 5.8%, reduces samples required to reach target performance by up to 42.5%, lowers cross-seed training variance by 61.0%, and reduces average power consumption by 28.1%. To our knowledge, this is the first end-to-end humanoid RL controller to adopt Mamba as the fusion backbone, demonstrating tangible gains in learning efficiency, training stability, and control economy.
|
| |
| 16:00-17:00, Paper TH1600-1.4 | Add to My Program |
| End-To-End Autonomous Parking Method Fusing CoViT Visual Encoding and State Transition-Assisted SAC (I) |
|
| Du, Fengjie | School of Transportation Science and Engineering, Beihang University |
| Yang, Shichun | Beihang University |
| Ren, Bingtao | Beihang University |
| Wang, ChuanYe | BeiHang University |
| Cao, Yaoguang | Beihang University |
| Guo, YuFan | BeiHang University |
Keywords: Deep Reinforcement Learning, Neural Networks, Embodied AI
Abstract: Aiming at the problems that (convolutional neural networks, CNNs) in traditional end-to-end autonomous parking methods are difficult to fully model global scene semantics, and the single (Soft Actor-Critic, SAC) algorithm is limited by sample utilization and training stability, this paper proposes an end-to-end autonomous parking framework that fuses a CoViT (Convolutional Neural Network-Combined Vision Transformer, CoViT) visual encoder and a VAE (variational autoencoder, VAE)-based state transition model. This method takes CoViT as the core of environmental perception to conduct local–global collaborative modeling of visual features: the global semantic information extracted by Vision Transformer from deep features is used as the query vector, and the shallow convolutional features are used as keys and values. Multi-layer feature fusion is realized through a cross-attention mechanism, thereby enhancing the representation ability of parking space boundaries, obstacles, and scene structure information. At the decision-making level, the state features encoded by CoViT are input into the SAC continuous control framework, and a VAE-based probabilistic state transition model is introduced as an auxiliary environmental dynamics modeling module. Joint optimization of policy learning and state transition prediction improves the sample efficiency and training stability of the model. Finally, a closed-loop vertical parking experiment is constructed on the CARLA simulation platform, and verification is carried out in static senario and senario with random dynamic pedestrian interference respectively. The experimental results show that the success rates of the proposed method in the two types of scenes reach 93.75% and 88.54% respectively, demonstrating superior parking accuracy, robustness, and dynamic environment adaptability
|
| |
| 16:00-17:00, Paper TH1600-1.5 | Add to My Program |
| Fast Safety-Critical Constrained Optimization for Integrated Vehicle Predictive Motion Control |
|
| Li, Zihan | Jilin University |
| Wang, Ping | Jilin University |
| Liu, Hanghang | Jilin University |
| |
| TH1600-2 Invited Sessions, Ballroom II |
Add to My Program |
| Session 7.2: Autonomous and Flexible Robotics towards Surgery |
|
| |
| Chair: Fang, Ge | Nankai University |
| Organizer: Fang, Ge | Nankai University |
| Organizer: Wang, Xiangyu | Nankai University |
| |
| 16:00-17:00, Paper TH1600-2.1 | Add to My Program |
| A Simultaneous Calibration Method for Dual-Arm Robots Based on Optical Trackers (I) |
|
| Min, Kang | Harbin Institute of Technology |
| Shi, Yudong | Hefei University of Technology |
| Mo, Hangjie | City University of HongKong |
| Li, Xiaojian | Hefei University of Technology |
| Hu, Xianghong | China Electronic ProductReliability and EnvironmentalTesting Research Institute |
| Yang, Shanlin | Hefei University of Technology |
Keywords: Kinematics, Modeling, Robotics and Automation Applications
Abstract: This paper presents a unified calibration method for the full parameters of the dual-arm robotic system, encompassing hand-eye, kinematic, and tool center point (TCP) parameters. Firstly, the closed-loop equation, AXP=YCQ, is established in three-dimensional space and solved by neglecting kinematic parameter errors, thereby obtaining the initial parameters. Subsequently, a calibration model incorporating all parameter errors is established and formulated as a constrained nonlinear least-squares optimization problem, thereby significantly enhancing the calibration precision and accuracy. The main contributions include: (1) constructing a spatial closed-loop equation and utilizing the Kronecker product for initial parameter calculation, which minimizes human annotation errors while improving data acquisition efficiency; and (2) establishing a unified error model that treats calibration as an optimization task to further refine accuracy. Simulation results demonstrate that the proposed method substantially suppresses positioning errors. In generalization tests, the maximum, mean, and root mean square (RMS) values of absolute position errors are reduced from 8.3871 mm, 3.1416 mm, and 3.4662 mm to 0.2983 mm, 0.1442 mm, and 0.1580 mm, respectively, validating the effectiveness and generalization capability of the method.
|
| |
| 16:00-17:00, Paper TH1600-2.2 | Add to My Program |
| LLM-Based Hierarchical Control Architecture for Robotic Operation |
|
| Ma, Mingyang | Jilin University |
Keywords: Computational Intelligence, Cooperative Systems and Control, Blockchain Technologies
Abstract: 抽象——非结构化环境中的机器人操作 仍是一个长期存在的关键挑战,驱动力是 固有的环境不确定性、任务复杂性以及 高层语义的冲突需求 推理和低层实时反应控制。虽然 大型语言模型(LLM)展现了卓越的表现 自然语言理解与常识推理 能够直接部署为端到端机器人 手柄存在严重的限制,包括高频 推断延迟,实时响应能力不足, 以及模型幻觉可能带来的安全风险。 为解决这些问题,本文提出了一本小说 基于LLM的三层分层控制架构, 这使大型语言模型的认知优势与 传统控制系统的实时可靠性, 设计了动态的闭环反馈机制 重新规划。多重实验 代表性的机器人任务验证所提议的 方法的任务成功率可达94%。 比最先进基线高出7–12%, 严格遵守实时性೦
|
| |
| 16:00-17:00, Paper TH1600-2.3 | Add to My Program |
| Uncertainty-Resilient Fuzzy Reinforcement Learning for Autonomous Endovascular Navigation of Magnetic Vascular Catheter Robots (Special Section Code 7h974) (I) |
|
| Li, Zhengyang | Nankai University |
| Fang, Ge | Nankai University |
| Qin, Yanding | Nankai University |
| Han, Jianda | Nankai University |
Keywords: Medical Robots and Systems, Deep Reinforcement Learning, Fuzzy Systems
Abstract: Magnetic vascular catheter robots (MVCRs) play a crucial role in minimally invasive vascular interventions, but their autonomous navigation faces significant challenges due to the complex and dynamic vascular environment, uncertain magnetic actuation, and strict safety constraints. To address these issues, this paper proposes a fuzzy reinforcement learning (FRL) algorithm for the autonomous navigation of MVCRs. The proposed FRL algorithm integrates fuzzy logic (FL) with proximal policy optimization (PPO) to handle the ambiguity and uncertainty in vascular navigation tasks, while balancing navigation efficiency and safety. First, a vascular environment model considering hemodynamic effects and vascular anatomical characteristics is established, and the kinematic model of the MVCR is derived based on magnetic actuation principles. Then, the fuzzy logic system is designed to fuzzify the state space (including catheter position, orientation, vascular curvature, and blood flow velocity) and dynamically adjust the reward function of PPO, which enhances the algorithm’s adaptability to complex vascular environments. Finally, extensive simulations based on the Simulation Open Framework Architecture (SOFA) platform and in vitro experiments are conducted to verify the performance of the proposed algorithm. The results show that compared with traditional PPO and current RL algorithms, the proposed FRL algorithm achieves a higher navigation success rate (97.3%), lower collision force, and a shorter navigation time in complex vascular models. The algorithm can effectively adapt to vascular anatomical variations and dynamic blood flow changes, providing a reliable solution for the autonomous navigation of MVCRs in clinical interventional procedures.
|
| |
| 16:00-17:00, Paper TH1600-2.4 | Add to My Program |
| Active-Model Compensated Model Predictive Control for Ultrasonic Motors in MRI-Compatible Surgical Robots (I) |
|
| Xu, Ke | Nankai University |
| Wang, Longxin | Nankai University |
| Fang, Ge | Nankai University |
| Qin, Yanding | Nankai University |
| Wang, Hongpeng | Nankai University |
| Han, Jianda | Nankai University |
Keywords: Medical Robots and Systems, Motion Control, Robotics and Automation Applications
Abstract: To address the challenge of high-precision control of ultrasonic motors (USMs) under complex loads in magnetic resonance imaging (MRI)-compatible surgical robots, this paper proposes active-model compensated model predictive control (AM-MPC), a novel control strategy integrating MPC controller with active-model-based disturbance compensation. This method adopts Kalman filter to estimate external disturbances and unmodeled nonlinearities as an extended state, which can be dynamically accommodated via feedforward compensation. Experimental results of ultrasonic motor control demonstrate that under a heavy load of 1.047kg, the proposed strategy reduces the root mean square error (RMSE) and maximum absolute error (MAE) by 27.73% and 25.79%, respectively, compared to traditional MPC. Furthermore, the proposed algorithm was successfully applied in an MRI-compatible robot to perform simulated puncture surgery in agar phantom, verifying its feasibility in neurosurgical application.
|
| |
| 16:00-17:00, Paper TH1600-2.5 | Add to My Program |
| Design and Validation of an MRI-Compatible Puncture Needle with Fiber Bragg Grating Force Sensing for Neurosurgery (I) |
|
| Wang, Longxin | Nankai University |
| Guo, Hongzhe | NANKAI University |
| Fang, Ge | Nankai University |
| Qin, Yanding | Nankai University |
| Wang, Hongpeng | Nankai University |
| Han, Jianda | Nankai University |
Keywords: Sensor Design, Medical Robots and Systems
Abstract: In brain stereotactic surgery, real-time and highprecision measurement of the interaction forces between the puncture needle and brain tissue is essential for improving surgical accuracy and ensuring safety. However, traditional force-sensing technologies have limited biocompatibility and are susceptible to electromagnetic interference, making them unsuitable for application in MR (Magnetic Resonance) environment. To overcome these limitations, an MRI (Magnetic Resonance Imaging)-compatible puncture needle with FBG force sensor is presented in this work. Two Fiber Bragg Gratings (FBGs) are suspended within the force-sensing module at the proximal end of the puncture needle, serving to measure axial force and achieve temperature compensation, respectively. Finite element modeling (FEM)-based simulation was implemented to evaluate the static performance of the puncture needle. Calibration experiments were conducted to verify simulation outcomes. The results indicate that the designed puncture needle achieves a high force resolution of 5.95 mN. The results of temperature compensation tests validated the effectiveness of the proposed compensation algorithm. In addition, puncture experiments in agar phantom demonstrated the feasibility of proposed needle in neurosurgical applications.
|
| |
| TH1700-1 Regular Sessions, Ballroom I |
Add to My Program |
| Session 8.1: Deep Learning |
|
| |
| Chair: Ding, Shihong | Jiangsu University |
| |
| 17:00-18:00, Paper TH1700-1.1 | Add to My Program |
| Text-Supervised Semantic Segmentation for User-Specified Robotic Cloth Grasping |
|
| Tu, Zhiwen | Jilin University |
| Zhu, Xingyu | Jilin University |
| Zhang, Jiqian | Jilin University |
| Feng, Runyang | Zhejiang Gongshang University |
| Gao, Yixing | Jilin University |
Keywords: Robot Vision, Deep Learning, Personal and Service Robotics
Abstract: User-specified robotic cloth grasping, which requires robots to grasp specific pieces of cloth according to user instructions, is a fundamental task for household robots. Existing approaches typically rely on closed-set classification or fully supervised semantic segmentation methods that demand expensive pixel-level annotations, limiting their scalability and generalization to unseen cloth categories. In this paper, we propose a text-supervised semantic segmentation framework that learns to segment cloth regions using only image-level textual descriptions, completely eliminating the need for pixel-wise labels. To overcome the scarcity of paired image--text cloth data, we leverage large vision--language models (LVLMs) to generate structured textual annotations for cloth images, constructing a dataset of 20 common cloth categories with 8,972 image--text pairs. Built upon CLIP, our framework introduces a Text Adapter and an Image Adapter to bridge the domain gap to the cloth feature space, and employs a dual contrastive learning mechanism for both global sentence--image alignment and local region--noun alignment. A Pixel-based Refinement Module (PRM) further enhances segmentation boundary quality. On public benchmarks, our method achieves mIoU scores of 58.0%, 29.1%, and 32.4% on PASCAL VOC, Cityscapes, and COCO Object, respectively, establishing competitive results among text-supervised methods. We further integrate the segmentation model into a complete robotic grasping system, achieving an 87.0% grasping success rate on 16 unseen cloth items in 300 real-world trials. The dataset is publicly available at https://github.com/ubun23/ClothesDescription.
|
| |
| 17:00-18:00, Paper TH1700-1.2 | Add to My Program |
| A Stable Self-Guided Framework for Semi-Supervised Semantic Segmentation |
|
| Zhang, Jiqian | Jilin University |
| Yin, Yue | Jilin University |
| Gao, Yixing | Jilin University |
Keywords: Deep Learning, Image Processing, Neural Networks
Abstract: In semi-supervised semantic segmentation (SSSS), self-guided frameworks have become mainstream and are widely used in many state-of-the-art methods. They leverage pseudo-labels generated by the model itself at each iteration as learning targets for consistency regularization and typically divide the task into two branches: a supervised branch with labeled data and an unsupervised branch with unlabeled data. However, such a paradigm essentially suffers from label ambiguity in the unsupervised branch, where the model may generate substantially disparate even contradictory pseudo-labels for a same visual sample.Directly leveraging such unstable guidance signals dramatically hinders the convergence of the unsupervised loss and leads to overfitting on the labeled data. To address this issue, we propose a stable self-guided framework for semi-supervised semantic segmentation. We first introduce a Knowledge Bank to preserve the knowledge learned by previous models during training, and then engage a Knowledge-Constrained Guidance Signal Modulation method that utilizes prior knowledge from the Knowledge Bank to refine the guidance signals of the current model. Within this framework, a pixel in pseudo-labels will be selected as the learning target for the unsupervised branch only when it achieves consensus between multiple sophisticated models from the Knowledge Bank and the current model. This limits the drift of the target distribution in the unsupervised branch and significantly improves the quality of pseudo-labels, while ensuring the effectiveness of consistency regularization.Extensive experiments demonstrate that our proposed framework shows superior performance across nearly all evaluation protocols on Pascal, Cityscapes, and COCO benchmarks.
|
| |
| 17:00-18:00, Paper TH1700-1.3 | Add to My Program |
| Defect-Aware Neural Compression for Cable Joint Surface Defect Images |
|
| Liu, Bin | China Energy Engineering Group Shaanxi Electric Power Design Institute Co., Ltd |
| Liu, Gen | State Grid Shaanxi Construction Branch |
| Yunxia, Haoyue | State Grid Shaanxi Electric Power Research Institute |
| Zhao, Xuefeng | State Grid Shaanxi Electric Power Research Institute |
| Zheng, Xinyu | Xi'an Jiaotong University |
| Ge, Chenyang | Xi'an Jiaotong University |
Keywords: Deep Learning, Neural Networks
Abstract: In recent years, with the advancement of deep learning technologies, neural network-based image compression methods have surpassed many traditional coding standards in rate–distortion performance. However, existing approaches are primarily designed for natural images and perform suboptimally when directly applied to cable joint surface defect images. This is because defect images are typically weak-textured, characterized by a high proportion of low-frequency background and sparse high-frequency defect regions, which significantly differ from the statistical properties of natural images. To address this, this paper proposes a defect-aware neural image compression method specif ically optimized for such characteristics. The method consists of three core modules: a Feature Complementary Module (FCM) that collaboratively exploits local and global information redun dancy; a Frequency Gating Module (FGM) that adaptively filters and modulates low-frequency and high-frequency features; and a Defect-Aware Attention (DAA) module that promotes rational bit allocation. Experiments on a self-built cable defect dataset demonstrate that the proposed method outperforms existing mainstream compression approaches in both quantitative metrics and visual quality, validating its effectiveness and superiority.
|
| |
| 17:00-18:00, Paper TH1700-1.4 | Add to My Program |
| Identification of Nonlinear Dynamic Networks Via Integral-Based Sparse Bayesian Learning and Stability Screening |
|
| Yao, Yige | Huazhong University of Science and Technology |
| Zheng, Yaozhong | Huazhong Univeristy of Science and Technology |
| Shi, Ran | Huazhong University of Science and Technology |
| Zhang, Hai-Tao | Huazhong University of Science AndTechnology |
Keywords: Networked Dynamical Systems, Systems Modeling & Control, Computational Intelligence
Abstract: Reconstructing the topology and internal dynamics of nonlinear networked systems from noisy observational data remains a challenging inverse problem. To tackle this challenge, we develop a robust two-stage identification framework driven by an integral formulation. To mitigate the noise amplification inherent in numerical differentiation, we construct a regression dictionary in integral form via the trapezoidal rule, transforming parameter estimation into a more robust integral problem. Subsequently, a decoupled structure screening and parameter refinement strategy is deployed. In the first stage, sparse Bayesian learning (SBL) recovers the sparse network structure. By leveraging the prior that multi-dimensional states share governing equations, we introduce a group feature aggregation mechanism to ensure consistent feature retention and prevent excessive sparsification. In the second stage, constrained ordinary least squares (OLS) yields unbiased parameter estimates strictly for the identified active basis functions. Finally, a statistical stability criterion is applied to iteratively prune noise-induced spurious terms. Simulations across Barabási-Albert networks with diverse nonlinearities validate that our method robustly identifies both adjacency matrices and dynamic parameters, showing strong generalizability.
|
| |
| TH1700-2 Regular Sessions, Ballroom II |
Add to My Program |
| Session 8.2: Robotics Systems |
|
| |
| Chair: Henglai, Wei | Beihang University |
| |
| 17:00-18:00, Paper TH1700-2.1 | Add to My Program |
| Hybrid Subgraph-Vector Retrieval with Planner-Solver-Critic for Wastewater Treatment Plant Operation and Maintenance Assistance |
|
| Pang, Long | China University of Geosciences |
| Ng Man-po, Jeff | Drainage Services Department, the Government of the HKSAR |
| Wong, Tsz-kin | Drainage Services Department |
| Wong, Kevin Chi Wai | Drainage Services Department |
| Hu, Wenkai | China University of Geosciences |
| Zhang, Lijun | Ningbo StrideTech Co., Ltd |
Keywords: Large Language Models, Decision Support Systems, Resource Management
Abstract: Operation and maintenance (O&M) of wastewater treatment plants relies on scattered manuals, procedures, expert knowledge, and historical records. Existing Retrieval-Augmented Generation methods struggle to jointly use structured relations and detailed texts for providing solutions to scenarios that require multi-step reasoning upon the retrieved information. This paper proposes a hybrid subgraph-vector RAG framework with a Planner-Solver-Critic answer generation mechanism to address the aforementioned gap. The contributions are twofold. First, a multi-path subgraph construction and fusion method combined with vector retrieval is developed. Second, a PSC generation scheme is designed to enforce explicit planning and iterative reasoning over the retrieved evidence. Experiments on a wastewater treatment O&M corpus demonstrate that the proposed framework improves both retrieval accuracy and answer quality.
|
| |
| 17:00-18:00, Paper TH1700-2.2 | Add to My Program |
| MPPI Cost Function Design for Quadrupedal Locomotion Via Hierarchical LLM-CEM Optimization |
|
| Tao, Chuyuan | Shandong Univeristy |
| Wu, Hanbo | Shandong Univeristy |
| Wang, Fanxin | Xi'an Jiaotong-Liverpool University (XJTLU) |
| Jiang, Haolong | Xi'an Jiaotong-Liverpool University (XJTLU) |
| Li, Zehao | Shandong Univeristy |
| Gao, Yuefeng | China Logistics Group Automotive Supply Chain Technology Co., Ltd |
Keywords: Legged Robots, Deep Learning, Motion Control
Abstract: Designing cost functions for legged robot locomotion in MPC is challenging. Balancing stability, agility and energy efficiency requires expert knowledge. While MPPI performs well in complex contact-rich dynamics, it depends on calibrated cost parameters. Direct parameter learning is computationally expensive, leading to low sample efficiency in naive optimization and reinforcement learning. In this work, we propose a hierarchical framework that leverages large language models (LLMs) to provide semantically meaningful initializations of cost parameters from high-level task descriptions, and refines them using the Cross-Entropy Method (CEM) based on closed-loop performance. The LLM serves as a structured prior that encodes task-relevant objectives, while CEM performs data-driven optimization in the parameter space, with MPPI acting as the inner-loop trajectory optimizer. This approach addresses the limitations of both manual cost design and purely data-driven methods: LLMs alone are not sufficiently reliable for direct deployment due to ambiguity and embodiment mismatch, while unguided stochastic optimization can be inefficient or unstable. By combining semantic priors with stochastic refinement, the proposed method enables efficient and adaptive cost design for complex locomotion tasks. Experimental results on quadrupedal locomotion over challenging terrain demonstrate that LLM-CEM initialization significantly improves performance, while highlighting the importance of structured priors in stabilizing the optimization process.
|
| |
| 17:00-18:00, Paper TH1700-2.3 | Add to My Program |
| An Underactuated Robot Gripper with Slider Dual-Parallelogram Linkages for Pinching and Scooping in Environmental Constraints |
|
| Hou, Yuxuan | Windermere Preparatory School |
| Guo, Chunyu | Southern University of Science and Technology |
| Zhang, Wenzeng | Tsinghua University |
Keywords: Robotics and Automation in Unstructured Environment, Robotics and Automation Applications, Mechanics and Mechanism Design
Abstract: Traditional linkage-based underactuated grippers suffer from inherent kinematic limitations, as their closed-loop topologies fail to fully exploit environmental constraints. Such structural defects result in poor environmental adaptability and unreliable grasping performance for ultra-thin objects on non-planar surfaces, which restricts their application in precision manipulation. To address these drawbacks, this paper proposes a novel slider dual-parallelogram underactuated gripper. Mounted in an inverted suspended configuration, the device utilizes a horizontal guide rail that guides the dual-parallelogram fingertips along a strict linear path. Benefiting from the passive environmental adaption characteristics of the designed linkage topology, the gripper can actively contact with supporting surfaces and passively adjust fingertip postures through environmental reaction forces. Equipped with only one actuator, the proposed gripper features a concise structural layout, low production cost, and lightweight characteristics. Kinematic and static analysis demonstrate that the presented mechanism overcomes the gripper instability and environmental adaption defects of existing underactuated grippers, providing a low-cost and stable solution for robotic grasping capabilities.
|
| |
| 17:00-18:00, Paper TH1700-2.4 | Add to My Program |
| An Underactuated Robot Gripper with Double-Slot Linkages for Pinching and Passive Scooping in Environmental Constraints |
|
| Cao, Chenning | Yucai High School Education Group in Shekou |
| Ye, Yuxuan | Yucai High School Education Group in Shekou |
| Li, Zeyu | Yucai High School Education Group in Shekou |
| Zhang, Wenzeng | Tsinghua University |
Keywords: Robotics and Automation in Unstructured Environment, Mechanics and Mechanism Design, Methodologies for Robotics and Automation
Abstract: Underactuated robotic grippers with adaptive grasping capabilities are of significant value for tabletop manipulation tasks. Existing grippers struggle to simultaneously adapt to inclined or irregular surfaces while maintaining stable scooping performance for thin objects and precise pinching accuracy, often relying on multi-motor actuation that increases system complexity. This paper presents a single-actuator underactuated robotic gripper (DSL Gripper) based on double-slot link mechanisms. The design achieves near-linear fingertip motion through coupled constraints between dual slots and linkage elements, while integrating spring-stopper mechanisms to enable surface-adaptive compliance. This allows the finger to perform three grasping modes—parallel pinching, passive scooping, and enveloping grasping—under single-motor actuation. Physical prototype experiments demonstrate that the DSL gripper exhibits excellent adaptability to inclined surfaces, particularly showing stable scooping capability and precise positioning for thin objects. With its simple structure and controllable cost, this design is suitable for daily service and industrial precision assembly applications, providing a novel technical solution for adaptive grasping in complex environments.
|
| |
| 17:00-18:00, Paper TH1700-2.5 | Add to My Program |
| DLIRO: Dual-Antenna RTK-Aided LiDAR-Inertial Odometry for Large-Scale Outdoor Localization |
|
| Sun, Zewen | National University of Singapore |
| Li, Dongen | National University of Singapore |
| Huang, Zefan | National University of Singapore |
| Wang, Wuqiang | Sichuan University |
| Lee, Christina Dao Wen | National University of Singapore |
| Liu, Junqi | College of Computer Science, Sichuan University, Chengdu 610065, China |
| Yuan, Chengran | National Universtiy of Singapore |
| Sun, Shuo | Singapore-MIT Alliance for Research and Technology (SMART) |
| Guo, Hongliang | Agency for Science Technology and Research |
| Tay, Francis | NUS |
| Ang Jr, Marcelo H | National University of Singapore |
Keywords: Localization & Tracking, Robotics and Automation Applications, Methodologies for Robotics and Automation
Abstract: High-accuracy outdoor SLAM often benefits from the fusion of RTK, LiDAR, and IMU, which combines global positioning capability with robust local geometric and inertial constraints. However, many existing methods mainly rely on single-antenna RTK or use GNSS primarily as position-level constraints, while the additional heading observability provided by dual-antenna RTK remains underexplored. To address this issue, we propose a tightly coupled GNSS–LiDAR–IMU SLAM framework that explicitly exploits dual-antenna RTK position and heading measurements. The proposed system first performs LiDAR–IMU fusion within an IESKF-based LIO module to obtain high-rate local state estimates, and then uses these estimates as the prediction in an ESKF, where dual-antenna RTK measurements are incorporated as observations to obtain posterior pose estimates. To better fuse low-rate RTK measurements with high-rate LIO outputs, LiDAR observations are temporally aligned according to RTK timestamps. A measurement quality-check mechanism is further introduced to reduce the adverse impact of degraded RTK signals. Meanwhile, an IMU-based latency compensation module propagates delayed RTK observations to the fusion time before the ESKF update, improving their temporal consistency with the high-rate LIO priors. In addition, we build a GNSS-LiDAR–IMU platform for comprehensive evaluation, motivated by the limited availability of public datasets with dual-antenna RTK. Experimental results demonstrate that introducing the RTK-based correction module improves localization accuracy and global trajectory consistency in large-scale scenarios.
|
| |