| | |
Last updated on August 14, 2026. This conference program is tentative and subject to change
Technical Program for Friday August 21, 2026
| |
| FrS1A |
Yi Hui Hall 4+5 |
| Award Session I |
Regular Sessions |
| Chair: Wang, Zhidong | Chiba Institute of Technology |
| |
| 10:45-10:57, Paper FrS1A.1 | |
| A Novel Audio-Supervised Cross-Modal Framework with Wearable EMG--IMU Sensing for Emotion-Aware Expressive Music Performance Analysis (I) |
|
| Hu, Wanxin | Université Évry Paris-Saclay |
| He, Xiaotong | Université Évry Paris-Saclay |
| Wang, Feilong | Université Évry Paris-Saclay |
| Schmirander, Yunus | Université Paris-Saclay |
| Su, Hang | Paris Saclay University |
| Dychus, Eric | Sandyc |
| Alfayad, Samer | Paris-Saclay Universit -Evry University |
| Couprie, Pierre | CHCSC, Université Evry Paris-Saclay |
Keywords: Sensor Networks
Abstract: Audio-based music emotion recognition describes the affective outcome of musical performance, but provides limited insight into the performer-side bodily process involved in expressive sound production. This paper proposes an audio-supervised cross-modal framework with wearable EMG–IMU sensing for emotion-aware expressive music performance analysis. Instead of directly annotating physiological or motion signals with subjective emotion labels, audio-derived emotion labels are used as weak supervision. A Music2Emotion-based audio-teacher model generates valence, arousal, and discrete emotion labels from piano performance audio, which are temporally aligned with synchronized EMG and IMU windows. Statistical muscular and motion features are then evaluated under EMG-only, IMU-only, and EMG–IMU fusion settings. Experiments on 29 piano pieces produced 34,011 window-level samples. The best internal result reached 91.82% accuracy and 91.77% macro F1-score, while held-out piece-level validation reached 89.30% accuracy. The results suggest that wearable EMG and IMU signals provide complementary bodily cues for predicting audio-derived affective categories.
|
| |
| 10:57-11:09, Paper FrS1A.2 | |
| Task Planning and Control Using Generative Adversarial Tri-Model (GAT) Method for Electrical Cable Installation |
|
| Lei, Hejun | The University of Hong Kong |
| Ma, Xin | The University of Hong Kong |
| Chen, Heping | The University of Hong Kong |
| Xi, Ning | The University of Hong Kong |
Keywords: Artificial Intelligence, Grasping and Manipulation, Industrial Robotics and Factory Automation
Abstract: Robotic manipulation of deformable linear objects (DLOs) has the potential to replace tedious manual shaping and installation tasks in industrial electrical assembly. However, applying pure neural networks (NNs) to DLO manipulation remains challenging due to data scarcity, limited transferability, and poor interpretability. Physics-based analytical models (AMs) serve as a powerful complement that helps bridge these gaps, yet existing hybrid approaches couple the two only in a unidirectional manner, leaving at least one of these issues unresolved. We propose the Generative Adversarial Tri-model (GAT) framework, which establishes a bidirectional coupling between an NN and an AM through two symmetric forward--inverse closed loops, turning their interaction into a positive-sum game that converges to mutual agreement. Experiments on a variety of electrical cables and target shapes demonstrate the accuracy (MSE of 81.8), efficiency (8 steps to converge on average), robustness, and transferability of GAT for cable manipulation.
|
| |
| 11:09-11:21, Paper FrS1A.3 | |
| Vision-Force ACT: Multimodal Action Chunking for Robotic Bimanual Peg-In-Hole Assembly |
|
| Chen, Mingqi | Shenzhen Technology University |
| Chen, Meijie | Shenzhen Technology University |
| Li, Qiang | Shenzhen Technology University |
| Ming, Zhong | Shenzhen University |
| Liu, Chao | LIRMM - French National Center for Scientific Research (CNRS) |
Keywords: Grasping and Manipulation, Deep Learning, Industrial Robotics and Factory Automation
Abstract: Robotic bimanual peg-in-hole assembly requires accurate geometric alignment and compliant manipulation. Vision-based imitation learning is widely used in such kind of bimanual manipulation, which can efficiently locate and approach the target. However, visual observations could be ambiguous when the peg, hole or contact surface is occluded by the manipulators. This paper presents Vision-Force ACT, a mutimodal extension of Action Chunking with Transformers (ACT) that incorporates bilateral wrist force/torque observations into both the conditional latent inference and the policy representation. The proposed method successfully outputs temporally coherent bimanual action chunks from synchronized images, joints state, and wrist F/T signals. A Mujoco-based simulated bimanual assembly platform equipped with 6-axis force/torque (F/T) sensor at both wrists is used in the experiment, and the proposed method is compared with the vision-only ACT baseline under identical training settings. The results show that Vision-Force ACT improves the task success rate from 40% to 68% with 50 demonstrations and from 60% to 78% with 200 demonstrations. The recorded F/T curves are also smoother than those of the vision-only policy and show closer agreement with the force demonstrations. These preliminary results support force sensing as an effective complement to visual observations for contact-rich bimanual imitation learning, while also motivating broader multi-task and real-robot validation.
|
| |
| 11:21-11:33, Paper FrS1A.4 | |
| SOPD-SocialNav: Selective On-Policy Distillation for Vision-Language Social Navigation |
|
| Zhang, Xinyu | Hokkaido University |
| Wang, Zishuo | Hokkaido University |
| Xiao, Ling | Hokkaido University |
Keywords: Human-Robot Interaction and Cooperation, Artificial Intelligence, Emerging Technologies and Applications
Abstract: Vision-language models have shown strong potential for social robot navigation by leveraging rich semantic understanding of complex environments and human behaviors. However, large scale VLMs are difficult to deploy on resource-constrained robotic platforms, while lightweight VLMs often lack sufficient social reasoning capability. To address this problem, we propose SOPD-SocialNav, a selective on-policy distillation (SOPD) method that transfers social navigation knowledge from a large teacher VLM to a lightweight student VLM. SOPD introduces an entropy-based token selection mechanism that uses teacher uncertainty to identify socially informative decision tokens, while suppressing gradients from low-entropy tokens corresponding to trivial navigation states. A temperature-controlled Jensen-Shannon divergence objective is then used to align the student and teacher distributions on the selected tokens. Experiments on the SNEI and MUSON benchmarks demonstrate that SOPD consistently outperforms supervised fine-tuning, off-policy distillation, and standard on-policy distillation baselines in action prediction, perception consistency, and reasoning consistency. Real-world deployment on a Scout Mini robot further shows that the distilled model can generate more socially appropriate navigation behaviors in conversational and queuing scenarios. These results suggest that SOPD is an effective strategy for building lightweight yet socially aware VLM-based navigation systems.
|
| |
| 11:33-11:45, Paper FrS1A.5 | |
| Differential Evolution-Based Airborne Converge Mission Scheduling for UAV Swarms in Large-Scale Forestry Inspection |
|
| Chen, Zhitao | Shanghai Jiao Tong University |
| Dong, Kangsheng | China Aerodynamics Research and Development Centre |
| Sun, Yinshuai | Shanghai Jiao Tong University |
| Li, Zhengxiong | Shanghai Jiao Tong University |
| Dong, Peng | Shanghai Jiao Tong University |
Keywords: Intelligent Control and Systems, Multi-Robot Systems, Agricultural Robotics
Abstract: Airborne carrier platforms can extend the operational range and deployment flexibility of multi-formation UAV swarms in large-scale forestry inspection by enabling a cyclic deployment-inspection-converge mode. In such missions, multiple UAV formations inspect different forest blocks for canopy mapping, tree inventory, pest and disease monitoring, fire-risk patrol, and communication relay. After inspection, the formations are spatially dispersed and must rendezvous with the carrier for sequential converge. With the increasing scale of UAV swarms and the growing demand for formation-cooperative missions, existing studies are insufficient to satisfy the flexibility and integrity requirements in converge mission forest farm inspection missions. To address these challenges, this paper proposes a converge scheduling framework considering the formation constrains. The elitist differential evolution algorithm is developed as the mission scheduling optimizer, in which a formation keeping reward enforces formation sequential converge over continuous time intervals, and a core priority reward ensures preferential scheduling of high value UAVs. In addition, a velocity gradient strategy was adopted to optimize the solution space. Together, these mechanisms enable flexible mission schedule while preserving the converge order at the formation level. Extensive simulations demonstrate that the framework efficiently generates feasible schedules across diverse formation distribution conditions.
|
| |
| 11:45-11:57, Paper FrS1A.6 | |
| ODD-SEC: Onboard Drone Detection with a Spinning Event Camera |
|
| Dai, Kuan | Hunan University |
| Zhang, Hongxin | School of Artificial Intelligence and Robotics, Hunan University |
| Zhong, Sheng | Hunan University |
| Zhou, Yi | Hunan University |
Keywords: Robot Vision and Computer Vision, Deep Learning, Artificial Intelligence
Abstract: The rapid proliferation of drones requires balancing innovation with regulation. To address security and privacy concerns, techniques for drone detection have attracted significant attention. Passive solutions, such as frame camera-based systems, offer versatility and energy efficiency under typical conditions but are fundamentally constrained by their operational principles in scenarios involving fast-moving targets or adverse illumination. Inspired by biological vision, event cameras asynchronously detect per-pixel brightness changes, offering high dynamic range and microsecond-level responsiveness that make them uniquely suited for drone detection in conditions beyond the reach of conventional frame-based cameras. However, the design of most existing event-based solutions assumes a static camera, greatly limiting their applicability to moving carriers—such as quadrupedal robots or unmanned ground vehicles—during field operations. In this paper, we introduce a real-time drone detection system designed for deployment on moving carriers. The system utilizes a spinning event-based camera, providing a 360° horizontal field of view and enabling bearing estimation of detected drones. A key contribution is a novel image-like event representation that operates without motion compensation, coupled with a lightweight neural network architecture for efficient spatiotemporal learning. Implemented on an onboard Jetson Orin NX, the system can operate in real time. Outdoor experimental re
|
| |
| FrS1B |
Yi Hui Hall 6 |
| Manipulation and Grasping |
Regular Sessions |
| Chair: Chen, Fei | T-Stone Robotics Institute, the Chinese University of Hong Kong |
| |
| 10:45-10:57, Paper FrS1B.1 | |
| HAS-Bench: A Benchmark Suite for Vision-Denied Tactile Object Search |
|
| Fu, Xiangyu | Technical University of Munich |
| Armleder, Simon | Technische Universität München |
| Cheng, Gordon | Technical University of Munich |
Keywords: Sensing, Haptic System, Grasping and Manipulation
Abstract: Tactile perception enables robots to identify and retrieve objects when vision is unavailable, yet standardized benchmarks for vision-denied tactile learning remain limited. We present HAS-Bench, a benchmark suite built on the released Hide-and-Seek dataset. HAS-Bench defines three tasks that follow a tactile search-and-retrieval interaction: i) object– weight classification, ii) target retrieval, and iii) early tactile recognition. All tasks share synchronized observations from a bimanual robot, including tactile point clouds, raw skin signals, virtual skin wrenches, proprioceptive poses, and wrist force/torque measurements, and are evaluated with fixed splits, metrics, and protocols. We report reference results using a multimodal fusion baseline and show that HAS-Bench exposes three key challenges: 1) weight-aware joint recognition, 2) target verification under same-object/different-weight distractors, and 3) calibration of early decisions under limited observation time. HAS-Bench further supports evaluation across modality configurations, enabling comparisons between skin-derived tactile representations and configurations that additionally include proprioceptive pose and wrist force/torque sensing. The dataset, evaluation code, and pretrained baseline checkpoints are publicly available.
|
| |
| 10:57-11:09, Paper FrS1B.2 | |
| Flexible Robotic Order Picking System with Zero-Shot Vision for Customized Manufacturing |
|
| Liu, Yichang | The Hong Kong University of Science and Technology (Guangzhou) |
| Xie, Zhenxing | HKUST(GZ) |
| Hu, Ruolin | The Hong Kong University of Science and Technology(Guangzhou) |
| Qin, Zixiang | The Hong Kong University of Science and Technology ( Guangzhou ) |
| Guo, Daqiang | The Hong Kong University of Science and Technology (Guangzhou) |
Keywords: Industrial Robotics and Factory Automation, Artificial Intelligence, Grasping and Manipulation
Abstract: Flexible order picking and kitting are critical for customized manufacturing, yet current robotic systems often struggle to adapt to changing orders, inventory variations, and quality requirements without engineering intervention. The gap lies between shop-floor knowledge and robot-executable plans: operators understand real-time conditions, whereas conventional automation remains tied to fixed programs. This paper presents OZPS, an order-aware zero-shot picking system that converts natural-language production orders into executable component-level picking tasks through closed-loop integration of order reasoning, open-vocabulary scene grounding, and deterministic manipulation. Instead of relying on end-to-end action generation or task-specific visual retraining, OZPS combines configurable production rules with RGB-D/VLM-based grounding to enable adaptive picking under dynamic conditions. The system was validated in 150 physical trials across normal fulfillment, shortage-driven substitution, and defective-component rejection, achieving task success rates of 86%, 76%, and 82%, respectively, together with 98% instruction-parsing accuracy, a 6% counting error, and a 19.8 s average setup time. The results show that modular integration of semantic order interpretation, zero-shot perception, and deterministic robot execution can reduce reprogramming effort and improve the adaptability of robotic picking systems in customized manufacturing.
|
| |
| 11:09-11:21, Paper FrS1B.3 | |
| PJN-Based Finite-Width Path Planning and End-Effector Posture Smoothing for Robotic Fibre Placement on Doubly Curved Surfaces |
|
| Jin, Ze | Harbin University of Science and Technology |
| Liu, Meijun | Harbin University of Science and Technology |
| Cao, Shanyu | Harbin University of Science and Technology |
|
|
| |
| 11:21-11:33, Paper FrS1B.4 | |
| VLM-Guided Robotic Bin Picking: A Cognitive-Kinematic Architecture for Eye-To-Hand Manipulation in Unstructured Environments |
|
| Liu, Junqi | The Hong Kong University of Science and Technology (Guangzhou) |
| Guo, Daqiang | The Hong Kong University of Science and Technology (Guangzhou) |
Keywords: Industrial Robotics and Factory Automation, Artificial Intelligence, Grasping and Manipulation
Abstract: Robotic bin-picking in unstructured environments presents significant challenges, including severe occlusions, object entanglement, and complex collision dynamics. Conventional geometric approaches and purely data-driven methods often lack robust semantic reasoning and deterministic safety guarantees. To address these limitations, this paper proposes a Vision-Language Model (VLM) guided hierarchical framework that effectively bridges the cognitive-perceptual gap and the kinematic-execution gap. Leveraging an eye-to-hand RGB-D sensor setup, our Cognitive-Kinematic Architecture explicitly decouples high-level semantic reasoning from low-level deterministic control. Visual observations are processed through this proposed bottleneck to enable robust, open-vocabulary perception while mitigating the risk of spatial hallucinations. The refined spatial information is subsequently utilized by a pose estimation network to generate 6-DoF grasp poses, which are executed via the MoveIt2 motion planning framework with collision-free guarantees. Extensive experimental results demonstrate significant improvements in both grasp success rate (GSR) and success task achievement (STA) metrics. Finally, we critically analyze the inherent limitations of our approach and outline future research directions, including integration with Digital Twin systems and haptic feedback mechanisms.
|
| |
| 11:33-11:45, Paper FrS1B.5 | |
| A Physical–Optical Coupled Framework for Target-Valid Autofocus in Robotic Cell Micromanipulation |
|
| Chen, Dingfu | Zhejiang University |
| Long, Yun | Zhejiang University |
| Chew, Ting Gang | Zhejiang University-University of Edinburgh (ZJU-UoE) Institute |
| Yang, Liangjing | Zhejiang University |
| Tan, U-Xuan | Singapore University of Techonlogy and Design |
Keywords: Micro/Nano Robotics, Robot Vision and Computer Vision, Intelligent Control
Abstract: Robotic cell micromanipulation requires autofocus under a narrow microscopic view that is sensitive to target drift. Sharpness-based methods reduce autofocus to axial maximization of a focus measure, but the score is meaningful only when it is computed on target-relevant content. Once disturbance leaves the intended cell only partially or unstably represented within the focus-evaluation region of interest (ROI), sharpness maximization no longer defines target-valid autofocus. This paper presents a physical–optical coupled framework for target-valid autofocus and instantiates it in a microscopic robotic simulation environment. The framework couples robot motion, target disturbance, optical defocus, image formation, and camera observation for repeatable closed-loop evaluation. OS-Autofocus (observation-stabilized autofocus) learns image based lateral–axial motion commands that stabilize the target observation while recovering focus. A target-valid VWD-Tenengrad reward credits sharpness improvement only while the intended cell remains inside the ROI, preventing off-target views from driving the policy. Ablations identify PPO as the best policy and a VWD-Tenengrad reward as the best focus measure, and static and moving-cell experiments show that OS-Autofocus improves target-valid autofocus over depth-from focus and visual-servoing-assisted baselines.
|
| |
| 11:45-11:57, Paper FrS1B.6 | |
| Mechanism Design and Performance Evaluation of an Electromagnetic-Driven Flexible Microgripper for Small Target |
|
| Lyu, Zekui | Southeast University |
| Zhang, Tairui | Southeast University |
| Gu, Wenwen | Nanjing Special Equipment Safety Supervision and Inspection Research Institute |
Keywords: Grasping and Manipulation, Smart Structures, Materials, Actuators
Abstract: Microgrippers are essential end-effector devices for micromanipulation and microassembly. This paper presents a microgripper composed of a voice-coil motor and a compliant mechanism, which offers excellent characteristics such as a large stroke, high precision, compatibility with force sensing, and strong adaptability. The mechanical design, preliminary analysis, finite element evaluation, and experimental testing of the proposed microgripper are completed and presented. The input displacement of the gripper is limited to a specific range to prevent overshoot impacts from affecting the target. The designed microgripper supports both normally open and normally closed operating modes to accommodate microtargets of different sizes. Ultimately, experimental performance tests of the microgripper met the expected requirements, and it has successfully grasped a variety of microtargets.
|
| |
| FrS2A |
Yi Hui Hall 4+5 |
| Award Session II |
Regular Sessions |
| Chair: Tao, Yong | Beijing University of Aeronautics and Astronautics |
| |
| 14:10-14:22, Paper FrS2A.1 | |
| Wearable Knee Assistance Robot Using Artificial Muscles |
|
| Zhou, Changqiu | The University of Hong Kong |
| Zhao, Yafei | The University of Hong Kong |
| Ling, Zi-qin | The University of Hong Kong |
| Yuan, Wenbo | The University of Hong Kong |
| Zhang, Qingqing | The University of Hong Kong |
| Ma, Xin | The University of Hong Kong |
| Wang, Haiyang | The University of HongKong |
| Chen, Jiangcheng | The University of Hong Kong |
| Lou, Vivian Weiqun | The University of Hong Kong |
| Xi, Ning | The University of Hong Kong |
Keywords: Rehabilitation and Assistive Robotics, Intelligent Control and Systems, Soft Robotics
Abstract: Rigid knee exoskeletons suffer from three fundamental incompatibilities: mechanical, kinematic, and control. This paper presents a wearable knee assistance robot driven by twisted string actuators (TSA) that systematically addresses all three. For mechanical incompatibility, a TSA impregnated with shear-thickening gel (STG) serves as an artificial muscle, emulating the variable-stiffness and force-generation properties of biological muscle. For kinematic incompatibility, a compliant exosuit aligns the actuation path with muscle fiber orientation, removing the need for a rigid joint constraint. For control incompatibility, sEMG-phase-onset detection with real-time tracking replaces IMU-based triggering, anticipating exertion intention before joint motion and rejecting false triggers from signal interference. Evaluated with 18 older adults (mean age 70.2 ± 4.9 years) across a three-level framework: at the physiological level, activation of vastus lateralis, rectus femoris, and vastus medialis decreased by 14% (SD = 4.6%), 11% (SD = 2.5%), and 17% (SD = 5.7%), respectively, during isometric knee extension; at the functional level, VL median frequency decline slowed under assistance, indicating delayed muscle fatigue; at the behavioral level, seven participants regained sit-to-stand capability from a height previously impossible without assistance, while all others completed STS from any height. These results demonstrate that simultaneously resolving mechanical, kinematic, and co
|
| |
| 14:22-14:34, Paper FrS2A.2 | |
| Development of a High-Stiffness Tendon-Driven Robotic Finger with Synchronous Tendon Routing |
|
| Yuan, Quan | ShanghaiTech University |
| Du, Zhenting | King's College London |
| Cao, Daqian | ShanghaiTech University |
| Bai, Weibang | ShanghaiTech University |
Keywords: Robot Design, Biologically Inspired Robotics, Grasping and Manipulation
Abstract: Multi-finger robotic hands typically require many actuators to control their joints independently, which increases weight, volume, wiring complexity, and cost, and makes it hard to combine high load-bearing capacity with adaptive compliance in a compact form. This paper presents an under-actuated tendon-driven robotic finger (UTRF) featuring a synchronous tendon-routing mechanism that mechanically couples all three joints at fixed angular ratios, allowing the entire finger to be actuated by a single tendon and a single actuator while ensuring coordinated joint motion with predetermined angular ratios. Meanwhile, an antagonistic tendon pair enables bidirectional actuation, whereas the spring-free synchronous tendon-routing mechanism, together with tendon elasticity, provides the finger with predictable stiffness and substantial load-bearing capability while preserving the compliance required for adaptive grasping. Kinematic and static models incorporating tendon elasticity are developed. A single-finger prototype was fabricated and tested under static loading, showing a mean deflection-prediction error of 0.551 mm (0.322% of the finger length) and a measured tip stiffness of bm{ 1.2 times 10^{3}} N/m under a 3 kg load. Furthermore, integrating it into a five-finger hand (UTRF-RoboHand) demonstrates the reliable grasping of objects with diverse shapes, weights, and fragility.
|
| |
| 14:34-14:46, Paper FrS2A.3 | |
| A Bio-Inspired Dexterous Finger with Lateral Flexibility and Variable Stiffness |
|
| Ye, Jiayin | Shanghai Jiao Tong University |
| Zhou, Guozhen | Shanghai Jiao Tong University |
| He, Haowen | Shanghai Jiao Tong University |
| Lei, Xuyang | Shanghai Jiao Tong University |
| Chen, Feifei | Shanghai Jiao Tong University |
Keywords: Biologically Inspired Robotics, Robot Design, Grasping and Manipulation
Abstract: Human hands exhibit remarkable manipulation performance owing to their unique musculoskeletal and ligamentous biomechanics. However, it remains challenging for current robotic dexterous hands to have both high control accuracy and human-like compliance and stiffness characteristics. This work develops a bio-inspired anthropomorphic dexterous finger. Antagonistic tendons and a variable stiffness module are integrated to implement active joint stiffness control, and elastic structures are used to replicate the passive lateral compliance of ligaments. Theoretical models are predicted to describe the lateral stiffness and stiffness control capability of the finger joint. Well in line with the experiment results, the finger realizes a trade-off between precise motion control and passive compliance, delivering a 24-fold stiffness adjustment range, demonstrating great potential for interactive manipulation tasks.
|
| |
| 14:46-14:58, Paper FrS2A.4 | |
| High-Rate Full-Hand Tactile Sensing for Sim-To-Real Recognition and Grasping |
|
| Qiu, Silin | Shenzhen Technology University |
| Zhong, ZhengYang | Shenzhen Technology University |
| Wang, Yanyi | Shenzhen Technology University |
| Zhenyuan, Zhang | Shenzhen Technology University |
| Enrui, Zhang | Shenzhen Technology University |
| Lyu, Jingke | School of Artificial Intelligence, Shenzhen Technology Univer-Sity |
| Li, Qiang | Shenzhen Technology University |
Keywords: Grasping and Manipulation, Sensing, Haptic System, Humanoid Robots
Abstract: Full-hand tactile sensing provides contact information that vision can hardly capture after a dexterous hand encloses an object. This is important for stable grasping and shape recognition under occlusion. This paper presents a full-hand tactile perception framework for sim-to-real dexterous grasping. The platform uses a four-finger dexterous hand with 16 degrees of freedom. It is equipped with an in-house, low-cost piezoresistive tactile sensing system that is densely integrated across the entire hand. It also provides stable high-rate tactile signals at 300 Hz. In simulation, MuJoCo is used to match the positions of different tactile sensing taxels. Ray-normal projection and local spacing correction are used to align the tactile layout on the hand surface. The simulated tactile responses are then mapped to the real voltage domain through force-voltage calibration. A multi-patch tactile classifier is trained and evaluated with simulated and real data. In experiments, the model achieves 73.33% accuracy on a real dataset with 150 object-grasp samples. After fine-tuning with another 150 real samples, the accuracy reaches 92.67% on the test set. These results show that high-rate full-hand piezoresistive sensing, geometric taxel alignment, and calibrated tactile simulation can support practical grasp-based object shape recognition.
|
| |
| FrS2B |
Yi Hui Hall 6 |
| Human-Robot Interaction |
Regular Sessions |
| Chair: Chen, Heping | The University of Hong Kong |
| |
| 14:10-14:22, Paper FrS2B.1 | |
| SMILE: No-Code Mission Authoring for Field Service Robots by Integrating Google Blockly with Behavior Trees |
|
| Nangsue, Norawit | King Mongkut's University of Technology Thonburi |
| Rommueang, Siwakon | King Mongkut's University of Technology Thonburi Institute of Field Robotics |
| Malimai, Polakrit | King Mongkut's University of Technology Thonburi |
| Julkananusart, Arthit | Institute of Field Robotics (FIBO) King Mongkut’s University of Technology Thonburi |
| Visarnkuna, Wuttichai | King Mongkut's University of Technology Thonburi |
Keywords: Intelligent Control and Systems, ROS, Software System for Robotics Application, Cognitive Robotics
Abstract: This paper presents SMILE (Service-robot Mission Interface with bLock-based Editing), a no-code mission-authoring system for an Autonomous Indoor Service Robot (AISR) developed at FIBO, KMUTT. SMILE combines Google Blockly visual programming with Behavior Trees (BT), allowing non-programmer operators to compose, edit, and deploy robot missions without writing code or XML. The architecture has three layers and separates authoring from execution: one Blockly workspace produces a .bxml document for round-trip visual editing and a standard BehaviorTree.CPP 4.0 .xml document for runtime execution. An execution engine based on pluginlib, bt_runner, ticks the active tree at 10 Hz and loads leaf nodes declaratively, supported by a reusable SubTree library and an SQLite mission store. In a formative usability study (n=5), non-programmer staff were able to author a basic patrol mission in about four minutes after two hours of training. This replaced an earlier workflow in which every change needed an off-site engineering request that took hours or even days. The system has been running continuously in a large Bangkok shopping mall since 2023, and on-site operators author and maintain the deployed missions by themselves. We report the architecture, a generator-validity stress test, and the formative results, and we also discuss the limitations of the current evaluation.
|
| |
| 14:22-14:34, Paper FrS2B.2 | |
| GOVBOT: A Conceptual Architecture for Trustworthy Embodied AI Service Robots in Government Knowledge Assistance |
|
| Alsofyani, Roaa | Saudi Data & AI Authority |
| K. Alshammari, Reem | King Abdulaziz City for Science and Technology |
| AlJohani, Tahani | Education and Training Evaluation Commission (ETEC) |
Keywords: Human-Robot Interaction and Cooperation, Humanoid Robots
Abstract: As governments expand digital and physical service delivery, citizens continue to face challenges accessing government services due to digital literacy limitations, accessibility needs, language barriers, and difficulties navigating complex service processes. This paper proposes GOVBOT, a conceptual architecture for an embodied AI-powered government service robot designed for citizen-facing environments, including service centers, public offices, and reception areas. Through multimodal human–robot interaction, the proposed architecture is designed to provide trusted government information, service guidance, and procedural assistance for citizens and public-sector employees. By integrating Generative AI, Retrieval-Augmented Generation (RAG), and the proposed Trustworthy AI Governance (TAG) Framework, GOVBOT embeds governance-by-design principles throughout the AI lifecycle, aiming to enable transparent, explainable, accountable, and policy-compliant government service delivery.
|
| |
| 14:34-14:46, Paper FrS2B.3 | |
| Public Discourse on Robot Performances at China's Spring Festival Gala: A Multi-Platform Social Media Analysis |
|
| Tu, Yangjun | Hunan University |
| Wu, Tongyang | Hunan University |
| Wang, Likang | Hunan University |
| Chen, Shaoxuan | Hunan Normal University in Changsha, China |
| Tan, Xiaoyu | Hunan University |
| Yang, Zhi | Hunan University |
Keywords: Human-Robot Interaction and Cooperation, Emerging Technologies and Applications
Abstract: As robots increasingly feature at large-scale events, public reactions extend beyond entertainment. Using the 2025 and 2026 CCTV Spring Festival Gala robot performances (audiences over 700 million) as a natural setting to track public discourse, we analyze 16,392 texts from five social media platforms, applying a multi-label scheme of 15 themes and 7 narrative frames to 2,333 non-danmaku texts and using 14,059 timestamped danmaku to examine real-time dynamics, alongside analyses of evaluative tension, like-weighting, and institutional/grassroots sourcing. We find that: (1) reactions diverge by platform, forming "divisions of discourse" (e.g., Zhihu stresses technological industrialization); (2) attitudes shifted from a sole "appreciation of the technological spectacle" to a "coexistence of pride and anxiety," with evaluative tension surging from 5.8% to 36.0% (Z = 18.49, p < 0.001); (3) occupation-related automation anxiety showed a persistent yet malleable pattern — Job Displacement was the only frame showing no significant year-over-year change (p = 0.54), yet anxiety fell short-term after broadcast (13.9% → 7.6%); (4) focal issues show a dual mismatch—anxiety is grassroots-dominated but institutionally neglected (agenda-setting mismatch), while the most-discussed topics are not the most-liked (engagement mismatch). These findings advance research on ambivalent technology attitudes and their dynamics, and inform the robot industry's socialization and national policy.
|
| |
| 14:46-14:58, Paper FrS2B.4 | |
| AFBResNet-TCNN-Based Robust SSVEP Brain-Controlled UAV System for Complex Operational Environments |
|
| Wu, Jiaxuan | Shenyang Ligong University |
| Dong, Tianyu | Shenyang Ligong University |
| Wang, Wenxue | Shenyang Institute of Automation, Chinese Academy of Sciences |
| Ma, Shuang | Shenyang Ligong University |
Keywords: Brain-Machine Interface, Human-Robot Interaction and Cooperation
Abstract: Abstract—Managing single-operator unmanned aerial vehicles (UAVs) in complex operational settings is challenging because of distraction, high engagement and heavy interaction burden. Existing steady-state visual evoked potential (SSVEP) decoding methods still suffer from recognition drifts and false activations under low signal-to-noise ratio, inter-individual variability and asynchronous control scenarios. To address these issues, this work proposes an adaptive filter bank residual temporal convolutional neural network (AFBResNet-TCNN) for brain-controlled UAV decoding, integrating multi-frequency and multi-phase stimuli and a maximum a posteriori probability threshold for idle-state detection. Experiments with 15 participants achieved 93.1% classification accuracy, a 64.5 bit/min information transfer rate (ITR), 95.4% idle-state detection accuracy, 92.5% online correct execution rate and 0.884 s post-decision execution latency. Compared with the evaluated FBTCNN baseline, AFBResNet-TCNN achieved higher cross-participant classification accuracy and ITR, while the idle-state and online experiments supported the safety and feasibility of low-effort UAV interaction.
|
| |
| FrS3A |
Yi Hui Hall 4+5 |
| Humanoid Locomotion |
Regular Sessions |
| Chair: Chen, Fei | T-Stone Robotics Institute, the Chinese University of Hong Kong |
| |
| 15:00-15:12, Paper FrS3A.1 | |
| Learning to Run a Half-Marathon: Human-Like Humanoid Locomotion Via Adversarial Motion Priors |
|
| Armleder, Simon | Technische Universität München |
| Fu, Xiangyu | Technical University of Munich |
| Guadarrama-Olvera, J. Rogelio | Technical University of Munich |
| Cheng, Gordon | Technical University of Munich |
Keywords: Humanoid Robots, Machine Learning
Abstract: We present a learning-based running controller that carried a full-sized humanoid robot through the 21 km course of the 2026 Beijing E-Town Humanoid Robot Half-Marathon, cruising at approximately 2.1 m/s and finishing in 3 h 05 min. Rather than hand-engineering the gait, we learn it from human motion capture using adversarial motion priors: a style reward learned from retargeted human clips complements a simple velocity-tracking objective, with commands sampled from a kinematic hull of the same data. Because the gait is specified by data rather than by robot-specific reward design, the same pipeline can carry humanoids of different sizes and morphologies from motion capture to real-world outdoor deployment without per-platform gait engineering. A walk-run transition and a flight phase emerge from the motion prior alone; the policy tracks velocity commands up to 4.5 m/s in simulation and 3.2 m/s on a physical robot. Across the deployed speed range, we characterize mechanical power, cost of transport (0.5 system-level during the race), and motor temperatures, showing that sustained running speed on air-cooled platforms is bounded by heat dissipation, not by the locomotion policy.
|
| |
| 15:12-15:24, Paper FrS3A.2 | |
| Locomotion Control of a Full-Size Humanoid Robot by Multi-Terrain Expert-Preserving Fine-Tuning Algorithm |
|
| Tian, Haoyu | Zhejiang University |
| Lei, Hua | Zhejiang University |
| Lu, GuoDong | Zhejiang University |
| Gan, Chunbiao | Zhejiang University |
Keywords: Humanoid Robots, Deep Learning
Abstract: Full-size humanoid robots have greater limb inertia and broader mass distribution, making stable multi-terrain control significantly more challenging. Moreover, conventional single-terrain pre-training policies struggle to avoid catastrophic forgetting. To improve the walking balance of the full-size humanoid robot on flat ground and multi-grade slope terrains, we propose a Multi-Expert Preserving Fine-tuning (MEPF) algorithm. First, walking sequences extracted from the Motion Capture Database are retargeted to the joint space; Adversarial Motion Priors (AMP) pre-trains a flat expert and a slope expert separately, endowing policies with human-like gait priors. Second, during fine-tuning, Terrain-conditioned Multi-expert Regularization (T-MAR) uses precise binary terrain labels from the simulator, eliminating cross-terrain knowledge interference and achieving accurate multi-terrain behavior preservation. Finally, an adaptive weight scheduling strategy based on the dual-expert mean squared error (MSE) gap dynamically allocates the gradient budget between the MAR regularization term and the Proximal Policy Optimization (PPO) objective, achieving parameter-free adaptive balance between expert behavior preservation and task performance. Experiments show that MEPF maintains stable walking on flat ground and slopes up to 9°, outperforming AMP mixed training, data aggregation, and policy distillation in both expert action preservation and terrain traversal speed.
|
| |
| 15:24-15:36, Paper FrS3A.3 | |
| Persistent Occlusion Disproportionately Degrades Feedforward Height-Map-Guided Humanoid Locomotion |
|
| Cheng, Yifeng | Zhejiang University |
| Huang, Weidi | Zhejiang University |
| Lu, GuoDong | Zhejiang University |
| Gan, Chunbiao | Zhejiang University |
Keywords: Humanoid Robots, Robot Vision and Computer Vision, Machine Learning
Abstract: Height-map-guided locomotion policies are widely adopted for humanoid robots traversing unstructured terrain, yet their robustness under sensor occlusion remains underexplored. Existing studies consider stochastic and temporally extended sensor dropouts, but do not isolate occlusion persistence under matched marginal corruption. Here, we trained a height-map-guided locomotion policy for the T1-Pro humanoid and constructed a structured four-factor perturbation framework spanning occlusion persistence, occlusion ratio, occlusion spatial distribution, and terrain type. Under matched occlusion ratios and mask-generation rules, episode-fixed persistent occlusion reduced survival by up to 37.2 percentage points relative to per-frame occlusion. Under our fixed-area truncated-window mask generator, greater spatial aggregation increased the observed persistence gap by approximately sixfold, whereas at alpha = 0.08 the gap varied by only 0.031 across six terrain types. These results identify temporal persistence as an independent and practically consequential robustness dimension for feedforward policies without explicit height-map memory.
|
| |
| 15:36-15:48, Paper FrS3A.4 | |
| Model Predictive Control for Bipedal Robot Walking Based on Foot-Ground Contact Awareness and State Estimation |
|
| Lu, Mengyue | Zhejiang University |
| Chen, Yongming | Zhejiang University |
| Gan, Chunbiao | Zhejiang University |
Keywords: Humanoid Robots
Abstract: Simplified-model errors, inertial measurement unit (IMU) drift, and discrete foot-ground contact switching jointly degrade feedback control of bipedal walking. This paper develops a proprioceptive closed-loop framework that integrates contact-aware state estimation, model predictive control (MPC), and executable compensation for a position-controlled robot. IMU, joint-encoder, and joint-current measurements are fused to compute contact probability. The contact result selects and weights leg-kinematic pseudo-measurements in an invariant extended Kalman filter (IEKF), which estimates base orientation, position, and velocity without resetting the state or covariance at support transitions. Using the estimated base and support states, a linear inverted pendulum model (LIPM) generates center-of-mass (CoM), zero-moment-point (ZMP), and swing-foot references, while MPC performs receding-horizon optimization under support constraints. A phase-dependent, bounded hierarchical mapping then converts the two-dimensional ZMP feedback into corrections for foot placement, pelvis attitude, swing-foot trajectory, and stance-leg attitude, providing a clear interface to inverse kinematics and joint-position control. Simulation and hardware experiments show lower base-velocity estimation error and improved closed-loop performance during the tested flat-ground acceleration, steady-speed, deceleration, and hardware-walking cases.
|
| |
| 15:48-16:00, Paper FrS3A.5 | |
| Plantar Pressure Guided Contact Aware Motion Tracking for Human-To-Humanoid Imitation |
|
| Yue, Boyu | Beihang University |
| Tao, Yong | Beijing University of Aeronautics and Astronautics |
| Ren, Fan | Nankai University |
Keywords: Humanoid Robots, Intelligent Control and Systems, Machine Learning
Abstract: Humanoid motion tracking has made promising human-to-humanoid imitation possible, yet kinematic references alone under-constrain foot-ground interaction. In dynamically demanding motions, a policy can match the desired body trajectory while relying on unintended contacts for stability, which makes the robot’s movements look very unlike. We propose a contact-aware humanoid motion tracking framework that incorporates measured plantar pressure into both retargeting and RL-based control. From synchronized human motion and pressure data, we extract per-foot contact labels, vertical ground reaction forces, and discretized center-of-pressure regions. Based on these signals, we first refine IK-based retargeted trajectories to enforce contact consistency and reduce reference artifacts such as foot sliding, floating, and penetration. We then introduce a progressive contact-aware training scheme with support-mode-conditioned rewards, and coverage-preserving adaptive sampling, enabling the policy to acquire stable whole-body tracking before learning contact timing, load transfer, and pressure migration. In Unitree G1 simulation, our method achieves kinematic tracking accuracy comparable to the baseline while substantially reducing contact, force, and CoP errors. These results indicate that measured foot–ground interaction provides critical physical supervision for resolving under-constrained humanoid tracking and improving the realism and physical consistency of human-to-humanoid motion
|
| |
| FrS3B |
Yi Hui Hall 6 |
| Navigation and Planning |
Regular Sessions |
| Chair: Guo, Di | Beijing University of Posts and Telecommunications |
| |
| 15:00-15:12, Paper FrS3B.1 | |
| Learning Semantic Navigation Primitives for Human Habitats Using a Recurrent State-Space World Model |
|
| Elksnis, Arturs | University of Bristol, University of the West of England, Bristol Robotics Laboratory |
| Ziqi, Luo | University of Bristol and the University of the West of England |
| Wang, Ning | Sheffield Hallam University |
| Pearson, Martin | Bristol Robotics Laboratory |
Keywords: SLAM and Navigation, Home and Personal Robot Systems, Robot Vision and Computer Vision
Abstract: Dividing complex tasks into sequences of simple primitives is a well established strategy in general problem solving. We apply this idea to Semantic Navigation (SN) tasks and propose Semantic Navigation Primitives (SNPs) that we learn and refine using the DreamerV3 World-Model algorithm applied in human domestic habitats. In this work, we train two SNPs, "move to the middle of the room" and "go through a door", which we combine to solve the higher order task of navigating from room to room to explore generic habitats. As data privacy is a key constraint for robots operating in domestic environments, we deploy our models using a single Nvidia Jetson Orin AGX embedded platform with no access to cloud-based resources. We also train our SNPs using modest commercial compute resources (RTX 3090) demonstrating the resource friendly nature of the approach to replicate and develop. We demonstrate a peak performance of 94% for our individual SNPs trained using a relatively small number of training steps (2.7M - 3.5M) highlighting the potential for this approach as the basis for a more complete embedded semantic navigation pipeline for domestic robots.
|
| |
| 15:12-15:24, Paper FrS3B.2 | |
| FlyAda: Belief-State Adaptation for Diffusion Policies under Partial Observation |
|
| Lyu, Lingxue | University of Pennsylvania |
Keywords: Intelligent Transportation Systems, Path and Motion Planning, Intelligent Control and Systems
Abstract: We study when a small online-updated latent helps an action-chunk diffusion policy on a UAV goal-reaching task under mass/drag/wind/delay perturbations. With the full 12-dim state visible, vanilla diffusion already handles the conventional sweep well, including extrapolation (4× mass, 10× drag, 8-step delay) — likely because receding-horizon re-planning at K=4 lets the next action chunk correct most dynamics error. An adaptation latent adds nothing here, and sometimes hurts at extreme extrapolation. The gap reappears under partial observation: zeroing velocity in the policy's input collapses success on the same sweep (0.3% mean), and a 3-frame-stack baseline does almost as badly (1.1%). FlyAda closes this gap by adding a small observer head trained with an auxiliary velocity-decoding loss — supervised with true velocity, a privileged training-time signal never observed at deployment — whose latent is updated online via an EMA. FlyAda reaches near-full success within the training perturbation range, degrading only at wind well beyond it (0.95 at 2 m/s, 0 at 3 m/s). A same-weights ablation isolating the test-time update rule shows continuous online updating is what matters; a linear probe recovers the hidden velocity from the latent at R²=0.992. The same gap and fix reproduce when we re-evaluate on a 6-DoF MuJoCo quadrotor.
|
| |
| 15:24-15:36, Paper FrS3B.3 | |
| Robust LiDAR-Visual-Inertial SLAM with Hierarchical Cross-Modal Matching for Quadruped Inspection Robots |
|
| Li, Kai | Southeast University |
| Wang, Zhongyuan | China Unicom Digital Technology Co., Ltd |
Keywords: Artificial Intelligence, SLAM and Navigation
Abstract: Quadruped inspection robots are increasingly used in industrial and infrastructure environments because of their strong mobility on stairs, slopes, narrow corridors, and uneven terrain. However, their locomotion usually introduces severe body vibration, which causes image blur, LiDAR scan distortion, and unstable odometry estimation. These problems become more critical in long-distance inspection scenarios with weak texture, repetitive structures, illumination changes, and LiDAR-degenerated geometry. This paper presents a robust LiDAR-visual-inertial SLAM framework for quadruped inspection robot mapping. The proposed method tightly fuses LiDAR, camera, and IMU measurements through an error-state iterated Kalman filter. To improve visual tracking under motion blur, an adaptive graph reasoning based visual front-end is adopted for reliable optical-flow-guided feature association. In addition, a hierarchical cross-modal matching strategy is designed for large-scale loop detection. Visual place recognition is first used to retrieve candidate historical regions, and LiDAR geometric registration is then performed for accurate verification. Finally, loop constraints are incorporated into an incremental factor graph to obtain a globally consistent 3D point cloud map.
|
| |
| 15:36-15:48, Paper FrS3B.4 | |
| HRG-PPO: Historical Residual Gating PPO for Privilege-Free Underwater Robot Navigation in Complex Currents |
|
| Meng, Linghan | Shenyang Institute of Automation, University of Chinese Academy of Sciences |
| Yao, Qingfeng | Shenyang Institute of Automation, Chinese Academy of Sciences |
| Zhang, Qifeng | Shenyang Institute of Automation, CAS |
Keywords: Machine Learning, Path and Motion Planning, Intelligent Control
Abstract: Underwater robot navigation in complex flow fields is challenging because strong current disturbances and partial observability limit reliable decision making. Existing methods often require global flow-field information, accurate models, or additional flow sensors, which are difficult to obtain in practical deployments. This paper investigates goal-directed navigation of a remotely operated vehicle (ROV) in a strong static double-gyre current field without privileged flow information and proposes Historical Residual Gating PPO (HRG-PPO). The key idea is to selectively use recent historical observations rather than treating all history as equally reliable. HRG-PPO encodes recent history as a residual correction to the current-observation representation and introduces a context-conditioned gate to filter outdated or less relevant historical information. The fused representation is used by the actor and critic for thruster-level control and value estimation. Simulation results based on the BlueROV2 Heavy model in MarineGym show that HRG-PPO improves the success rate and reduces episode length compared with standard PPO and a history-augmented PPO baseline using a fixed 10-step observation window.
|
| |
| 15:48-16:00, Paper FrS3B.5 | |
| Dynamic Payload Release and Transportation by a Quadrotor with Cable-Suspended Payload System Using Bilevel Optimization |
|
| Nai, Yuxuan | Zhejiang University |
| Jiang, Ding | University of Illinois Urbana-Champaign |
| Han, Ziyin | University of Illinois at Urbana Champaign |
| Yang, Chengyu | University of Illinois Urbana-Champaign |
| Cheng, Sheng | Beihang University |
| Hovakimyan, Naira | University of Illinois at Urbana-Champaign |
Keywords: Intelligent Transportation Systems, Path and Motion Planning, Emerging Technologies and Applications
Abstract: The quadrotor with a cable-suspended payload system offers an attractive alternative to ground transportation in hard-to-reach terrain. Yet they still lack the agility and efficiency seen in human-operated heli-logging, where pilots execute dynamic maneuvers to simultaneously release the payload and rapidly return for the next pickup. In contrast, most existing trajectory planning methods for such systems treat the release as the task’s endpoint, overlooking the cyclic nature of real transport missions. This paper introduces a bilevel optimization framework that jointly determines the payload’s release time and state along with the quadrotor's pre- and post-release trajectories. The problem is formulated as a two-phase hybrid system with continuity and spatial feasibility constraints. To manage the nonconvexity resulting from continuity constraints, we decompose the problem into a high-level optimization over release decisions and a low-level trajectory optimization expressible as quadratic programs. Simulation and hardware experiments demonstrate that the framework produces diverse, dynamically feasible release maneuvers from steady hover-and-release to aggressive trajectories resembling heli-logging behaviors.
|
| |
| FrS4A |
Yi Hui Hall 4+5 |
| Assistive Robots and Intelligent Control |
Regular Sessions |
| Chair: Chen, Heping | The University of Hong Kong |
| |
| 16:20-16:32, Paper FrS4A.1 | |
| Tool, Colleague, or Child? a Grounded Theory of Relational Divergence between Frontline Employees and Service Robots |
|
| Tu, Yangjun | Hunan University |
| Tan, Xiaoyu | Hunan University |
| Chen, Shaoxuan | Hunan Normal University in Changsha, China |
| Wu, Tongyang | Hunan University |
| Yang, Zhi | Hunan University |
Keywords: Human-Robot Interaction and Cooperation
Abstract: Service robots are rapidly entering the food service industry, yet a puzzling phenomenon persists: within the same enterprise and a shared service-robot ecosystem, employees performing comparable tasks construct fundamentally different work relationships—some regard robots as "friends" and "children," others as "colleagues," and still others as "machines" or "nothing." This study conducted semi-structured interviews with 18 employees at a large robot-themed restaurant in China, employing systematic grounded theory to investigate how employee—robot work relationships evolve, diverge, and are driven. The findings reveal that: (1) relationships evolve through novelty, habituation transition, and relational divergence, with the first two phases broadly shared across participants; (2) after habituation, relationships diverge into Affective (friends/children/partners), Functional (colleagues/work partners), and Detached (tools/equipment) types; (3) divergence is driven by personal disposition, role cognition, and social environment; and (4) robot breakdowns primarily function as relational "developers," revealing or reinforcing existing relational forms rather than automatically operating as immediate turning points. This study proposes "Post-Habituation Relational Divergence," filling the theoretical gap regarding "what happens after novelty fades" in longitudinal HRI research.
|
| |
| 16:32-16:44, Paper FrS4A.2 | |
| Implementation and Experimental Evaluation of a Real-Time Vision-Based Human-Following Robot |
|
| Satapanakul, Dhamwarich | Chulalongkorn University |
| Kullatee, Thee | Chulalongkorn University |
| Lerdtada, Patteera | Chulalongkorn University |
| Suraamornkul, Atit | Chulalongkorn University |
| Tang, Jing | Chulalongkorn University |
| Chaichaowarat, Ronnapee | Chulalongkorn University |
Keywords: Mobile Robotics, Human-Robot Interaction and Cooperation, Home and Personal Robot Systems
Abstract: Human-Following Robots (HFRs) are autonomous mobile systems that detect, track, and follow a person in real time. Although modern perception models provide reliable human detection, implementing a complete HFR on a physical robot remains challenging because visual measurements must be converted into stable motion commands under constraints such as perception latency, RGB-D depth noise, static friction, and chassis inertia. This paper presents the implementation and experimental evaluation of a vision-based human-following robot using an NVIDIA Jetson Orin Nano, an Intel RealSense D455 RGB-D camera, YOLOv8-based person detection, target acquisition logic, and a decoupled PID control architecture. The controller independently regulates heading and following distance, while nonlinear angular saturation, deadband filtering, anti-windup clamping, velocity limiting, and emergency-stop functions enhance stability and safety. The system was evaluated indoors through three experiments: linear tracking, dynamic S-walk, and rotational pivot testing. Results show that the most practical following distance is 1.5 to 2.0 m, where the distance RMSE remains below 0.50 m. The S-walk test achieved a distance RMSE of 0.493 m and a centering RMSE of 22°, while the rotational pivot test achieved a centering RMSE of 11.17° with a settling time of 4.4 s. The results demonstrate stable indoor human-following performance and highlight the trade-off between tracking accuracy and mechanical stability.
|
| |
| 16:44-16:56, Paper FrS4A.3 | |
| Comparative Evaluation of Tracking and Re-Identification Methods for Human-Following Robots |
|
| Suraamornkul, Atit | Chulalongkorn University |
| Satapanakul, Dhamwarich | Chulalongkorn University |
| Lerdtada, Patteera | Chulalongkorn University |
| Kullatee, Thee | Chulalongkorn University |
| Tang, Jing | Chulalongkorn University |
| Chaichaowarat, Ronnapee | Chulalongkorn University |
Keywords: Robot Vision and Computer Vision, Human-Robot Interaction and Cooperation, Home and Personal Robot Systems
Abstract: Human-Following Robots (HFRs) require continuous estimation of the correct target rather than generic tracking of all pedestrians. Standard multi-object tracking benchmarks measure detection, localization, and association accuracy, but do not directly represent the safety-critical objective of an HFR: maintaining the identity and position of the designated target across occlusion, crowd interaction, and visual ambiguity. This paper presents a curated proof-of-concept benchmark for evaluating HFR perception pipelines using synchronized RGB-D and LiDAR data. The benchmark evaluates tracker, sensor, and re-identification configurations, including simple 2D association trackers, motion-based trackers, appearance assisted trackers, LiDAR-based trackers, and fusion-oriented variants. A custom Following Score evaluates HFR suitability using target coverage, permanent ID switching, loss severity, identity stability, and fragmentation. Results show no single tracker family dominates all conditions. Appearance-assisted vision trackers perform strongly in normal visibility and dense crowds, while LiDAR-based or fusion-oriented methods are more valuable under severe occlusion. DeepSORT with color-histogram Re-ID achieved the best overall balance, with a mean Following Score of 0.9495 and a worst-case score of 0.8643. The results demonstrate that HFR perception should be selected according to scenario failure mode, not generic tracking accuracy alone.
|
| |
| 16:56-17:08, Paper FrS4A.4 | |
| Natural Manipulator Motion: An Approach Based on Principles of Biomechanics and Animation |
|
| Tchinda Henoch, Wakam | University of Jinan |
| Li, Guangwei | University of Jinan |
| Suanon Ifon Felix, Marcellin | University of Jinan |
| Qu, Haojun | University of Jinan |
| Li, Jinping | University of Jinan |
Keywords: Human-Robot Interaction and Cooperation, Rehabilitation and Assistive Robotics, Path and Motion Planning
Abstract: Natural, human-like motion is not only elegant but also essential for the safe integration of robots into human spaces. While learning-driven approaches such as imitation learning and reinforcement learning have shown promising results, optimizing them for human-like behavior remains challenging due to high data requirements and the difficulty of formalizing 'naturalness' within reward functions. In contrast, principle-driven methods based on biomechanics or animation offer more interpretable solutions that extend to new tasks without retraining, but often focus on isolated aspects of motion, neglecting the structural coupling between physical efficiency and perceptual quality. This paper introduces a principle-driven framework for generating natural manipulator motion that operates independently of costly data collection or blind reward engineering. We formulate a definition of natural motion based on three fundamental pillars: expressiveness, efficiency, and smoothness. These are mapped into a three-tier architecture in which (i) expressive intent is encoded using animation principles in structured motion primitives, (ii) efficient configurations are obtained through constrained inverse kinematics that minimizes unnecessary joint motion, and (iii) smooth execution is achieved via jerk-limited trajectory generation. The proposed method is validated on a 6-degree-of-freedom collaborative robot (Fairino FR5) across a range of interactive tasks.
|
| |
| 17:08-17:20, Paper FrS4A.5 | |
| User Authentication Via Sequence Modeling: Analyzing Task Execution Patterns in Teleoperated Motion Commands |
|
| Tan, Xiye | Hong Kong Baptist University |
| Wu, Xiaoyi | Smith College |
| Huang, Kevin | Smith College |
Keywords: Tele-Robotics/Networked, Cloud Robotics, Human-Robot Interaction and Cooperation, Machine Learning
Abstract: As robotic systems become increasingly integrated into sensitive teleoperated domains (such as surgery and healthcare, hazardous material handling, disaster response, autonomous driving intervention, and military applications), ensuring secure and continuous user authentication is a critical concern. This paper presents a machine learning framework for robust identification of users through behavioral patterns captured in teleoperated robot control sequences. The RoboTurk dataset [1], which comprises end-effector motion data over time from a 7-degree-of-freedom robot manipulator, is utilized to train Long Short-Term Memory (LSTM) networks capable of modeling user-specific motion signatures. The sliding window technique is applied to segment sequences into fixed-length inputs, enabling continuous analysis and simultaneously avoids the limitations of zero-padding and interpolation methods, which require access to complete trajectories. The proposed approach is promising, as it achieves a high average classification accuracy of 93.4% and a stable 87.36% accuracy and 89.16% F1 via five-fold cross-validation, demonstrating its effectiveness in distinguishing between a large number of users. The method generally offers a scalable and real-time compatible solution for secure user identification in human-robot interaction systems
|
| |
| 17:20-17:32, Paper FrS4A.6 | |
| A Partitioned Edge-Region Compensation Method for Magnetically Levitated Planar Motors Using Two-Dimensional Linear Interpolation |
|
| Yuan, Shuai | Shenyang Jianzhu University |
| Wang, Wenjie | Shenyang Jianzhu University |
| Zhang, Feng | Shenyang Jianzhu University |
| Yang, Yongliang | Shenyang Institute of Automation, CAS |
Keywords: Industrial Robotics and Factory Automation, Intelligent Control and Systems
Abstract: A partitioned edge-region compensationtable method based on two-dimensional linear interpolation is proposed to address the increased force/torque modeling error of magnetically levitated planar motors during long-stroke motion and the output jumps caused by discrete lookup-table compensation at offgrid positions. The proposed method is built on a center-region analytical model, while the workspace is divided into center, edge, and corner regions. In the edge and corner regions, a two-dimensional compensation table is constructed from offline high-accuracy numerical results, and bilinear interpolation is used to smoothly reconstruct the compensation coefficients at off-grid positions. A finite element method (FEM)—based numerical co-simulatio co-simulation model is established based on structural parameters reported in the literature. The results show that the twodimensional compensation table effectively reduces the static prediction error of the pure analytical model in edge regions. On the selected off-grid validation path, two-dimensional linear interpolation improves output continuity and achieves lower errors with respect to the FEM reference for most wrench components, reducing the overall RMSE from 66.563 to 37.250 compared with nearest-neighbor lookup. These results indicate that the proposed method improves edge-region modeling accuracy and off-grid output continuity for most wrench components under the tested simulation conditions.
|
| |
| FrS4B |
Yi Hui Hall 6 |
| Robot Perception and Learning |
Regular Sessions |
| Chair: Guo, Di | Beijing University of Posts and Telecommunications |
| |
| 16:20-16:32, Paper FrS4B.1 | |
| Control-Oriented Simulation Framework for Baseline Tracking of Electro-Hydraulic Robotic Arm Joints (I) |
|
| Sharikadze, Lizi | Université Paris-Saclay |
| Marshoud, Abd Alrahman | Université Évry Paris-Saclay |
| AlSharif, Marwa | Higher Institute for Applied Sciences and Technology (HIAST) |
| Sleiman, Maya | KALYSTA |
| Alkasm, Sulaf | Engineering School of Informatics, Electronics and Automations (ESIEA), Paris |
Keywords: Humanoid Robots, Intelligent Control and Systems, Robot Design
Abstract: This paper presents a control-oriented Simscape framework for pre-commissioning assessment of the three electro-hydraulic joints of a fabricated 3-DOF upper-body humanoid arm. The model retains the joint-specific cylinder areas, actuator strokes, crank geometry, servo-valve dynamics, chamber pressures, and CAD-derived mass properties. One fixed controller is retained for each joint across all tests. Joint 1 uses filtered PD feedback, reference-velocity feedforward, and a local holding bias. The bias reduces the simulated 3 s zero-reference drift from 0.306 rad to 4.15 × 10⁻⁵ rad. Under a common 0.15 rad smooth-step command, Joints 1 to 3 achieve RMSE values of 0.0057, 0.0034, and 0.0059 rad, respectively. At 0.5 Hz, normalized RMSE is 3.01% for Joint 1, 1.81% for Joint 2, and 14.57% for Joint 3; the latter also exhibits a 0.18 s phase lag that identifies its present bandwidth limitation. All tested commands remain below 0.77 mA under the specified ±10 mA valve range. Peak chamber pressures do not exceed 177.5 bar, while the most demanding Joint 1 return-to-zero test leaves 0.7 mm of stroke margin. A separate CAD self-weight audit produces less than 6.0 × 10⁻⁶ rad change in zero-reference position at the selected posture. The resulting metrics provide a reproducible baseline for hardware commissioning and later controller comparisons.
|
| |
| 16:32-16:44, Paper FrS4B.2 | |
| Balancing Efficiency and Accuracy: A Comparative Study of UNet-Based Models for Epiphysis Segmentation in Robotic Surgery |
|
| Wang, Yueyang | Beihang University |
| Ling, Yongjian | Capital Medical University |
| Luo, Yanzhong | Beijing Children's Hospital |
| Liu, Haonan | Beijing Children's Hospital |
| Dong, Guo | Beijing Children'sHospital, Capital Medical University, National Center forChildren's Health |
| Liu, Wenyong | Beihang University |
Keywords: Medical Robotics, Deep Learning
Abstract: Accurate segmentation of bone tissues from pre-operative CT images is a prerequisite for automatic path planning in robot-assisted figure-eight plate guided growth surgery (FEPGS). However, pediatric epiphyseal structures exhibit weak boundaries,posing significant challenges for deep learning-based segmentation. In this study, an automatic planning framework for FEPGS was developed, in which representative deep learning segmentation methods were systematically compared to identify the most suitable segmentation strategy for subsequent path planning. Experimental results demonstrated that 2D nnUNet achieved the best balance between segmentation accuracy, boundary consistency, and computational efficiency, making it more suitable for preoperative planning in FEPGS. Furthermore, path evaluation demonstrated that all planned screw paths satisfied the predefined safety criterion, confirming the feasibility of the proposed planning framework. These findings demonstrate that anatomically reliable segmentation is essential for clinically feasible automatic path planning and provide practical guidance for selecting segmentation frameworks in robot-assisted surgery of FEPGS.
|
| |
| 16:44-16:56, Paper FrS4B.3 | |
| Benchmarking Self-Supervised Features for Unsupervised Wind Turbine Blade Anomaly Detection |
|
| Cai, Zhihui | Shenzhen Institutes of Advanced Technology,Chinese Academy of Sciences |
| Zhang, Qingwei | Beijing University of Posts and Telecommunications |
| Liang, Guoyuan | Shenzhen Institutes of Advanced Technology |
Keywords: Deep Learning, Robot Vision and Computer Vision, Machine Learning
Abstract: Although Unsupervised Anomaly Detection (UAD) performs well on academic benchmarks, its effectiveness can diminish in real-world industrial settings. This study evaluates two DINOv2+PatchCore variants and ten representative baseline configurations on a composite wind turbine blade dataset. The results show that the quality of the feature representation has a substantial influence on image-level anomaly detection performance in the evaluated setting. We investigate a parameter-training-free framework that integrates the DINOv2 foundation model with a PatchCore memory bank. By using frozen DINOv2 representations, the framework captures fine-grained visual deviations without domain-specific parameter optimization. On a dataset of 13,439 images, the patch-level variant achieves an AUROC of 0.9710, outperforming Reverse Distillation by 4.5 percentage points among the evaluated configurations. The comparison suggests that general-purpose self-supervised representations can provide a strong basis for wind turbine blade anomaly screening. However, the observed improvement may jointly result from the pretraining objective, model architecture, and pretraining scale, and the current framework detects visual deviations rather than directly determining whether an observed region represents a hazardous defect. The study therefore establishes a strong empirical baseline on the evaluated composite dataset and motivates further controlled and cross-domain evaluation.
|
| |
| 16:56-17:08, Paper FrS4B.4 | |
| CBAM-YOLOv5-Based Rail Surface Defect Detection for Railway Inspection Robots |
|
| Hou, Yuxuan | Yanshan University |
| Sun, Kaijing | Yanshan University |
| Li, Jiaqi | University |
| Wen, Shuhuan | Yanshan University |
| Liu, Xin | Tsinghua University |
| Liu, Chenglin | China Railway Shanhaiguan Bridge Group Co., Ltd |
Keywords: Robot Vision and Computer Vision, Deep Learning, Field Robotics
Abstract: Rail surface defect detection is a key perception task for railway inspection robots. In practical robot-based inspection, the detector must handle multiple defect categories, weak visual features, complex rail textures, and limited onboard computational resources. To address these requirements, this paper presents a CBAM-YOLOv5-based rail surface defect detection method. A rail surface defect dataset is constructed from public and open-source rail images, covering 11 defect categories. After manual filtering, random rotation, brightness augmentation, and LabelImg annotation, the dataset contains 1500 images and is divided into training, validation, and test subsets at a ratio of 8:1:1. To improve recognition of weak and texture-like defects while maintaining a lightweight one-stage detection pipeline, the CBAM attention mechanism is embedded into YOLOv5. The channel and spatial attention branches enhance defect-related feature responses and suppress background interference. Experimental results show that the proposed CBAM-YOLOv5 achieves a mAP@0.5 of 0.799 and an inference speed of 217 FPS. Although YOLOv8 obtains a higher mAP@0.5 of 0.812, its speed is 83 FPS. Therefore, the proposed method provides a practical accuracy-speed tradeoff for railway inspection robots, where real-time inference, lightweight deployment, and stable visual output are all required.
|
| |
| 17:08-17:20, Paper FrS4B.5 | |
| Research on Dynamic Hip Joint Angle Measurement Based on Monocular Vision and IMU Fusion |
|
| Xia, Chun | Shenyang Aerospace University |
| Deng, Dequan | Shenyang Aerospace University |
| Zhirui, Zhao | Shenyang Aerospace University |
| Zeng, Xinyu | Shenyang Aerospace University |
| Liu, Yining | Shenyang Aerospace University |
Keywords: Sensor Networks, Intelligent Control, Machine Learning
Abstract: Accurate measurement of human joint angles is a key fundamental technology in motion rehabilitation assessment and human-machine collaborative control. Addressing the demand for low-cost and portable human motion capture, this study establishes a synchronous measurement system based on monocular vision (MediaPipe) and an inertial measurement unit (IMU). Taking the lower limb stationary leg-raising movement as the typical operating condition, the performance differences between the two sensors in identifying human hip joint angles are systematically compared. Experimental results demonstrate that the visual and inertial data exhibit extremely high consistency in the time domain, with a waveform correlation coefficient ℛ of 0.9748 and an overall root-mean-square error (RMSE) of 6.27°. Nevertheless, limitations exist for each single sensor: at the extreme points of maximum leg elevation (where the thigh is close to the torso), the vision algorithm is susceptible to body self-occlusion, leading to underestimated measurements and signal edge fluctuations. Conversely, the raw IMU signals tend to generate high-frequency jagged burrs during the instantaneous phase of motion direction switching due to soft tissue artifacts. To address these issues, a complementary filtering fusion algorithm based on dynamic confidence is introduced. By real-time evaluation of the dynamic residuals and confidence levelCvis(t)of the vision signals, the fusion weight is dynamically adjusted to trac
|
| |
| 17:20-17:32, Paper FrS4B.6 | |
| A Study on the Recognition of Drivable Areas Based on MGDeepLabv3+ |
|
| Li, Rui | Xi'an Jiaotong University |
| Wang, Qiming | Xi’an University of Technology |
| Yang, Shiqiang | Xi'an University of Technology |
| Deng, Yaxi | Xi'an University of Technology |
| Wei, Xiaoqing | Xi’an University of Technology |
| Liu, Lingtong | Xi’an University of Technology |
| Zhang, Xiaodong | Xi'an Jiaotong University |
Keywords: Deep Learning, Artificial Intelligence, Intelligent Control
Abstract: With the rapid development of deep learning, vision-based environmental perception has become a core technology for fully autonomous driving. However, existing semantic segmentation methods still face challenges related to multi-scale feature representation and model complexity. To address these issues, this study proposes MGDeepLabv3+, an improved semantic segmentation network based on DeepLabv3+. First, the original backbone is replaced with MobileNetV2 to reduce model parameters and computational cost. Second, a dense atrous spatial pyramid pooling structure and hybrid pooling are integrated to construct the MdenseASPP module, enhancing multi-scale feature extraction. An attention-guided upsampling mechanism is further introduced to strengthen the representation of salient features during decoding. In addition, standard convolutions in the decoder are replaced with dilated convolutions to better exploit spatial location and boundary information. Experimental results on public datasets and real-world driving scenes demonstrate that MGDeepLabv3+ achieves high segmentation accuracy with a relatively small number of parameters. The proposed network enables accurate and real-time detection of drivable areas across diverse driving scenarios.
|
| |
| FrPP1 |
|
| Poster Session |
Poster Sessions |
| |
| 10:35-16:40, Paper FrPP1.1 | |
| When Bigger Is Not Better: LLM Scale and User Perception in Short-Duration Social Human-Robot Interaction |
|
| Frederiksen, Morten Roed | IT-University of Copenhagen |
| Volf, Amanda Rasille Røn | IT-University of Copenhagen Denmark |
Keywords: Human-Robot Interaction and Cooperation, Human-Machine Interface, Home and Personal Robot Systems
Abstract: As Large Language Models increasingly power embodied agents, it remains unclear whether computational scaling directly enhances user perception during brief encounters. This paper investigates the impact of model parameter size 4B, 8B, and 30B on perceived intelligence and likability during short-duration, open-ended interactions. These exchanges approximate the brief social encounters found in many service-robot applications. Through a within-subjects study (N=19), we found no statistically significant overall preference for the 30B model over the 4B or 8B variants in intelligence, naturalness, enjoyment, or humor. A significant negative correlation was observed between AI interaction frequency and intelligence ranking for the 30B model (p=.005), suggesting that more experienced users may be more sensitive to differences in model capability. Within the tested interaction regime, the results indicate that increasing parameter scale alone did not produce a detectable perceptual advantage. Interaction-level dynamics such as conversational flow, responsiveness, and socially appropriate behavior may therefore be at least as important as model scale during brief, low-stakes social encounters.
|
| |
| 10:35-16:40, Paper FrPP1.2 | |
| SEA-Prompt: Structured Expert Adaptation Prompting for Zero-Shot Anomaly Detection |
|
| Guo, Hongxi | Changchun University of Architecture and Civil Engineering |
| Zhang, Haozhe | Changchun University of Science and Technology |
| Yang, Yang | Changchun University of Architecture and Civil Engineering |
| Zhao, Guangyu | Changchun University of Science and Technology |
| Tang, Jilong | Changchun University of Science and Technology |
Keywords: Deep Learning, Robot Vision and Computer Vision, Artificial Intelligence
Abstract: Zero-shot anomaly detection (ZSAD) classifies and localizes anomalies in unseen categories without target-domain training images, and recent CLIP-based methods align visual patches with normal and abnormal text prompts. However, existing prompt designs occupy two unsatisfactory extremes. A single learnable abnormality prompt is too rigid to capture the visual heterogeneity of real defects, while independent prompts drift toward the dominant abnormal mode of the auxiliary data and lose the shared semantic that ZSAD relies on for cross-category generalization. We argue that anomalies are simultaneously shared and heterogeneous, and propose SEA-Prompt, a Structured Expert Adaptation Prompting framework on a frozen CLIP backbone. SEA-Prompt decomposes each abnormal prompt into shared normal-context tokens, a shared abnormal basis anchoring the common abnormality concept, and expert-specific residual tokens spanning complementary directions. A multi-level patch fusion module combines shallow and deep ViT features, and an expert-conditioned adaptation mechanism routes this evidence into expert residuals only, with updates suppressed on normal images and experts aggregated through a score-adaptive readout. Across 14 industrial and medical benchmarks, SEA-Prompt achieves competitive performance at both image and pixel levels.
|
| |
| 10:35-16:40, Paper FrPP1.3 | |
| Topology-Aware Cooperation Thresholds in Networked Public Goods Games |
|
| Shangguan, Naiwen | Changchun University of Science and Technology |
| Yang, Yang | Changchun University of Science and Technology |
| Han, Zhengran | Changchun University of Science and Technology |
| Zhang, Haozhe | Changchun University of Science and Technology |
| Tang, Jilong | Changchun University of Science and Technology |
Keywords: Multi-Robot Systems, Intelligent Control and Systems, Sensor Networks
Abstract: Shared sensing, relay communication, and local task allocation in networked autonomous systems often generate rewards at the group level rather than through isolated pairwise links. This paper investigates when such group-level cooperation can be favored by a given interaction topology. A finite population is modeled as an arbitrary graph, and a networked public goods game is formulated in which each node participates in overlapping neighborhood games. Under weak selection, cooperation is evaluated by the fixation probability of a single cooperator and by the critical synergy factor required for cooperation to be favored. Deterministic topologies, ER/WS/BA random networks, enumerated small graphs, and empirical networks are examined. The results show that dense connectivity alone can raise the required synergy, because larger local groups dilute the marginal return of individual contribution. In contrast, separated local groups, rule-dependent clustering, and moderate heterogeneity can lower the cooperation threshold. Compared with pairwise donation games, the public-goods formulation yields more stable and interpretable conditions across heterogeneous structures, providing topology-level guidance for designing cooperative mechanisms in networked multi-agent systems.
|
| |
| 10:35-16:40, Paper FrPP1.4 | |
| Cascaded Domain-Adaptive Vision-Language Model for Transmission Line Hazard Detection |
|
| Li, Zhuyun | Waseda University |
| Yang, Yingyi | China Southern Power Grid Technology Co., Ltd., |
| Mai, Xiaoming | China Southern Power Grid Technology Co., Ltd., |
| Zhang, Lei | Guangzhou Guoyan Zhiyun Technology Co., Ltd |
| Lai, Jiayang | China Southern Power Grid Technology Co., Ltd |
| Qu, Xian | China Southern Power Grid Technology Co., Ltd |
| Li, Shiran | China Southern Power Grid Technology Co., Ltd |
| Liu, Minghao | China Southern Power Grid Technology Co., Ltd., |
| Feng, Jianhao | Smart Operation and Maintenance Business Unit China Southern Power Grid Technology Co., Ltd |
| Xu, Peixin | China Southern Power Grid Technology Co., Ltd |
| Ieiri, Yuya | Waseda University |
| Yoshie, Osamu | Waseda University |
Keywords: Robot Vision and Computer Vision, Emerging Technologies and Applications, Artificial Intelligence
Abstract: Unauthorized construction machinery near highvoltage transmission lines causes line tripping and unplanned outages. Existing monocular surveillance systems suffer from high false-alarm rates because single-lens cameras lack depth information, making it difficult to distinguish genuine proximity threats from perspective-induced visual coincidences. We propose a cascaded hazard detection framework that combines YOLOv12 pre-filtering, dynamic perspective-based electronic fencing, and a domain-adapted Vision-Language Model (VLM) for fine-grained spatial reasoning. We apply LowRank Adaptation (LoRA) to fine-tune Qwen3-VL-8B-Thinking on a semi-automatically constructed, class-balanced dataset of 2,065 image-text pairs, achieving 98.0% F1-score on a heldout balanced test set. The cascaded architecture reduces VLM inference load by 20–100×, making real-time 24/7 corridor monitoring feasible. Systematic ablation studies across model scale, LoRA rank, input resolution, and data balancing reveal the critical role of each design choice.
|
| |
| 10:35-16:40, Paper FrPP1.5 | |
| Toward User-Mediated Self-Repair in Ubiquitous Robots through Goal-Oriented Agentic AI |
|
| Frederiksen, Morten Roed | IT-University of Copenhagen |
Keywords: Human-Robot Interaction and Cooperation, Human-Machine Interface, Home and Personal Robot Systems
Abstract: Ubiquitous robotic systems often lack traditional visual interfaces, necessitating resilient natural language interaction for maintenance and repair tasks. This paper presents a goal oriented agentic AI architecture designed to enable non-expert users to perform technical repairs through situated dialogue. The framework utilizes a multi-layered approach that decouples high-level strategic planning from reactive conversational execution to transform unconstrained human instructions into a structured hierarchy of goals. We conducted a study involving twenty participants to evaluate the system's efficacy using a physical hardware testbed. The architecture achieved a 95% task completion rate, and participants reported positive self-efficacy following real-time guidance that adapted to conversational diversions and linguistic variations. A comparative analysis with an online baseline revealed that the transition to a physical environment significantly decreased perceived social presence (p=.0005), and trust and competence, (p=.037), while the agentic framework remained robust throughout the interaction. These findings indicate that goal oriented agentic AI can support the sustainability of body-worn technologies by empowering users to perform critical maintenance in ubiquitous contexts.
|
| |
| 10:35-16:40, Paper FrPP1.6 | |
| QueryKAN a Variate-Centric and Edge-Parameterized Framework for Remaining Useful Life Prediction |
|
| Han, Zhengran | Changchun University of Science and Technology |
| Yang, Yang | Changchun University of Science and Technology |
| Huang, Fuzhong | Changchun University of Science and Technology |
| Ji, Zhao | Changchun University of Science and Technology |
| Tang, Jilong | Changchun University of Science and Technology |
| Shangguan, Naiwen | Changchun University of Science and Technology |
| Zhang, Haozhe | Changchun University of Science and Technology |
| Lan, Lixiang | Changchun University of Science and Technology |
Keywords: Deep Learning, Machine Learning, Artificial Intelligence
Abstract: To address the challenge of aero-engine remaining useful life prediction under complex operating conditions, this paper proposes QueryKAN-RUL, an end-to-end predictive framework. The proposed method first refines the underlying degradation trend through hierarchical temporal disentanglement, and then employs an inverted self-attention mechanism to explicitly capture cross-sensor fault dependencies in multivariate monitoring signals. On this basis, a state-query aggregation strategy is introduced to distill global health information into a compact representation, which is subsequently fed into a Kolmogorov–Arnold Network for expressive nonlinear life mapping. Experimental results on the C-MAPSS benchmark demonstrate that QueryKAN-RUL consistently outperforms representative baselines in both RMSE and NASA Score, while effectively reducing the high-risk problem of RUL overestimation. These results verify the effectiveness of the proposed framework for accurate and reliable aero-engine prognostics.
|
| |
| 10:35-16:40, Paper FrPP1.7 | |
| A Smooth Locomotion Policy for Humanoid Robots Based on Structural Priors and Continuity Constraints |
|
| Li, Mingqiu | Changchun University of Science and Technology |
| Li, Hengxu | Changchun University of Science and Technology |
| Yang, Yang | Changchun University of Science and Technology |
| Tang, Jilong | Changchun University of Science and Technology |
Keywords: Machine Learning, Humanoid Robots, Intelligent Control
Abstract: PPO-based humanoid locomotion policies often achieve stable walking in simulation but generate high-frequency joint chattering, which reduces motion quality and increases the sim-to-real gap. This paper proposes a smooth locomotion framework that improves the control signal at two levels. First, a phase-driven inverse-kinematics (IK) reference is used to reformulate the action as a residual correction, thereby reducing the effective action amplitude. Second, a Lipschitz continuity policy-gradient penalty (LCP) is added to the PPO actor loss to constrain the local sensitivity of the policy to observation perturbations. Reward terms with overlapping chattering-suppression effects are further simplified. Ablation, heuristic-baseline, parameter-sensitivity, multi-seed and complex-terrain experiments on a JVRC humanoid in MuJoCo show that the proposed method substantially reduces action and joint jitter while preserving walking capability.
|
| |
| 10:35-16:40, Paper FrPP1.8 | |
| WEL: Lightweight Traffic Sign Detection Via Ghost Neck and ECA Attention |
|
| Yang, Yang | Changchun University of Science and Technology |
| Liu, Ruidong | ChangChun University of Science and Technology |
| Li, Mingqiu | Changchun University of Science and Technology |
| Tang, Jilong | Changchun University of Science and Technology |
Keywords: Deep Learning, Artificial Intelligence, Robot Vision and Computer Vision
Abstract: Traffic sign detection on resource-constrained in vehicle platforms demands simultaneous accuracy and compu tational efficiency. We propose WEL, a lightweight detector built upon YOLO11 and TSD-Net with two complementary neck modifications targeting computation, parameter count, and accuracy preservation. First, all C3k2 modules in the PAN-FPN neck are replaced by C3k2 Ghost, which generates intrinsic feature maps via standard convolution and ghost maps via cheap depthwise convolution, substantially reducing neck FLOPs. Second, Squeeze-and-Excitation (SE) channel attention is replaced by Efficient Channel Attention (ECA), which models inter-channel interactions via a 1-D convolution with adaptive kernel size, avoiding the dimensionality-reduction bottleneck of SE at a cost of only k ≈ 5 parameters per module. On TT100K, WEL achieves 92.33% mAP@50 at 57.5 GFLOPs with 7.27M parameters, reducing computation by 16.8% and parameters by 6.2% while maintaining accuracy (+0.10% mAP@50) over TSD-Net (69.1 GFLOPs, 92.23%, 7.75M). Ablation experiments confirm that Ghost and ECA are complementary: Ghost reduces neck FLOPs at a small accuracy cost; ECA compensates for that cost at negligible parameter overhead, preserving the lightweight gains end-to-end. Among YOLO-based detectors on TT100K from 2023–2025 reporting >90% mAP@50, WEL attains the smallest parameter count while matching or exceeding their accuracy.
|
| |
| 10:35-16:40, Paper FrPP1.9 | |
| PhysD(R, O): Physics-Informed Diverse Dexterous Grasp Generation |
|
| Zhu, Zifan | Shanghai Jiao Tong University |
| Zheng, Jianxin | Shanghai Jiao Tong University |
| Zhuang, Chungang | Shanghai Jiao Tong University |
Keywords: Grasping and Manipulation, Machine Learning, Robot Vision and Computer Vision
Abstract: Generating dexterous grasps requires balancing configuration diversity with geometric and physical feasibility. Existing CVAE-based approaches commonly use condition-independent latent priors, which may produce concentrated grasp distributions when applied to objects with substantially different geometries. This paper presents a generative framework built on the mathcal{D(R,O)} hand-object interaction representation. A dynamic conditional prior and Feature-wise Linear Modulation are introduced to adapt latent sampling and feature decoding to the input geometry. Distance-aware reconstruction is further used to emphasize potential contact regions, while a differentiable physics regularization term penalizes contact-wrench imbalance and hand-object penetration. Experiments on 100 objects from the Google Scanned Objects dataset using the DexHand platform show that the proposed method achieves the highest analytical Q_1 value and force-closure rate among the evaluated methods and produces a broader distribution of finger configurations than the evaluated learning-based baselines. The simulation results further characterize the trade-off between configuration dispersion and dynamic grasp robustness.
|
| |
| 10:35-16:40, Paper FrPP1.10 | |
| Constraint-Aware Trajectory Planning for Continuous Robotic Surface Coverage Via Kinematic Consistency and Hierarchical GTSP |
|
| Gao, Xianglin | The Research Institute for Special Structures of Aeronautical Composite the Aviation Industry Corporation of China |
| Sun, Xu | Shandong University |
| Zhuang, Yi | Shandong University |
| Zhang, Zhaofeng | Shandong University |
| Zhou, Lelai | Shandong University |
Keywords: Path and Motion Planning, Grasping and Manipulation, ROS, Software System for Robotics Application
Abstract: This paper presents a constraint-aware coverage trajectory planning algorithm for composite manipulators operating on complex free-form surfaces. First, a hierarchical graph based route optimization framework formulates the joint space coverage sequence as a generalized traveling salesman problem. A coarse-to-fine propagation strategy coupled with reconfiguration-aware edge costs reduces the graph size while retaining kinematically valid transitions. Second, an anti tangling inverse kinematics selection mechanism uses joint angle unwrapping and terminal-joint envelope constraints to limit excessive terminal-joint excursions. Finally, a gravity aware pose feasibility filter enforces surface-normal alignment and rejects candidates that violate the prescribed orientation threshold. Simulations show reduced joint displacement and reconfiguration counts under the evaluated configuration. The resulting pose–configuration sequence provides a kinematically informed input to downstream motion planning with MoveIt.
|
| |
| 10:35-16:40, Paper FrPP1.11 | |
| Ultrasonic Defect Classification in Composite Materials Via Dynamic Modality Aggregation and Graph Attention Network |
|
| Zhang, Qi | AVIC Research Institute for Special Structures of Aeronautical Composite |
| Liu, Shengjie | Shandong University |
| Mao, Zihao | Shandong University |
| Zhang, Zhaofeng | Shandong University |
| Chen, Qizhi | Shandong University |
| Zhou, Lelai | Shandong University |
Keywords: Deep Learning, Artificial Intelligence
Abstract: Reliable defect classification in carbon fiber reinforced polymer structures is crucial for automated robotic inspection, but this classification remains challenging due to the heterogeneous acoustic responses of different defect types and the inherent noise of ultrasonic A-scan signals. Existing deep learning classifiers either rely on single-domain feature representations or process each signal independently, producing significant classification ambiguities between morphologically similar defects. This study proposes DMGAT, a framework integrating a dynamic modality aggregation mechanism with a Graph Attention Network (GAT) for multi-domain ultrasonic defect classification. Intra-sample cross-modal edges encode cross-domain interactions within each signal, while inter-sample modality-specific KNN edges connect acoustically similar samples within each feature domain. The multi-head GAT algorithm propagates information between the two edge types, while dynamic modality aggregation captures the most discriminative post-propagation representation for each sample. Evaluations on a simulated dataset containing 1200 samples covering four structural states (normal, inclusion, delamination, and porosity) show that the proposed framework achieves an accuracy of 98.33%, outperforming baseline methods. Dynamic modality aggregation improves accuracy by 5.4% compared to the standard GAT baseline method, demonstrating the necessity of adaptive multi-domain feature weighting.
|
| |
| 10:35-16:40, Paper FrPP1.12 | |
| Color Image Encryption Based on a 2D Hyperchaotic Map and Plaintext-Bound Perturbation |
|
| Li, Peiyuan | Changchun University of Science and Technology |
| Yang, Yang | Changchun University of Science and Technology |
| Li, Yifeng | Changchun University of Science and Technology |
| Zhang, Haozhe | Changchun University of Science and Technology |
| Tang, Jilong | Changchun University of Science and Technology |
|
|
| |
| 10:35-16:40, Paper FrPP1.13 | |
| Real-Time Weld Seam Extraction and Tracking Based on Incremental Point Cloud |
|
| Yang, Yifan | Shandong University |
| Chen, Baining | Shandong University |
| Tian, Xincheng | Shandong University |
| Li, Yibin | Shandong University |
| Shi, Zhibin | Sinopec Petroleum Engineering Co., Ltd., China |
| Zhang, Xiaoming | Army Arms University of PLA, National Key Laboratory of Intelligent Parallel Technology, Beijing, 100072, China |
| Zhou, Lelai | Shandong University |
Keywords: Industrial Robotics and Factory Automation, Robot Vision and Computer Vision
Abstract: This paper proposes a real-time weld seam extraction and tracking method based on incremental point cloud. Continuously acquired line-profile point clouds are organized into an incremental window, which is updated using an asynchronous buffering mechanism to support continuous welding reference points output. A historical-prior-based 3D ROI is constructed using the seam line and extension direction fitted from the previous window. Within the ROI, DBSCAN is used to segment the seam-side point clouds, and RANSAC is adopted to fit two local planes. The intersection of the fitted planes is taken as the seam line of the current window, and welding reference points are generated according to the number of buffered frames. Simulation experiments on a parameterized v-groove butt workpiece show that the proposed method can achieve stable seam extraction, smooth welding reference points generation, and real-time processing under controlled point cloud disturbances. The results also demonstrate the trade-off between computational efficiency and extraction stability under different window lengths.
|
| |
| 10:35-16:40, Paper FrPP1.14 | |
| Fusing Visual and Perceptual Priors: An Anthropomorphic Grasping Control Method for Underactuated Dexterous Hands |
|
| Du, Tingyan | Nanjing University of Aeronautics and Astronautics |
| Duan, Jinjun | Nanjing University of Aeronautics and Astronaut |
| Zhuang, Ming | China International Engineering Consulting Corporation |
| Yu, Yanzhao | Nanjing University of Aeronautics and Astronautics, Qinhuai District, Nanjing, Jiangsu Province |
| Wu, Chendong | Nanjing University of Aeronautics and Astronautics |
| Bin, YiMing | Nanjing University of Aeronautics and Astronautics |
| Qi, Zeyu | Nanjing University of Aeronautics and Astronautics |
| Wang, Lingyu | Nanjing University of Aeronautics and Astronautics |
| Wang, Zhengwei | Nanjing University of Aeronautics and Astronautics |
| Miao, Yunfei | Nanjing University of Aeronautics and Astronautics |
| Zhang, Jiaming | Nanjing University of Aeronautics and Astronautics |
| Tian, Wei | Nanjing University of Aeronautics and Astronautics |
Keywords: Grasping and Manipulation, Deep Learning, Human-Robot Interaction and Cooperation
Abstract: Significant challenges arise for underactuated dexterous hands when grasping thin, flat, and fragile objects without tactile sensing. To address these challenges, a cross-modal anthropomorphic grasping control method integrating visual and proprioceptive priors is proposed to enable dexterous manipulation of such items. First, a prior knowledge base is constructed from human demonstration data collected using data gloves. A "Thumb-Reference" method is then introduced to temporally decouple synchronous multi-finger motions, enabling accurate parsing of human operational experience. Second, a safe workspace is established by integrating YOLOv12-based visual instance segmentation. Concurrently, a 1D-CNN-LSTM temporal observer, augmented with first-order derivatives of physical features, is designed to enable rapid prediction of contact states during human-robot interaction. Furthermore, corresponding intervention behaviors are adaptively planned based on the predicted states (e.g., Slip, Crush, and Secure), thereby achieving an anthropomorphic, compliant grasp that is grounded in human experience. Extensive experiments demonstrate that the proposed framework accurately recognizes thin and fragile objects and grasps them in an anthropomorphic manner, while effectively ensuring the safety of human-robot interaction. This work provides crucial technical support for safe human-robot collaboration in applications such as domestic care and precision assembly.
|
| |
| 10:35-16:40, Paper FrPP1.15 | |
| DOCC-Based Unsupervised Learning System for Detecting Stator Pin Insertion Defects in Blower Motors |
|
| Kim, Donghun | Dept. of Electrical Engineering, Kyungnam University |
Keywords: Industrial Robotics and Factory Automation, Intelligent Control and Systems, Machine Learning
Abstract: This study, in order to overcome the limited data environment experienced in the actual industrial sites where it is difficult to obtain defective image data of defective products, a system for detecting defective automobile blower motors based on DOCC (Deep One-Class Classification) that determines defects by learning only normal data is proposed. In order to verify the hardware simplification and economic efficiency of the inspection system, dual side camera configurations is analyzed to solve the viewing angle problem caused by lateral curvature. As a result of the experiment, the proposed inspection system shows that it efficiently detects defective stator pin insertion of the blower motor.
|
| |
| 10:35-16:40, Paper FrPP1.16 | |
| Bivariate Synergistic Adaptive Control for All-Position Pipeline Welding Robots Based on ANFIS |
|
| Chen, Baining | Shandong University |
| Yang, Yifan | Shandong University |
| Tian, Xincheng | Shandong University |
| Wang, Yingwei | Sinopec Petroleum Engineering Co., Ltd |
| Zhang, Xiaoming | Army Arms University of PLA, National Key Laboratory of Intelligent Parallel Technology, Beijing, 100072, China |
| Li, Yibin | Shandong University |
| Zhou, Lelai | Shandong University |
Keywords: Industrial Robotics and Factory Automation, Intelligent Control, ROS, Software System for Robotics Application
Abstract: During the construction of long-distance oil and gas pipelines, all-position welding robots face challenges from variations in groove morphology. Constant parameter control under these conditions frequently causes depression and protrusion defects. This paper proposes an adaptive control strategy based on the Takagi-Sugeno-Kang (TSK) Adaptive Neuro-Fuzzy Inference System (ANFIS) to address these welding defects. A simulation platform constructed using ROS and Gazebo establishes the mapping relationship between nonlinear disturbances and actuator control commands. Using the real-time geometric error and spatial posture angle as inputs, the proposed ANFIS implements a decoupled adjustment of the wire feed speed and the travel speed. An error backpropagation algorithm optimizes the TSK consequent parameters using offline experimental datasets. This process extracts a compensation trend to counter the effects of gravity on the molten pool from the flat position to the overhead position. Continuous dynamic tracking simulations demonstrate that the trained controller exhibits smooth and prompt responses to sudden morphological defects across the full circumferential trajectory. To maintain arc stability, the nonlinear decoupling strategy restricts changes to the wire feed speed under minor disturbances, activating dual-velocity adjustments only for severe defects. This control approach improves the formation quality and reliability of all-position pipeline welding.
|
| |
| 10:35-16:40, Paper FrPP1.17 | |
| Reinforcement Learning for Multi-Level High Platforms Crossing with Vision-Based Legged Robot |
|
| Huang, Rundong | Shandong University |
| Jingyu, Sun | Shandong University |
| Wang, Zixuan | Shandong University |
| Li, Yibin | Shandong University |
| Zhang, Xiaoming | Army Arms University of PLA, National Key Laboratory of Intelligent Parallel Technology, Beijing, 100072, China |
| Zhou, Lelai | Shandong University |
Keywords: Deep Learning, Biologically Inspired Robotics, Artificial Intelligence
Abstract: Legged robots can adapt to complex and continuously changing terrains. However, accurate terrain recognition and large-height jumping remain challenging tasks. Traditional frameworks rely on fixed predefined actions, leading to limited flexibility and suboptimal timing performance. To address these limitations, this study proposes a reinforcement learning control policy based on depth images. A front-mounted depth camera captures depth images, which are incorporated into the observation queue to enable proactive robot responses. Furthermore, a transitional teacher-student encoder structure is introduced to facilitate smooth transfer of the capabilities of the teacher encoder, which leverages privileged information, to a student encoder based on proprioceptive and image observations. To improve adaptation to step platforms of varying heights, contrastive learning is further incorporated to enhance the environmental discrimination capability of the robot. Finally, a reward function based on position differences enables the robot to freely explore diverse strategies for traversing platforms of different heights. Real-world experiments deploy the RL control policy on the DeepRobotics Lite3 robot and demonstrate successful jumps onto platforms with heights of up to 50 cm.
|
| |
| 10:35-16:40, Paper FrPP1.18 | |
| Research on Surface Defect Detection Algorithm for Steel Strips Based on Improved YOLOv11 |
|
| Li, Mingqiu | Changchun University of Science and Technology |
| Yang, Peng | Changchun University of Science and Technology |
Keywords: Robot Vision and Computer Vision, Deep Learning, Industrial Robotics and Factory Automation
Abstract: To address the challenges of severe scale variations, difficult small-object detection, and complex background interference in steel strip surface inspection , this paper introduces an improved YOLOv11-based defect detection algorithm designated as YOLOv11-Steel. First, a lightweight multi-scale fusion structure, termed the C3K2-Vim module, is designed to enrich multi-scale feature representations by integrating the EfficientViM mechanism. Second, a Dynamic Hyperbolic Tangent (DyT) mechanism is incorporated into the backbone network's C2PSA module to enhance training stability and feature discrimination. Finally, the DySample dynamic upsampling operator is adopted to boost the localization accuracy of subtle anomalies. Experimental results demonstrate that the proposed model achieves a 78.2% mAP@50 on the augmented NEU-DET dataset , yielding a 3.5% improvement over the baseline YOLOv11. These findings validate the effectiveness of YOLOv11-Steel and its substantial engineering potential for high-precision industrial quality control
|
| |
| 10:35-16:40, Paper FrPP1.19 | |
| CAPose: Robust Human Pose Estimation for Robotic Vision in Crowded and Occluded Scenes |
|
| Jianxin, Cui | Changchun University of Science and Technology |
| Bai, Xuemei | Changchun University of Science and Technology |
| Hu, Hanping | Changchun University of Science and Technology |
Keywords: Deep Learning, Robot Vision and Computer Vision
Abstract: 一个名为CAPose的稳健人体姿态估计框架是建议用于拥挤和遮蔽场景。该方法的目标 通过减少特征来提升关键点本地化 错位、跨人干扰与跨尺度 语义不一致。它首先缓解了肤浅 由于局部几何变形引起的错位,通过 可变形残余适应。然后它会被强化 通过建模实现骨干中的空间信息响应 长距离依赖关系和抑制干扰 邻居和杂乱的背景。进一步说明 提高音阶间的语义一致性,颈部 执行注意力引导的双向功能 聚合。此外,还有大内核上
|
| |
| 10:35-16:40, Paper FrPP1.20 | |
| An Efficient Spatial–Temporal Modeling Framework for Low-Channel EEG-Based Affective State Classification |
|
| Dong, Haiyu | Changchun University of Science and Technology |
| Bai, Xuemei | Changchun University of Science and Technology |
Keywords: Artificial Intelligence, Brain-Machine Interface, Deep Learning
Abstract: 电图(EEG)的情绪识别播放 在情感计算和人机合作中扮演着重要角色 互动。然而,脑电信号表现出复杂的空间特征 通道间的依赖关系与动态时间 特性,实现准确的时空建模 很有挑战性。为了解决这个问题,改进了 基于SGCRNN的时空框架被提出 低通道脑电情绪识别。谱图 卷积用于捕捉空间关系 在脑电图通道中,而双向门控循环则是 单位(BiGRU)取代了原有的GRU,以增强时间表现 依赖建模。此外,还有一个高效通道 注意力(ECA)机制被
|
| |
| 10:35-16:40, Paper FrPP1.21 | |
| DGL-YOLO-PD: A Lightweight Detection Algorithm for Distracted Driving Based on Gated Fusion and Structure-Tailored Pruning |
|
| Ba, Kunze | Changchun University of Science and Technology |
| Zhang, Chenjie | Changchun University of Science and Technology |
| Hu, Hanping | Changchun University of Science and Technology |
Keywords: Deep Learning, Artificial Intelligence
Abstract: Distracted driving detection requires fine hand-object localization, robust behavior understanding, and low-cost edge inference. Existing lightweight detectors often enhance representation and compress the network separately, so subtle behavioral cues are easily weakened after pruning. This paper proposes DGL-YOLO-PD, a dual-stage detector built on YOLOv8n. First, Adaptive Multi-scale Gated Fusion (AMGF) extracts local hand-object details and global cabin context, while C2f-iMLCA suppresses redundant channels with an inverted-residual attention structure. Second, fusion- and attention-aware dependency-guided pruning compresses the enhanced topology without breaking residual, concatenation, attention, high-resolution, or gated-fusion dependencies. Balanced knowledge distillation recovers the pruned student without changing its inference graph. On StateFarm and AUC datasets, unpruned DGL-YOLO reaches 98.83% mAP50 and 96.79% mAP50–95. After 5.0× pruning and distillation, DGL-YOLO-PD keeps only 0.647 M parameters, 2.4 GFLOPs, and a 1.5 MB model size, while recovering to 98.07% mAP50 and maintaining 40–60 ms model-side inference on Jetson Nano.
|
| |
| 10:35-16:40, Paper FrPP1.22 | |
| Touch-Based Teleoperated Control Framework for the Underwater Manipulator |
|
| Yang, Dekun | Wuhan University of Technology |
| Luo, Jing | Wuhan University of Technology |
Keywords: Tele-Robotics/Networked, Cloud Robotics
Abstract: This paper presents a touch-based teleoperation system for controlling an underwater robotic manipulator mounted on a remotely operated vehicle (ROV). A Monte Carlo workspace sampling method is proposed to characterize the reachable spaces of both the 3D Systems Touch haptic device and the Schilling Orion 7P manipulator, establishing a linear mapping between the two heterogeneous devices. The system is implemented through a ROS-based cascaded control pipeline (outer Cartesian proportional / inner jointvelocity PID) and validated in the UUV Simulator (Gazebo) environment. Two categories of experiments are conducted: (1) a trajectory tracking experiment that quantifies open-loop endeffector following accuracy, and (2) pick-and-place experiments under zero-current and vc = 1 m/s current disturbance conditions. Hydrodynamic effects on control accuracy are analyzed by correlating Fossen’s equations of motion with experimentally observed ROV attitude perturbations and end-effector tracking errors. Results show that the zero-current overall RMSE is 7.96 cm, increasing to 8.87 cm under 1 m/s current, while the ROV roll standard deviation rises from 0.27◦ to 5.11◦, demonstrating the significant influence of hydrodynamic disturbances on teleoperation precision.
|
| |
| 10:35-16:40, Paper FrPP1.23 | |
| GUIDE: Global-Uncertainty Integrated Depth-Topology Navigation |
|
| Li, Yilin | Boston University |
Keywords: SLAM and Navigation, Path and Motion Planning, ROS, Software System for Robotics Application
Abstract: To address the problems of local perception instability, visual depth error accumulation, topology node mismatch, and conflicts between local obstacle avoidance and global paths in long-distance autonomous navigation of robots in scenarios such as park inspection, warehouse transportation, and underground utility tunnel navigation, this study proposes a long-distance autonomous navigation method for robots based on uncertain depth perception and topology consistency constraints. First, this method uses a visual depth estimation network to obtain the local spatial structure from monocular RGB images and constructs a depth uncertainty map using multi-frame consistency and edge depth perturbation. Second, the navigation environment is represented as a topological graph structure composed of key nodes, traversable edges, and dynamic edge weights, and topology node matching is completed by combining visual descriptors, local geometry, and historical motion constraints. Finally, a joint cost function integrating obstacle distance, depth uncertainty, topology direction deviation, path smoothness, and target attraction was constructed to achieve consistent coordination between local traversability decisions and global topology paths. Simulation experiments showed that the proposed method achieved a navigation success rate of 94.6% in long-distance navigation tasks, with an average path efficiency of 0.91 and an average collision count reduced to 0.18 per task.
|
| |
| 10:35-16:40, Paper FrPP1.24 | |
| An Exposure-Robust 3D Gaussian Splatting Framework for Overexposed Scenes |
|
| Li, Maochen | Beihang University |
| Zhang, Yonghong | Beihang University |
| Ji, Xuquan | Beihang University |
| Wang, Guikai | Beihang University |
| Hu, Lei | Beihang University |
Keywords: Robot Vision and Computer Vision, Machine Learning, Artificial Intelligence
Abstract: 3D Gaussian Splatting (3DGS) achieves remarkable novel view synthesis performance, yet its training relies heavily on the photometric consistency of input images. When fed with overexposed images, texture and gradient information in saturated regions is severely lost, significantly degrading the reconstruction quality. To address this, we propose an Exposure-Robust 3DGS Training Framework. First, an inverse Retinex preprocessing module without learnable parameters is designed to convert overexposed images into enhanced ones with restored intrinsic colors. Second, a photometric-geometric decoupled supervision strategy is constructed, where the enhanced images supervise color and texture, while the gradient field of the original overexposed images supervises the geometric structure, thus preventing enhancement artifacts from interfering with 3D optimization. Furthermore, three pixel-wise weights—luminance weight, local contrast weight, and adaptive gradient weight—are introduced to adaptively regulate the supervision contribution from different regions. Experiments on the LOM overexposure dataset demonstrate that our method achieves the best overall performance in terms of PSNR, SSIM, and LPIPS, offering an effective solution for high-quality 3D reconstruction under overexposed conditions.
|
| |
| 10:35-16:40, Paper FrPP1.25 | |
| Near-Field Blind-Zone Compensation and Voxel Evidence Fusion for Quadruped Terrain Assessment |
|
| Wei, Jinke | HangZhou City University |
| Liu, Yan | Hangzhou City University |
| Li, Yanjun | HangZhou City University |
| Zhang, Junfeng | Hangzhou Sotry Automatic Control Tech Co., Ltd |
| Qiu, Zhenbing | Hangzhou City University |
| Cui, Chenhuan | Hangzhou City University |
| Jiang, Xinze | Hangzhou City University |
| Cai, Yongbin | Hangzhou Sotry Automatic Control Tech Co., Ltd |
Keywords: Sensor Networks, Robot Vision and Computer Vision, ROS, Software System for Robotics Application
Abstract: Near-field terrain perception on quadruped robots is affected by LiDAR blind zones, depth-camera noise, and cross-sensor geometric conflicts. This paper presents a per-frame voxel evidence fusion pipeline for Livox Mid-360 and Intel RealSense D435i data. A Mid-360-triggered cache bounds depth-frame age and timestamp skew, while source-dependent weights and a conjunctive Z-conflict test select either a weighted centroid or a single sensor observation. Unlike occupancy or TSDF mapping, the method does not integrate a persistent map. We analyze 24 archived ROS 2 bags; because configurations were recorded in separate runs, their mode-level results are descriptive rather than paired. A prediction-blind author audit labelled 240 fixed-time RGB frames. Across all configurations, accuracy was 64.6% (95% bag-cluster bootstrap interval 49.2%–78.8%) and macro F1 over the three observed target classes was 0.361. The recalls for Flat, Stairs, and Obstacle were 0.702, 0.100, and 0.400, respectively; only five Obstacle frames and no Slope frames occurred. These results support near-field coverage analysis but not general semantic-accuracy or locomotion claims; the remaining limitations are reported explicitly.
|
| |
| 10:35-16:40, Paper FrPP1.26 | |
| Spatiotemporal Decoupling and Fusion Method for Fire Smoke Detection in Complex Forest Scenes |
|
| Yang, Fengshuo | Changchun University of Science and Technology |
| Zhang, Chenjie | Changchun University of Science and Technology |
| Hu, Hanping | Changchun University of Science and Technology |
Keywords: Deep Learning, Machine Learning, Artificial Intelligence
Abstract: 森林环境 0351;烟雾探测变得ࢱ 6;难 因为烟雾在全画 4133;时看起来像自 2;雾 时间建模引入了 2321;重的计算量和 无关背景运动。 6825;项工作提出了ߌ 8;个 探测器引导的ROI时 间区分流程 将任务解耦为烟 8654;候选人定位, 本地动员验证。 6731;量级 YOLOv11-FasterNet-BiFPN-C2PSA探测器ཛ 8;次提供 空间先验,以及 4102;有SE-ConvLSTM的光流 仅在检测到的投 6164;回报率内应用ߣ 7;区分类似烟雾 雾状干扰的扩散 2290;实验显示 静态定位模块实 9616;了94.80%的mAP@0.5,且 93.25%mAP@0.5:0.95,而完整 30340;管道则获得 二分类࠭
|
| |
| 10:35-16:40, Paper FrPP1.27 | |
| An Accurate Recognition and Localization Method for Randomly Stacked 3C Components Based on Subpixel Edge Detection under Complex Illumination Conditions |
|
| Yu, Yanzhao | Nanjing University of Aeronautics and Astronautics, Qinhuai District, Nanjing, Jiangsu Province |
| Duan, Jinjun | Nanjing University of Aeronautics and Astronaut |
| Qi, Zeyu | Nanjing University of Aeronautics and Astronautics |
| Bin, YiMing | Nanjing University of Aeronautics and Astronautics |
| Wu, Chendong | Nanjing University of Aeronautics and Astronautics |
| Du, Tingyan | Nanjing University of Aeronautics and Astronautics |
| Wang, Zhengwei | Nanjing University of Aeronautics and Astronautics |
| Wang, Lingyu | Nanjing University of Aeronautics and Astronautics |
| Miao, Yunfei | Nanjing University of Aeronautics and Astronautics |
| Zhang, Jiaming | Nanjing University of Aeronautics and Astronautics |
| Tian, Wei | Nanjing University of Aeronautics and Astronautics |
Keywords: Robot Vision and Computer Vision, Grasping and Manipulation
Abstract: During the automated assembly of 3C motherboards, delicate components such as CPUs and memory modules are frequently supplied in randomly stacked piles, resulting in issues such as mutual occlusion, metallic reflections, and complex backgrounds. To address these issues, a method for accurate recognition and localization of randomly stacked 3C components based on subpixel edge detection under complex illumination conditions is proposed. First, a dynamic-factor-based preprocessing fusion algorithm is proposed, in which local grayscale statistics are utilized to adaptively suppress metallic reflections and enhance edge contrast, thereby improving the image quality of randomly stacked components. Second, a subpixel edge detection method based on Zernike moments is introduced, by which the contour localization accuracy is elevated to the subpixel level; combined with an adaptive-threshold early-termination mechanism, rapid and accurate classification of the stacked components is achieved. Furthermore, a "matching-score-threshold and top-layer-priority" grasping strategy is designed, in which the shape matching score is employed as a proxy indicator of the stacking hierarchy and the grasping sequence is automatically determined. Finally, experiments were conducted to test the feasibility of the proposed algorithm. the experimental results show that the proposed algorithm achieves a recognition accuracy of 96.8% and a grasping success rate of 95.2%. The proposed method offers a prom
|
| |
| 10:35-16:40, Paper FrPP1.28 | |
| A Multimodal Intervention-Aware Reward Learning Method for Human-In-The-Loop Reinforcement Learning in Robotic Assembly |
|
| Wu, Chendong | Nanjing University of Aeronautics and Astronautics |
| Duan, Jinjun | Nanjing University of Aeronautics and Astronaut |
| Yu, Wenjin | Rokae (Beijing) Co. Ltd |
| Yu, Yanzhao | Nanjing University of Aeronautics and Astronautics, Qinhuai District, Nanjing, Jiangsu Province |
| Du, Tingyan | Nanjing University of Aeronautics and Astronautics |
| Bin, YiMing | Nanjing University of Aeronautics and Astronautics |
| Qi, Zeyu | Nanjing University of Aeronautics and Astronautics |
| Wang, Zhengwei | Nanjing University of Aeronautics and Astronautics |
| Wang, Lingyu | Nanjing University of Aeronautics and Astronautics |
| Miao, Yunfei | Nanjing University of Aeronautics and Astronautics |
| Zhang, Jiaming | Nanjing University of Aeronautics and Astronautics |
| Tian, Wei | Nanjing University of Aeronautics and Astronautics |
Keywords: Human-Robot Interaction and Cooperation, Intelligent Control, Machine Learning
Abstract: Robotic assembly is being widely adopted for precision assembly of 3C components. However, reinforcement-learning-based assembly methods still suffer from low data efficiency, long training time, and limited recovery capability under contact jamming. To address these issues, a multimodal intervention-aware reward learning method is proposed for human-in-the-loop reinforcement learning. First, a sparse reward model combining a ResNet-based visual success classifier and a contact-safety penalty is constructed to provide task-level supervision and contact-safety constraints. Second, a multimodal intervention-aware progress reward is learned from visual observations, robot states, and force/torque feedback, providing dense process-level feedback for policy learning. In addition, an alternating bilevel training framework is developed, where the inner loop updates the assembly policy and the outer loop refines the reward model to improve its ability to distinguish assembly progress. Finally, a robotic assembly platform is built to evaluate the proposed method on RJ45, USB, and RAM insertion tasks. Experimental results show that, under the same training budget, the proposed method improves the final success rate of the assembly policy by 28.5% compared with conventional human-in-the-loop reinforcement learning (HIL-RL). In the RAM insertion task, the robot achieves a 94.6% final success rate and an 83.9% jamming-recovery success rate, demonstrating the effectiveness of the proposed
|
| |
| 10:35-16:40, Paper FrPP1.29 | |
| Bundle Adjustment Optimization Method for Cabin Segment Docking Pose Estimation Based on Single Camera |
|
| Ju, Gang | Changchun University of Technology |
| Liu, Tianhao | Changchun University of Technology |
| Wang, Zhe | Changchun University of Technology |
| Jiang, Changhong | School of Electrical and Electronic Engineering, Changchun University of Technology |
| Xie, Mujun | Changchun University of Technology |
Keywords: Robot Vision and Computer Vision, Industrial Robotics and Factory Automation, Human-Robot Interaction and Cooperation
Abstract: This paper proposes a robust and constrained bundle adjustment optimization method for single camera cabin segment docking. A unified optimization framework is constructed by incorporating prior covariance weighting, GNC based robust optimization, and adjacent frame pose smoothness constraints into the same bundle adjustment objective, so as to suppress pose jitter and optimization instability caused by degraded observations and local measurement corruption. The EPnP solution is used as the initial estimate, while an outer GNC iteration with the McClure robust kernel and an inner LM iteration are jointly employed for nonlinear optimization. Real-time experiments under clean, illumination variation, and random occlusion conditions demonstrate that the proposed method can provide stable and smooth physical trajectory estimates for large-segment docking, with good real-time performance and practical engineering value as a vision-based measurement solution.
|
| |
| 10:35-16:40, Paper FrPP1.30 | |
| Bend-Grasper: Target-Conditioned Whole-Body Control for Cross-Reach-Domain Humanoid Grasping |
|
| Liu, Zijie | Wuhan University of Science and Technology |
| Wenhui, Huang | Wuhan University of Science and Technology |
| Lin, Yunhan | Wuhan University of Science and Technology |
| Liu, Yang | Wuhan University of Science and Technology |
| Min, Huasong | Robotics Institute of Beihang University of China |
Keywords: Humanoid Robots
Abstract: Abstract—Humanoid robots are increasingly being applied to home service scenarios. They require not only strong mobility but also effective whole-body stabilization to grasp targets beyond the reachable workspace of the humanoid robotic end-effectors in both upright and squatted postures. Therefore, the sample-efficient training of high-dimensional action policies and the high-fidelity tracking of the end-effector are crucial. To address these challenges, we present a target-constrained whole-body stabilization framework denoted as Bend-Grasper. Firstly, a three-stage Selective Policy Inheritance Curriculum (SPI-Curriculum) is proposed to master locomotion, postural compensation, and fine-grained end-effector control in a progressive manner. Secondly, to stabilize policy learning and mitigate catastrophic forgetting, we concurrently employ selective action-mean regularization alongside KL regularization. And again, we develop a target-conditioned whole-body control (TC-WBC) algorithm based on a novel goal-conditioned observation space formulation. Guided by a geometric reachable workspace constraint, the algorithm enables the robot to cooperatively generate full-body loco-manipulation trajectories while proactively executing compensatory behaviors such as squatting and torso bending. Finally, an end-effector residual compensation module is introduced to rectify grasping errors induced by these dynamic postural changes.
|
| |
| 10:35-16:40, Paper FrPP1.31 | |
| RegAct: A Registry-Constrained Text-To-Action Framework with Closed-Loop Execution for Edge Robots |
|
| Zhuchen, Zhong | Fuzhou University Zhicheng College |
| Xu, Wenjie | Fuzhou University Zhicheng College |
| Ma, Yunying | Fuzhou University Zhicheng College |
Keywords: Artificial Intelligence, Human-Machine Interface, Mobile Robotics
Abstract: Chinese voice entrances for smart homes and indoor robots must convert colloquial commands into executable action sequences while remaining responsive under edge-device budgets and offline operation. This paper proposes RegAct,a registry-constrained text-to-action parser that combines fast keyword routing, route classification, registry-bounded action selection, plan-library expansion, and offline guarded rejection. The edge parser has 1.57M parameters and uses a shared ngram encoder, a route head, a route-conditioned action head, a route-action legality mask, and a capped prototype memory for long-tail actions. We build a product-grounded Chinese dataset of 12,383 validated commands from smart-home, notification, and mobile-robot scenes, split by source plan to avoid paraphrase leakage. On 1,887 test commands, the parser obtains 0.6529 route accuracy, 0.4187 action-label accuracy,and 1.0000 route-action consistency. Inside the full closed-loop pipeline, it reaches 1.0000 structurally valid-or-guarded rate,0.5707 operation-sequence accuracy, and 1.03 ms P95 latency in a CPU-only local edge-profile measurement. A blinded dualLLM audit on 200 random test cases gives 94.0% agreement (κ = 0.879) on semantic correctness and a conservative both positive rate of 42.0%. RegAct therefore provides millisecond response and an explicit registry boundary, while semantic intent resolution remains the principal limitation.
|
| |
| 10:35-16:40, Paper FrPP1.32 | |
| Artificial Intelligence Quantum Science - Theoretical Basis and Practice |
|
| Zheng, Kuifei | Eternal Space-Time Artificial Intelligence Technology (Beijing) Co., Ltd |
| Lyu, Chenglin | Eternal Space-Time Artificial Intelligence Technology (Beijing) Co., Ltd |
Keywords: Artificial Intelligence, Education Robotics, Machine Learning
Abstract: In recent years, artificial intelligence has developed rapidly. Taking this as a breakthrough, this paper constructs a strong artificial intelligence model framework driven by quantum chips, aiming to empower humanoid robots with strong artificial intelligence models and build basic capabilities for their autonomous progress. To enable strong artificial intelligence to have the ability of independent judgment and progress and better serve human beings, this paper creates five truth theories that can train the correctness of strong artificial intelligence's thinking and decision-making: Trend, Absolute Coordinate System Theory, Extreme Number Theory, Consciousness and Power, and Yi Quan, covering truth systems such as the evolution of historical laws, prediction models of things, the foundation of mathematical logic, sociology, and world economics, and establishing the logical origin of strong artificial intelligence. After that, this paper applies the five truth theories to humanoid robots equipped with artificial intelligence. Through model deduction, it develops humanoid robots that conform to future trends and possess the capacity to collect and control neutrinos. It is further discovered that neutrinos controlled by humanoid robots can act on and affect human cells. The construction of underlying logic targeting humanoid robots can drive the development of artificial intelligence toward strong artificial intelligence.
|
| |
| 10:35-16:40, Paper FrPP1.33 | |
| IKFL-PSO: IKFlow-Based Latent-Space PSO for Inverse Kinematics of Redundant Manipulators |
|
| Yang, Tianle | Institute of Nuclear & New Energy Technology, Tsinghua University |
| Zhou, Qin | Tsinghua University |
| Yi, Yuanlin | Hefei University of Technology School of Civil Engineering |
| Chen, Haolong | Hefei University of Technology School of Civil Engineering |
| Li, Zhijie | Tsinghua University Institute of Nuclear and New Energy Technology |
Keywords: Path and Motion Planning, Deep Learning
Abstract: Redundant manipulators can provide multiple inverse kinematics (IK) solutions for the same end-effector pose, but practical IK solutions should not only satisfy pose accuracy but also task requirements such as joint-space continuity and obstacle avoidance. IKFlow can efficiently generate diverse IK candidates through latent-variable sampling and mapping, but randomly sampled candidates may contain residual pose errors and fail to meet task-specific requirements. To address this problem, this paper proposes IKFL-PSO for inverse kinematics of redundant manipulators. IKFL-PSO iteratively updates latent particles in the pose-conditioned latent space of a pretrained IKFlow model. At each iteration, the particles are decoded into candidate joint configurations and evaluated according to pose error, joint-space continuity, and a hard obstacle-safety constraint. Simulations on a 9-DOF redundant manipulator show that IKFL-PSO can generate accurate and collision-free IK solutions while maintaining joint-space continuity without relying on dense random sampling. These results indicate that latent-space optimization provides a practical strategy for extending IKFlow from passive IK candidate generation to task-oriented IK optimization.
|
| |
| 10:35-16:40, Paper FrPP1.34 | |
| A Service-Readiness Method for Low-Cost Quadruped Robots with Voice Activation and Autonomous Viewpoint Seeking |
|
| Su, Liting | Fuzhou University Zhicheng College |
| Zhang, Xinhan | Fuzhou University Zhicheng College |
| Cao, Buyujie | Fuzhou University Zhicheng College |
| Li, Qingcheng | Fuzhou University Zhicheng College |
| Ma, Yunying | Fuzhou University Zhicheng College |
Keywords: Home and Personal Robot Systems, Human-Robot Interaction and Cooperation, Robot Vision and Computer Vision
Abstract: Before low-cost quadruped robots provide close-proximity services, they must identify speech-triggered scenes and adjust viewpoints for task startup. We present a scene-triggered service-readiness pipeline for a quadruped robot on a Raspberry Pi 5. The pipeline integrates voice activation, ready-pose specification, autonomous viewpoint seeking, visibility/centering/range/stability checks, and readiness-gated task attachment into a lightweight pre-task process. A rule-based human-orientation component is evaluated offline as a candidate side-view verifier. Using home sit-up assistance as validation, the robot establishes task-ready observation after a speech trigger. Physical experiments show that in six viewpoint-seeking trials, the robot reached readiness in 3.2–8.3s (5.8s avg). On 380 annotated non-supine images, the offline rule achieved 69.74% five-class accuracy, and 93.23% precision with 62.00% recall for side vs. non-side decisions. The sit-up counting module achieved 93.94% accuracy across 12 test-sample groups (165 repetitions), with event-level recall, precision, and F1 of 94.55%, 99.36%, and 96.89%. On Pi 5, pose-processing ran at 13–20 FPS with 32–70ms single-frame latency. Results indicate the system completes a lightweight local process from speech activation to task attachment in home-like settings, with offline evaluation showing potential for side-view verification.
|
| |
| 10:35-16:40, Paper FrPP1.35 | |
| Mel Statistical Calibration for Silent EMG-To-Speech across Multiple Temporal Encoders |
|
| Luo, Yuyang | Changchun University of Science and Technology |
| Wang, Yang | Changchun University of Science and Technology |
| Bai, Yu | Changchun University of Science and Technology |
| Gao, Yuchao | Changchun University of Science and Technology |
| Yang, Zijian | Dong Campus of Changchun University of Science and Technology |
Keywords: Human-Machine Interface, Brain-Machine Interface, Deep Learning
Abstract: 静音肌电图生成语音旨在合成语音 根据面部表面肌电图记录的 默口型,提高语音的清晰度。 现有系统通常预测中间Mel 然后用神经声码器来处理波形 但预测的Mel特征可能会表现出 相对于参考Mel的统计不匹配 分销。本文探讨了Mel统计学 校准(StatCal),一种轻量级推断时间 后处理方法,执行每梅尔箱均值和 方差比对后再按α控制 插值。StatCal 不要求重新培训 声学模型或声码器后端或ASR后端的修改。 在固定的HiFi-GAN语音编码器和DeepSpeech v0.9.3下 评估协议,S
|
| |
| 10:35-16:40, Paper FrPP1.36 | |
| A Study on the Identification of Motor Imagery and Resting States Based on Aperiodic EEG Features |
|
| Zhou, Qingshu | Yanshan University |
| Hua, XinYu | YanShan University |
| Sun, Bowen | Yanshan University |
| Lu, Houlin | Yanshan University |
| Li, Jiaxin | Yanshan University |
| Zhao, Jing | Yanshan University |
Keywords: Brain-Machine Interface, Machine Learning
Abstract: In asynchronous motor imagery brain-computer interfaces (MI-BCI), accurately distinguishing motor imagery from idle states is essential to avoid false triggers. However, most existing methods rely on subject-specific calibration and exhibit limited cross-subject generalizability, hindering clinical adoption. This study introduces aperiodic EEG features into motor imagery versus idle state classification. Irregular-Resampling Auto-Spectral Analysis (IRASA) was employed to decompose the EEG power spectrum into oscillatory and aperiodic components, and each trial was further divided into three 1-second time windows to capture temporal dynamics. Experiments on BCI Competition IV-2a demonstrate that aperiodic parameters—particularly the Offset—are significantly lower during motor imagery than during rest (P < 0.001), with this difference being directionally consistent across all participants. In leave-one-subject-out cross-validation, window-based aperiodic features achieved mean accuracies of 81% (left hand vs. idle) and 83% (right hand vs. idle), substantially outperforming CSP and FBCSP, while performance was comparable across methods in within-subject settings. Each feature dimension carries explicit neurophysiological meaning, offering a promising approach for building interpretable MI-BCI systems that require no individual calibration.
|
| |
| 10:35-16:40, Paper FrPP1.37 | |
| A Semantic Perception and Prediction-Driven MPPI Navigation Method for Dynamic Human-Robot Interaction Scenarios |
|
| Zhang, Muyuan | Nanjing University of Science and Technology |
| Wei, Chen | Army Engineering University of PLA |
| Wang, Manyi | Nanjing University of Science and Technology |
| Liu, Yuhang | Nanjing University of Science and Technology |
|
|
| |
| 10:35-16:40, Paper FrPP1.38 | |
| Lightweight Visual-SLAM-Guided Autonomous Area Coverage Robot System: Design and Validation for Cluttered Indoor Environments |
|
| Zhou, Huimin | Tsinghua University |
| Chi, Xi'ang | Tsinghua University |
| Dong, Ge | Tsinghua University |
Keywords: SLAM and Navigation, Path and Motion Planning, ROS, Software System for Robotics Application
Abstract: Achieving reliable 3D perception, autonomous navigation, and high coverage on low-cost embedded platforms remains challenging for coverage robots. This paper presents a lightweight coverage robot integrating depth-visual SLAM, ROS 2 navigation, and an improved Boustrophedon Cellular Decomposition (BCD) planner on a Raspberry Pi 5. Deploying BCD on a real visual-SLAM and Nav2 stack reveals execu- tion failures not present in simulation: noisy maps fragment decomposition, costmap inflation blocks waypoints, and sparse path vertices cause local-planner deviation. To address these failures, application-driven BCD adaptations are presented and evaluated against three sensor configurations (visual SLAM, 2D LiDAR, fusion) and three coverage schemes. Results show that the visual SLAM configuration achieves the best trade- off between localization accuracy, safety, and system cost. Compared with the conventional BCD implementation, the proposed adaptations reduce repetition by 18% (54.6% to 44.7%), tighten run-to-run consistency by a factor of 3.1 at negligible overhead, while the 2D LiDAR configuration collided in every trial due to an undetected low-profile obstacle.
|
| |
| 10:35-16:40, Paper FrPP1.39 | |
| Extrinsic Calibration of 2D LiDAR--IMU Systems with QCQP-SDR and Motion-Increment Pairing |
|
| Yang, Mengshen | Yunnan Minzu University |
| Jia, Fuhua | University of Nottingham, Ningbo, China |
| Hou, Xing | University of Nottingham Ningbo China |
| Zhao, Li | Yunnan Minzu University |
| Chen, Yunhao | Yunnan Minzu University |
| Yan, Yunhai | Yunnan Minzu University |
| Yang, Jingkai | Yunnan Minzu University |
| Tang, Jianing | Yunnan Minzu University |
Keywords: SLAM and Navigation, Mobile Robotics, ROS, Software System for Robotics Application
Abstract: Accurate extrinsic calibration between a 2D LiDAR and an IMU is important for planar localization and mapping, yet remains challenging in practice because the calibration objective is non-convex and the available motion estimates are often degraded by motion degeneracy, scan-matching failures, and inertial drift. This paper presents a targetless, egomotion-based calibration pipeline for 2D LiDAR--IMU systems. We derive a planar hand-eye-style formulation from per-sensor motion increments, cast the estimation problem as a quadratically constrained quadratic program (QCQP), and solve its semidefinite relaxation (SDR). To improve the quality of the optimization input, we combine GICP-based LiDAR odometry, IMU integration with zero-velocity updates and planar-motion constraints, and segment-level rejection of low-information or inconsistent motion windows. Simulations and real-platform experiments show that the proposed method achieves lower translation and yaw errors than an analytical baseline, and produces usable extrinsic estimates when the optimization is restricted to informative motion segments selected by the frontend checks.
|
| |
| 10:35-16:40, Paper FrPP1.40 | |
| Strain-Engineered Amorphous FeSiB Helical Microrobots with Multimodal Locomotion |
|
| Huang, Junhao | ShenZhen University |
| Ning, Haosen | Shenzhen University |
| Cheng, Mingxing | City University of Hong Kong |
| Zhou, Yuxuan | Shenzhen University |
| Liu, Juncheng | Shenzhen University |
| Liu, HaoXuan | ShenZhen University |
| Wu, Zongze | Shenzhen University |
| Hou, Chaojian | Shenzhen University |
Keywords: Micro/Nano Robotics
Abstract: Magnetic microrobots provide a promising route for wireless actuation, precise manipulation, and programmable locomotion at the micro/nanoscale. However, most existing systems still rely on conventional magnetic materials, leaving limited space for material-level regulation of magnetic response. Here, we introduce FeSiB amorphous soft magnetic alloy as a multi-element magnetic layer for artificial bacterial flagellumlike helical microrobots. The FeSiB film was integrated with a strain-engineered silicon nitride bilayer and an anisotropic micro-tooth array, enabling planar precursors to self-roll into three-dimensional helical architectures. Magnetic characterization showed thickness-dependent soft magnetic behavior, with the high-field magnetic response increasing by 4.17-fold from 30 to 90 nm while maintaining low coercivity. The final helix diameter and pitch were further regulated by precursor width, FeSiB thickness, and stress-layer configuration. Under a three-dimensional Helmholtz coil system, the microrobots exhibited transverse rolling, radial rolling, and rotating-forward motion, with mode transitions governed by field strength, actuation frequency, and robot geometry. These results establish FeSiB as a tunable magnetic material for programmable helical microrobots.These results show that amorphous soft magnetic alloys can expand the material choices for magnetic microrobots and provide a practical basis for codesigning magnetic materials and field-driven locomotion.
|
| |
| 10:35-16:40, Paper FrPP1.41 | |
| Rehab-EmotionTrack: Toward Patient-Role-Aware Speech Evidence Aggregation for Therapist-Facing Rehabilitation State Monitoring (I) |
|
| He, Xiaotong | Université Évry Paris-Saclay |
| Hu, Wanxin | Université Évry Paris-Saclay |
| Wang, Feilong | Université Évry Paris-Saclay |
| Schmirander, Yunus | Université Paris-Saclay |
| Su, Hang | Paris Saclay University |
| Dychus, Eric | Sandyc |
| Alfayad, Samer | Paris-Saclay Universit -Evry University |
Keywords: Rehabilitation and Assistive Robotics, Sensor Networks
Abstract: Emotion recognition is relevant to rehabilitation because patient speech can reveal effort, fear, pain, fatigue, or willingness to continue training. However, models trained on clean text or clean speech do not transfer directly to this setting: rehabilitation dialogue is multispeaker, patient speech can be quiet, discontinuous, or affected by environmental noise, and therapist prompts often contain the same safety-related words that a patient would use to report distress. This paper uses public clinical dialogue resources to evaluate failure modes relevant to therapist-facing rehabilitation state monitoring. We present method, a patient-role-aware prototype that decouples speaker diarization from role inference, aggregates target-patient utterances across fragmented speaker clusters, scores rehabilitation semantic cues, extracts valence--arousal and prosodic voice features, and applies text--voice fusion with conservative safety rules. The system outputs three labels: stateA, stateB, and stateC. In public stress tests, the fusion setting supports the feasibility of patient-role-aware speech evidence aggregation for therapist-facing rehabilitation monitoring. These results should be interpreted as an early robustness study, not as deployment-ready clinical validation; future work will test the approach in real rehabilitation sessions.
|
| |
| 10:35-16:40, Paper FrPP1.42 | |
| Smart-Insole Foot Trajectory Reconstruction During Stair Locomotion (I) |
|
| Wang, Feilong | Université Évry Paris-Saclay |
| Schmirander, Yunus | Université Paris-Saclay |
| Largeteau, Etienne | University D'Évry |
| Sleiman, Maya | KALYSTA |
| Qi, Wen | Politecnico Di Milano |
| Su, Hang | Paris Saclay University |
| Dychus, Eric | Sandyc |
| Alfayad, Samer | Paris-Saclay Universit -Evry University |
Keywords: Rehabilitation and Assistive Robotics
Abstract: Foot-mounted inertial sensing is widely used for wearable trajectory reconstruction when external positioning is unavailable, but stair ascent introduces load transfer, foot flexion, and support transitions that make inertial-threshold-based zero-velocity updates less reliable. This paper presents a smart-insole framework that uses plantar pressure and bending signals to aid zero-velocity update selection and a lightweight windowed factor graph optimizer to improve support-phase consistency. The pressure/flex module extracts mechanical support evidence from cycle-normalized plantar loading and low bending-change cues to confirm or supplement inertial zero-velocity candidates. The estimator jointly optimizes inertial propagation, zero-velocity factors, a soft non-holonomic constraint, and bias smoothness. Preliminary repeated-trial validation with one participant shows that the proposed method reduces support-phase drift and improves reconstruction smoothness over inertial-only Kalman filtering baselines, supporting a physical-consistency view of stair-ascent foot trajectory reconstruction.
|
| |
| 10:35-16:40, Paper FrPP1.43 | |
| BronchoSAS: A State-Aware Agent with Skills for Phantom-Based Bronchoscopy Training |
|
| Tian, Chuan | University of Southern Denmark |
| Yu, Hao | University of Edinburgh |
| Oliveira, Bruno | SDU |
| Jiang, Haolin | SDU Robotics, Maersk Mc-Kinney Moller Institute, University of Southern Denmark |
| Konge, Lars | Copenhagen Academy for Medical Education and Simulation |
| Cold, Kristoffer | Copenhagen Academy for Medical Education and Simulation |
| Cheng, Zhuoqi | University of Southern Denmark |
Keywords: Education Robotics, Medical Robotics, Artificial Intelligence
Abstract: This paper presents BronchoSAS, a state-aware AI assistant designed to support novice learners during phantom-based bronchoscopy training. The system converts airway navigation state into real-time spoken guidance, landmark teaching, task-focused question answering, recovery support, and post-session formative feedback. Its hybrid single-agent architecture combines deterministic curriculum and priority rules with selective LLM support for eligible skill arbitration and response realisation. A preliminary formative evaluation with five novices and one domain expert found that the prototype was generally perceived as relevant and understandable, while response timing, conversational continuity and teaching usefulness remain priorities for improvement.
|
| |
| 10:35-16:40, Paper FrPP1.44 | |
| A Multi-Agent Service Robot for Smart Event Assistance |
|
| K. Alshammari, Reem | King Abdulaziz City for Science and Technology |
| G. Alanazi, Abdullah | Artificial Intelligence Association, |
| K. Alshammari, Razan | King Saud University |
Keywords: Human-Robot Interaction and Cooperation, Artificial Intelligence, Cognitive Robotics
Abstract: Conference and event venues frequently encounter operational challenges, including visitor disorientation, overloaded information desks, language barriers, and the need to support numerous simultaneous attendee requests. While large language models have enhanced conversational capabilities, monolithic assistants remain difficult to maintain, extend, and scale as event services evolve. This paper presents **Siraj**, a modular multi-agent robotic assistant that combines a service robot with specialized AI agents for navigation, scheduling, event information, and networking support. A routing mechanism dynamically dispatches user requests to the most appropriate agent, enabling modular updates, efficient task specialization, and scalable support across both robot-based and digital interaction channels. The system was evaluated through scenario-based testing during a three-day AI Agents Hackathon in Saudi Arabia. Experimental results show that the routing framework correctly assigned 97.4% of user requests to the appropriate agent, while enforcing strict grounding in event-specific knowledge improved task success from 73.7% to 89.5%. These results demonstrate that the proposed multi-agent architecture provides accurate, context-aware event assistance while improving system modularity, maintainability, and scalability for smart event environments.
|
| |
| 10:35-16:40, Paper FrPP1.45 | |
| A Domain-Structured Adversarial Alignment Model for EEG Emotion Recognition Based on Source-Domain Grouping |
|
| Chang, Jiang | Shanxi University |
| Guo, Xun | Shanxi University |
| Zhang, Dingyu | Shanxi University |
| Wang, Fang | Chengdu Technological University |
Keywords: Brain-Machine Interface, Artificial Intelligence, Deep Learning
Abstract: This paper presents domain-structured adversarial alignment (DSAA), a source-only model for cross-subject EEG emotion recognition. DSAA measures pairwise source-domain discrepancies using maximum mean discrepancy (MMD) and partitions the source subjects into relatively homogeneous and heterogeneous groups. During pre-training, the homogeneous group supports supervised reconstruction, whereas the heterogeneous group is used for subject-adversarial learning. The encoder and emotion classifier are then fine-tuned using labeled source-subject data. The held-out target subject is excluded from gradient-based training and used for evaluation. DSAA obtains average accuracies of 93.10%, 78.92%, and 80.69% on SEED, SEED-IV, and SEED-V, respectively. These results indicate that differentiated use of source domains can support cross-subject recognition without using target-subject samples for parameter updates.
|
| |
| 10:35-16:40, Paper FrPP1.46 | |
| A Two-Stage Sound Source Localization Method Based on Robust TDOA and Dynamic Space-Constrained StGCF |
|
| Pan, Ci | Henan Polytechnic University |
| Kan, Yue | Henan Polytechnic University |
| Zha, Fusheng | Harbin Institute of Technology |
Keywords: Sensor Networks, Human-Robot Interaction and Cooperation, Intelligent Control and Systems
Abstract: To address the degradation of sound source localization performance caused by reverberation, multipath propagation, and background noise in complex indoor environments, as well as the high computational cost of full-space searching in the conventional Spatio-Temporal Global Coherence Field (stGCF) algorithm, this paper proposes a two-stage localization method based on Robust TDOA-Guided Dynamic Space-Constrained stGCF (RTDS-stGCF). First, multi-channel time delay information is extracted using Generalized Cross Correlation with Phase Transform (GCC-PHAT), while dynamic reference microphone selection and Median Absolute Deviation (MAD) filtering are employed to suppress outlier TDOA observations. Subsequently, high-confidence observations are selected according to the cross-correlation peak magnitude to obtain a coarse source location estimate. Based on the coarse localization result, a dynamically constrained search region is constructed, and high-resolution stGCF searching is performed within the local region to achieve accurate source localization. Experimental results demonstrate that the proposed method effectively reduces the search space and computational burden, mitigates the performance degradation caused by redundant searching, and improves localization accuracy and robustness in complex acoustic environments. The proposed framework provides an effective solution for sound source localization in reverberant and noisy scenarios.
|
| |
| 10:35-16:40, Paper FrPP1.47 | |
| Velocity-Conditioned Implicit Phase Learning for Style-Aware Humanoid Locomotion with Adversarial Motion Priors |
|
| Wang, Ningejie | Zhejiang University |
| Huang, Weidi | Zhejiang University |
| Lu, GuoDong | Zhejiang University |
| Gan, Chunbiao | Zhejiang University |
Keywords: Humanoid Robots, Machine Learning
Abstract: Achieving natural humanoid locomotion remains challenging. Model-based methods require extensive manual tuning, while pure reinforcement learning produces unnatural gaits. This paper proposes a sim-to-real framework integrating Adversarial Motion Priors (AMP) with implicit phase encoding for style-aware bipedal walking on a 29-DoF robot T1PRO. A velocity-driven multi-motion blending sampler enables seamless gait transitions through conditional kernel-density weighting. A General Motion Retargeting pipeline transfers heterogeneous human motion-capture data to the robot morphology. Symmetric data augmentation mitigates dataset asymmetry. Trained with PPO and an LSGAN discriminator in IsaacSim, the policy demonstrates accurate velocity tracking, natural gait styles, and improved robustness against command perturbations. Results validate reduced manual reward shaping and enhanced motion naturalness compared to baselines.
|
| |
| 10:35-16:40, Paper FrPP1.48 | |
| Class-Selective Wavelet-Decoupled Heterogeneous Feature Knowledge Distillation for Aircraft Skin Small Object Defect Detection |
|
| Zheng, ChuanYao | Nanjing University of Aeronautics and Astronautics |
| Wang, Congqing | Nanjing University of Aeronautics and Astronautics |
| Gou, Jiawei | Nanjing University of Aeronautics and Astronautics |
Keywords: Deep Learning, Robot Vision and Computer Vision
Abstract: Precise aircraft skin defect detection is critical for ensuring aviation safety. To improve lightweight model accuracy for irregular, small-object defects amid background clutter, this paper proposes a multi-scale frequency-domain knowledge distillation framework. First, a compact student network is established by introducing a Dual-Domain Wavelet Downsampling (DWD) module to preserve high-frequency structural boundaries during downsampling. To resolve operator-induced semantic fractures between non-isomorphic networks resulting from this architectural divergence, a Class-Selective Wavelet-Decoupled Heterogeneous Feature Distillation (CWKD) module is developed. CWKD leverages layer adapters and 2D-DWT to decouple feature maps into frequency sub-bands, eliminating representation gaps under zero inference overhead. Furthermore, a Joint Weight Map Generation Mechanism (JWGM) is designed to suppress background interference and balance long-tailed category distributions by compounding a spatial mask, teacher high-frequency saliency, and class-selective weights. Experimental results show that the proposed DWD module achieves an 81.8% parameter compression (reducing parameters from 2.59M to 0.47M) while improving baseline mAP@0.5 to 85.9%. Furthermore, the distillation framework provides an additional accuracy gain of +0.6% mAP under zero inference overhead, reaching a peak precision of 86.5% at 140.8 FPS.
|
| |
| 10:35-16:40, Paper FrPP1.49 | |
| Mutual Information Guided Deep Feature Extraction Method for Driver Emotion Recognition |
|
| Hu, Fenghua | Xi'an Jiaotong University |
| Xiaodong, Zhang | Xian Jiaotong University |
| Zhao, Zirui | Xi'an Jiaotong University |
| Pan, Chengkun | Xi'an Jiaotong University |
| Zhang, Xueteng | Xi’an Jiaotong University |
| Li, Rui | Xi'an Jiaotong University |
Keywords: Intelligent Transportation Systems, Robot Vision and Computer Vision
Abstract: Driver emotions affect driving safety, making driver emotion recognition crucial. Deep learning has advanced facial expression recognition, however most methods directly use high dimensional pretrained features that contain redundancy and task-irrelevant components, which limits generalization, especially with scarce driving data. To address this issue, we propose a mutual information guided deep feature extraction method in this paper. Specifically, 1280-dimensional deep facial features are extracted using a pretrained lightweight convolutional neural network. A low-variance filtering step is applied to remove near-constant dimensions. Then, a mutual-information-guided feature extraction pipeline is introduced, including permutation-based significance testing, correlation-based redundancy removal, and MI–random forest joint scoring. This process produces a compact 100-dimensional feature representation. The CK+ dataset is used to identify discriminative feature indices, which are subsequently evaluated on a self-collected driver emotion dataset. Experimental results demonstrate that the proposed compact representation maintains comparable recognition performance with the original high-dimensional features in frame-level and temporal recognition tasks while reducing feature dimensionality by more than 90%.
|
| |
| 10:35-16:40, Paper FrPP1.50 | |
| Grasping Detection Network Based on Regional Multi-Scale Features and Physical Prior Knowledge |
|
| Liu, Jiajun | Shandong University |
| Tian, Xincheng | Shandong University |
| Dai, Xiaomeng | China Railway Construction Corporation Bridge Engineering Bureau Group Construction Assemble Technology Corporation |
| Zhou, Lelai | Shandong University |
Keywords: Deep Learning, Grasping and Manipulation
Abstract: Achieving reliable and efficient grasping in clutter remains a significant challenge in robotics. Existing approaches often focus solely on modeling grasp within the gripper scale, while neglecting the direct influence of environmental context on grasp quality. Moreover, they face difficulties in integrating physical priors to enable robust and antipodal grasps. To address these issues, we propose a novel Grasping Detection Network Based on Regional Multi-scale Features and Physical Prior Knowledge. Specifically, our method employs a farthest grid sampling strategy in conjunction with grasp heatmaps to ensure the diversity and effectiveness of candidate grasps. A Multi-scale Feature-based Grasp Generator is then utilized to produce high-quality 6-DoF grasp poses. Furthermore, we incorporate a physics-inspired prior regularization term into the loss design to enhance both generalization and stability. Experimental evaluations on the GraspNet-1Billion dataset demonstrate that our method achieves superior grasp accuracy and inference efficiency compared with SOTA baselines. Additionally, real-world experiments conducted on Franka Panda demonstrate the effectiveness of the proposed methods with a high grasp success rate.
|
| |
| 10:35-16:40, Paper FrPP1.51 | |
| Designing Task-Induced Arousal: A Multimodal Stress Induction Method for Interactive Experiments |
|
| Frederiksen, Morten Roed | IT-University of Copenhagen |
Keywords: Human-Robot Interaction and Cooperation, Rehabilitation and Assistive Robotics
Abstract: HCI and HRI studies often require short, repeatable arousal manipulations that can run while participants continue interacting with a device or robot. These experiments are often challenged by the need to induce arousal in settings that still resemble real interaction. Participants must continue using a device, touching a robot, or producing sensor data while the manipulation unfolds. We aimed to develop a compact and repeatable way to induce controlled task-related arousal during interactive experiments by combining a lightweight browser-based pacing task, escalating timing demands and urgency cues, and a concurrent physical hotwire-style challenge. In an A-B-A within-participant study, the induction condition significantly increased mental demand, temporal demand, effort, frustration, and SAM arousal (all p < .001), while perceived performance decreased (p < .001). GSR peak rate increased relative to both calm conditions (p = .030) and escalated over time (p < .001). Grip variability also increased (p = .036), as did release speed (both p < .001), while valence remained above the scale midpoint. These results provide initial evidence that the combined procedure induces controlled, relatively high-valence task-related arousal and may serve as a reusable experimental tool for future human-computer and human-robot interaction studies.
|
| |
| 10:35-16:40, Paper FrPP1.52 | |
| An Intelligent Multimodal Adaptive Disinfection Robot System for Public Certificates Based on Sensor Fusion and Closed-Loop Control |
|
| Zhang, Hexin | Changchun University of Science and Technology |
| Bai, Xuemei | Changchun University of Science and Technology |
| Xu, Tiantian | Company-Changchun Shuangzi Measurement and Control Technology Co., Ltd |
Keywords: Deep Learning, Artificial Intelligence
Abstract: To address the limitations of traditional certificate disinfection systems, including fixed operation procedures, poor environmental adaptability, and lack of sensor-feedback control mechanisms, This paper presents a sensor-driven multimodal adaptive disinfection robot system for public certificate sterilization scenarios.The proposed system integrates ultraviolet (UV) irradiation, ultrasonic atomization, and hot-air drying modules under a unified embedded robotic architecture. A multi-source sensing and input mechanism is introduced to construct the operating state, including temperature, humidity, liquid-level state, UV sensing signal, and user-selected certificate-size mode. Based on this state representation, a rule-based Multi-modal Adaptive Disinfection Policy (MADP) is designed to adjust disinfection parameters, including UV exposure time, atomization duration, and hot-air drying duration. The system is deployed on an STM32 embedded platform to realize real-time control and closed-loop feedback execution. Experimental results demonstrate that the proposed system achieves sterilization efficiency above 99.95% under different operating conditions while maintaining stable performance across varying certificate sizes.Material compatibility tests are conducted to evaluate the influence of repeated disinfection processes on certificate materials. The proposed method provides a lightweight, deployable, and sensor-feedback-driven robotic solution for public safety and automate
|
| |
| 10:35-16:40, Paper FrPP1.53 | |
| GNN-Based Joint Beamforming Design for Dual-RIS-Assisted ISAC Systems |
|
| Yin, Rongbin | Changchun University of Science and Technology |
| Bai, Xuemei | Changchun University of Science and Technology |
| Geng, Xiaofei | School of Electronic and Information Engineering, Changchun University of Science and Technology |
| Xu, Tiantian | Company-Changchun Shuangzi Measurement and Control Technology Co., Ltd |
Keywords: Deep Learning, Machine Learning, Sensor Networks
Abstract: The integrated sensing and communication (ISAC) system merged with reconfigurable intelligent surface (RIS) has recently received much attention. This paper proposes an innovative graph neural network (GNN)-based framework for beamforming optimization in ISAC systems with RIS. The ISAC system is designed to optimize signal transmission through the RIS by maximizing the worst-case target illumination power to increase overall performance. In the proposed GNN-based framework, we model the ISAC network as a heterogeneous graph. This enables the GNN to capture complex interactions among RIS nodes, communication user (CU) nodes, and sensing target nodes, thereby establishing an effective mapping from the channel state information (CSI) to the DFBS beamforming and dual-RIS phase shift vectors. Numerical results confirm that the proposed GNN-based framework outperforms existing baseline approaches in terms of computational efficiency, implementation feasibility, and generalization capability.
|
| |
| 10:35-16:40, Paper FrPP1.54 | |
| Multi-Sensor Fusion-Based Defect Localization for In-Pipe Robots |
|
| Ge, Jiaxin | Beijing University of Chemical Technology |
| Li, Guozheng | Beijing University of Chemical Technology |
| Li, Zhiqing | Beijing University of Chemical Technology |
|
|
| |
| 10:35-16:40, Paper FrPP1.55 | |
| Mamba-Transformer Hybrid Imitation Learning for Robotic Assembly Tasks |
|
| Zhang, Deyao | Beijing University of Chemical Technology |
| Ji, Yun | Chn Energy I&c Technology Co., Ltd |
| Li, Zhiqing | Beijing University of Chemical Technology |
Keywords: Machine Learning, Grasping and Manipulation, Industrial Robotics and Factory Automation
Abstract: Fine robotic assembly requires stable long-horizon action generation and stage-adaptive execution under tight geometric constraints. To address the insufficient modeling of long-horizon actions and the limited adaptability of fixed action chunks across different assembly stages, this paper proposes a Mamba-Transformer hybrid imitation learning method for assembly tasks. The proposed method adopts ACT as the baseline framework and incorporates Mamba-3 modules to enhance continuous action-sequence modeling. An adaptive action chunking mechanism based on prediction uncertainty is further designed to dynamically adjust the execution horizon according to action prediction variance and temporal prediction inconsistency. In simulation, dual-view images and end-effector pose data are collected through gamepad-based teleoperation, and the method is evaluated on square-slot and T-slot insertion tasks. Experimental results show that the proposed method achieves success rates of 84% and 72% on the square-slot and T-slot tasks, respectively, improving ACT by 14 and 8 percentage points. Ablation studies indicate that Mamba-3 performs best when jointly incorporated into multiple temporal modeling modules, and that the adaptive action chunking mechanism further improves stage-wise adaptability during assembly.
|
| |
| 10:35-16:40, Paper FrPP1.56 | |
| Design, Performance Analysis, and Experimental Verification of a Six Wheel Variable Diameter Pipeline Decontamination Robot |
|
| Li, Guiyao | Beijing University of Chemical Technology |
| Li, Zhiqing | Beijing University of Chemical Technology |
|
|
| |
| 10:35-16:40, Paper FrPP1.57 | |
| Uncertainty-Guided Adaptive Iterative Refinement for Scene Text Recognition |
|
| Li, Zhennan | Beijing University of Chemical Technology |
| Fan, Jing | Shandong Provincial Water Network Operation and Dispatching Center |
| Li, Zhiqing | Beijing University of Chemical Technology |
Keywords: Robot Vision and Computer Vision, Deep Learning
Abstract: 文本识别(STR)逐步从纯视觉建模到联合视觉-语言建模框架。ABINet通过以下方式实现竞争性能结合显式语言模型和迭代 精炼(IR)机制。然而,它的推广 在现实世界数据集上的能力仍然受限, 尤其是在复杂且噪声较大且鲁棒的场景中 语义建模至关重要。在最初的 ABINet 中,语言模块被实现 使用双向卷积网络(BCN)。它 主要建模本地上下文,并从零开始训练。因此,它捕捉远程依赖关系的能力 全局语义有限,这可能影响错误 在复杂的
|
| |
| 10:35-16:40, Paper FrPP1.58 | |
| Research on Dynamic Pressure Recognition Based on Flexible Sensor Arrays |
|
| Yang, Zijian | Dong Campus of Changchun University of Science and Technology |
| Bai, Yu | Changchun University of Science and Technology |
| Yang, Jia | Jilin Jianzhu University |
| Wang, Yang | Changchun University of Science and Technology |
| Luo, Yuyang | Changchun University of Science and Technology |
| Gao, Yuchao | Changchun University of Science and Technology |
Keywords: Machine Learning, Deep Learning
Abstract: ,动态压力识别在柔性传感器阵列中是 由于非线性、时间变化和 接触信号对噪声敏感性。这篇论文 提议 动态时空特征融合框架 压力状态识别。连续测量为 以顺序压力热图表示,以及压力 变异趋势图通过帧间构建 差异 以捕捉局部变化的方向和大小。A 3D 带有残差连接的卷积神经网络是 用于提取层级时空特征 从压力序列、空间模式的整合和 时间动力学用于区分压力状态(增加, 稳定,减弱)。 在自建数据集上的实验表明,该方
|
| |
| 10:35-16:40, Paper FrPP1.59 | |
| Enhancing Hindsight Experience Replay with Hindsight Sample Replacement |
|
| Li, Ruixiang | Beijing University of Chemical Technology |
| Li, Zhiqing | Beijing University of Chemical Technology |
Keywords: Machine Learning, Deep Learning, Grasping and Manipulation
Abstract: The Non-Negative Sparse Reward (NNSR) problem poses a critical challenge in applying Hindsight Experience Replay (HER) to sequential operation tasks, where the agent's limited physical interactions with environmental objects often yield non-negative target rewards—corresponding to uninformative samples that impede efficient learning. Dense rewards can quickly guide the agent to complete tasks, but designing an optimal reward function that fits the task is difficult. Combining dense rewards with HER technology, this paper proposes a reinforcement learning scheme that incorporates a hindsight sample replacement mechanism, which significantly alleviates the NNSR problem with a simple dense reward function design. Additionally, this paper augments this scheme with an Energy-Based HER prioritization strategy to further address the NNSR challenge. Empirical evaluations on robotic arm sequential tasks show that the proposed scheme achieves superior training efficiency over HER baselines, while maintaining comparable final performance.
|
| |
| 10:35-16:40, Paper FrPP1.60 | |
| A MobileNetV3-CBAM-Based Multi-Class Gait Recognition Model for Smart Insoles |
|
| Gao, Yuchao | Changchun University of Science and Technology |
| Wang, Yang | Changchun University of Science and Technology |
| Yang, Jia | Jilin Jianzhu University |
| Luo, Yuyang | Changchun University of Science and Technology |
| Yang, Zijian | Dong Campus of Changchun University of Science and Technology |
| Bai, Yu | Changchun University of Science and Technology |
Keywords: Sensor Networks, Machine Learning, Deep Learning
Abstract: Smart insoles play an important role in portable motion health monitoring and abnormal gait early warning. For edge wearable devices with limited computing power and strict real-time constraints, direct spatiotemporal feature modeling on raw multi-channel sensor data is of great academic and engineering value for low-power, high-precision embedded gait recognition. This paper proposes a lightweight MobileNetV3 gait recognition model integrated with the Convolutional Block Attention Module (CBAM). The 36-channel raw plantar pressure signals from bilateral insoles are arranged along the time dimension to form a 36×T two-dimensional spatiotemporal feature matrix. MobileNetV3 with depthwise separable convolution efficiently extracts spatiotemporal coupling features of plantar pressure, and CBAM adaptively recalibrates feature weights in both channel and spatial dimensions to enhance key gait information representation and suppress redundant interference. Experiments are conducted on a dataset covering six typical motion states: standing, normal walking, tiptoe walking, fatigued walking, left ankle sprain gait and right ankle sprain gait. Results show that the proposed MobileNetV3-CBAM achieves high classification accuracy with small parameters and low computational overhead, which satisfies edge device deployment constraints and can effectively distinguish normal gait from multiple abnormal patterns. This method strikes a sound balance between computatio
|
| |
| 10:35-16:40, Paper FrPP1.61 | |
| PAS-Net: A Framework for Robust Mental Strain Detection (I) |
|
| Chen, Qinhao | Jinan University |
| Wang, Yueming | South China University of Technology |
| Wang, Yunhan | The Chinese University of Hong Kong |
| Li, Wenjie | Beihang University |
| Ma, Shang | South China University of Technology |
| Qi, Wen | Politecnico Di Milano |
Keywords: Human-Robot Interaction and Cooperation, Sensing, Haptic System, Machine Learning
Abstract: Affective recognition exhibits broad application prospects in human–computer interaction, mental health monitoring, driving safety, and many other fields. Modern wearable devices integrate multimodal physiological sensors, facilitating continuous acquisition of physiological signals that support affective-state analysis under real-world conditions. However, practical translation faces two important challenges: cross-subject physiological variability degrades generalization, while coarse phase-level labels fail to represent rapid physiologi- cal fluctuations. We propose PAS-Net, a framework comprising three Multi-Layer Perceptrons (MLPs) for experimental-condition classification, arousal–valence proxy regression, and mental strain detection, respectively. Because per-window strain annotations are unavailable, an LLP-inspired proportion- constrained procedure is used to generate pseudo-labels, together with condition-dependent thresholds for inference. Experiments on five held-out subjects demonstrate a mean proxylabel discrimination ROC-AUC of 0.819 and preserve several intended condition-level trends.
|
| |
| 10:35-16:40, Paper FrPP1.62 | |
| An Infrared-Guided Retinex Image Fusion Method for Low-Light Robot Vision |
|
| Zhou, Hanqun | Xi’an Technological University |
| Bi, Yang | XIHANG UNIVERSITY |
| Li, Jiguang | XIHANG UNIVERSITY |
Keywords: Robot Vision and Computer Vision, Deep Learning
Abstract: For nighttime robotic vision, visible images offer rich scene details but severe illumination degradation, whereas infrared images highlight thermal targets yet lack texture. To address illumination estimation bias, imbalanced bimodal fusion, and blurred structural details in existing Retinex‑based methods, we propose IRR‑Fusion, an infrared‑guided Retinex network. It decomposes visible images into reflectance and illumination layers, uses infrared structural responses to adaptively compensate illumination inaccuracies, and dynamically integrates visible textures with infrared targets via an attention‑based reflectance fusion module, along with a gradient consistency constraint to preserve salient edges. On 50 LLVIP pairs, IRR‑Fusion achieves the best standard deviation and visual information fidelity, ranks second in entropy and spatial frequency among five methods, and uses only 0.507M parameters and 0.0427 s per pair, balancing fusion quality and efficiency. Evaluation on 50 unseen MSRS pairs confirms cross‑dataset generalization. Moreover, YOLO11s pedestrian detection on the LLVIP test set shows that fused images improve mAP@0.5 and mAP@0.5:0.95 over low‑light visible images by 5.25 and 9.82 percentage points, respectively. These results demonstrate the effectiveness, generalization, and downstream perception value of IRR‑Fusion for low‑light robotic vision.
|
| |
| 10:35-16:40, Paper FrPP1.63 | |
| Zero-Shot Decision Making for Mobile Manipulation Robots1nOpen-Scenario Autonomous Tidying |
|
| Xiao, Shuzhen | Beihang University |
| Tao, Yong | Beijing University of Aeronautics and Astronautics |
| Yian, Song | BUAA |
| Ren, Fan | Nankai University |
Keywords: Mobile Robotics, Artificial Intelligence, Grasping and Manipulation
Abstract: This paper proposes a zero-shot task planning pipeline for mobile manipulators. It is based on active visual observation and large language model (LLM) decision-making. The goal is to achieve unsupervised, long-horizon autonomous tidying through an "observation-driven" paradigm. First, we design a multi-station active inspection mechanism. The SAM3 algorithm is introduced to perform zero-shot instance segmentation and 3D spatial anchoring on unknown cluttered objects. Second, local zero-shot visual features are directly converted into lightweight, structured scene state logs. This eliminates the computing power consumption of high-frequency global state maintenance. Finally, the LLM uses its common sense reasoning to logically cluster unknown objects based on these serialized logs. This open-vocabulary clustering requires no predefined rules. It closed-loop generates a global long-horizon action flow of "active inspection, dynamic classification, and cross-station placement". Full-process physical experiments were conducted in a real scene with four cluttered workstations. Results show the system accurately completes autonomous classification and cross-region tidying of various unknown objects. It relies solely on real-time visual observations, requiring no task-specific object training, predefined object categories, or object-level 3D models. This research significantly improves the operational flexibility and robustness of the system in highly dynamic, unstructured scenario
|
| |
| 10:35-16:40, Paper FrPP1.64 | |
| OSDaR-IR-MOT: Workstation-Grade Trident Infrared Detection and Multi-Object Tracking for Rail Intrusion Monitoring |
|
| Zhang, Xuyang | Changchun University of Science and Technology |
| Hu, Haoqi | Changchun University of Science and Technology |
| Gao, Yuchao | Changchun University of Science and Technology |
| Yang, Haoran | Changchun University of Science and Technology |
| Zheng, Zegang | Xi'an Jiaotong University |
|
|
| |
| 10:35-16:40, Paper FrPP1.65 | |
| An Improved YOLOv11-Based Algorithm for Infrared Small Target Detection |
|
| Hu, Haoqi | Changchun University of Science and Technology |
| Zhang, Xuyang | Changchun University of Science and Technology |
| Zheng, Zegang | Xi'an Jiaotong University |
| Zhu, Mingshuo | Changchun University of Science and Technology |
| Wang, Jiepeng | Changchun University of Science and Technology |
| Wang, Jinchi | Changchun University of Science and Technology |
| Han, Kexu | Changchun University of Science and Technology |
Keywords: Robot Vision and Computer Vision, Intelligent Transportation Systems, Field Robotics
Abstract: With the widespread adoption of unmanned aerial vehicles (UAVs), anti-UAV technologies have become essential for security surveillance and airspace management. While thermal infrared vision enables continuous monitoring, exist ing detection methods consistently struggle with low thermal contrast, dramatic scale variations, and extreme positional sen sitivity when targets occupy only a few pixels. Standard deep learning detectors often suffer from high false-negative rates and regression instability when faced with complex background clutter and morphologically unstable drone signatures. To address these limitations, this paper proposes an im proved small object-aware detection framework built upon the YOLOv11 architecture. We integrate a Dynamic Detection Head to adaptively fuse multi-level scale, spatial, and task aware features, improving recall for tiny drone targets. Fur thermore, we devise a joint loss function combining Normalized Wasserstein Distance (NWD) with dynamic shape-aware Scale and-Shape (SSQ) penalties. By modeling bounding boxes as 2D Gaussian distributions, this loss provides smooth gradient supervision, mitigating the extreme positional sensitivity of traditional Intersection over Union (IoU) metrics in sub-pixel regimes. Comprehensive evaluations on the Anti-UAV410 thermal benchmark demonstrate the effectiveness of our framework. At 640 × 640 input, the proposed detector reaches AP50 of 88.2%, AP50-95 of 49.1%, and Recall@0.5 of 91.5%, outper forming YOLO
|
| |