| | |
Last updated on September 13, 2026. This conference program is tentative and subject to change
Technical Program for Friday October 30, 2026
| |
| FrAT1 |
Palace Hall East, 3F |
| Control Theory and Applications 3 |
Oral Session |
| |
| 09:00-09:15, Paper FrAT1.1 | |
| Prescribed-Time Consensus for Uncertain Multi-Agent System Via Volterra-Operator Approach |
|
| Yang, Shun | Huazhong University of Science and Technology |
| Xiao, Qiang | Huazhong University of Science and Technology |
Keywords: Control Theory and Applications, Information and Networking, Autonomous Vehicle Systems
Abstract: The problem of prescribed-time consensus tracking is studied in this paper for second-order nonlinear multi-agent systems in the presence of parametric uncertainties, external disturbances, and time-varying delays. To address the challenges arising from unknown dynamics and delayed information, the distributed control task is decomposed into two tightly coupled components: parameter estimation and tracking control. A novel Volterra-operator-based estimation scheme is developed, which enables prescribed-time convergence of unknown parameters even in the presence of disturbances, thereby overcoming the limitations of conventional asymptotic or finite-time estimation methods. On this basis, an adaptive sliding-mode control law is constructed to ensure that the system realizes prescribed-time consensus tracking. By constructing an appropriate Lyapunov functional and analyzing the delay-dependent terms, the consensus tracking errors are rigorously proven to be prescribed-time stable and the boundedness of control inputs is ensured. Simulation is conducted in a two-dimensional scenario and further validates the effectiveness, robustness, and prescribed-time convergence capabilities of the presented scheme.
|
| |
| 09:15-09:30, Paper FrAT1.2 | |
| Cosserat Rod Theory-Based Graph Neural Networks for Winding Manipulation of Deformable Linear Objects |
|
| Lee, Giwan | Chonnam National University |
| Hong, Ayoung | Chonnam National University |
Keywords: Control Theory and Applications, Navigation, Guidance and Control
Abstract: Robotic manipulation of deformable linear objects (DLOs), such as winding cables around obstacles, is highly challenging due to their infinite degrees of freedom and complex external interactions. Traditional physics-based modeling is computationally intensive, while data-driven methods struggle with physical consistency and generalization. To address these limitations, this paper proposes a Cosserat rod-based Graph Neural Network (C-GNN) framework for predicting DLO shapes and external contact forces during winding tasks. By embedding a physics-informed loss function derived from Cosserat rod theory, the network enforces mechanical constraints, thereby enhancing prediction accuracy and generalization. We integrate the C-GNN into a Model Predictive Control (MPC) framework to optimize robotic winding operations. Simulation experiments in PyElastica demonstrate that the proposed C-GNN successfully estimates DLO configurations and contact forces, achieving a node position RMSE of 7.04 mm. Under MPC control, the C-GNN controller outperforms a pure GNN by reducing geometric deviation to 0.10 mm and significantly suppressing average contact forces. The framework exhibits extrapolated condition robustness in winding scenarios, ensuring stable manipulation with reduced risk of excessive deformation.
|
| |
| 09:30-09:45, Paper FrAT1.3 | |
| Probabilistic Forced Backward Reachable Sets for Stochastic Avoid Problems |
|
| Solanki, Prashant | Delft University of Technology |
| van Beers, Jasper | Delft University of Technology |
| de Visser, Coen | TU Delft |
Keywords: Control Theory and Applications, Navigation, Guidance and Control, Autonomous Vehicle Systems
Abstract: In safety-critical systems, reachability provides a principled way to certify unsafe operating regions. Deterministic and robust forced backward reachable sets characterize states from which a system is driven into an avoid set regardless of the control action. This paper introduces a stochastic counterpart, the probabilistic forced backward reachable set, defined as the set of states from which every admissible Markov policy reaches the avoid set within a finite horizon with probability above a prescribed threshold. We develop a dynamic-programming characterization and prove measurability and monotonicity properties. For finite state and control spaces, the proposed method computes the set exactly. For continuous state and control spaces, we construct a finite abstraction and derive an error bound showing convergence of the approximate value function as the grid size tends to zero. The resulting threshold-shifted abstraction provides a certified inner approximation of the probabilistic forced backward reachable set.
|
| |
| 09:45-10:00, Paper FrAT1.4 | |
| Discrete-Time D-Type and PD-Type ILC: An LMI-Based Design Approach |
|
| Mochizuki, Kosei | Shibaura Institute of Technology |
| Zhai, Guisheng | Shibaura Institute of Techbology |
Keywords: Control Theory and Applications, Navigation, Guidance and Control, Industrial Applications of Control
Abstract: Iterative Learning Control (ILC) is a highly effective control strategy for improving the tracking performance of systems that execute repetitive tasks. While continuous-time ILC has been extensively studied, the development of discrete-time ILC is of great practical significance due to the widespread implementation of digital controllers. This paper investigates discrete-time D-type and PD-type ILC schemes. In discrete-time frameworks, the difference of the tracking error is analyzed as a comparison to the continuous-time derivative operation. The proposed learning laws generate the control input for the upcoming iteration based solely on the control input and the tracking error from the previous iteration. To guarantee the monotonic convergence of the tracking error, sufficient design conditions are mathematically derived and formulated in terms of Linear Matrix Inequalities (LMIs). Finally, numerical simulations using MATLAB are provided to demonstrate the effectiveness, convergence, and properties of the proposed discrete-time D-type and PD-type ILC approaches.
|
| |
| 10:00-10:15, Paper FrAT1.5 | |
| Robust Controller Design for a Hybrid Dynamic Vibration Absorber with Limited Information |
|
| Goto, Junta | Shibaura Institute of Technology |
| Zhai, Guisheng | Shibaura Institute of Techbology |
Keywords: Control Theory and Applications, Navigation, Guidance and Control, Industrial Applications of Control
Abstract: Passive dynamic vibration absorbers have limited effectiveness in suppressing vibrations in complex systems, especially under fluctuating operating conditions. To overcome this limitation, this paper proposes a robust design method for hybrid dynamic vibration absorbers, which integrate active control to enhance performance. In practical applications, physical parameters such as spring constants and damping coefficients inevitably deviate from their nominal values due to long-term aging and modeling errors. Therefore, the core challenge is designing robust state feedback controllers that can maintain asymptotic stability and control performance despite these parameter uncertainties, while strictly adhering to structural constraints, such as distributed access to state information. The design problem is formulated as a linear matrix inequality (LMI) that incorporates the robust performance objectives of a linear quadratic regulator (LQR). The paper then details a homotopy-based algorithm to solve this complex, constrained LMI problem. Furthermore, to demonstrate the practical applicability of the proposed methodology, this study evaluates the controller design on a higher-order five-mass system. Simulation results confirm that the designed robust controllers effectively attenuate vibrations despite structural limitations and uncertain conditions.
|
| |
| 10:15-10:30, Paper FrAT1.6 | |
| Decentralized H_{infty} Filtering for Interconnected Systems with Structured Lyapunov Matrices |
|
| Chang, Yufang | Hubei University of Technology |
| He, Ziyi | Dalian University of Technology |
| Zhai, Guisheng | Shibaura Institute of Techbology |
| Huang, Wencong | Hubei University of Technology |
Keywords: Control Theory and Applications, Navigation, Guidance and Control, Information and Networking
Abstract: The problem of decentralized H_infinity filtering for interconnected systems is addressed, in which each subsystem is coupled with others through interconnection signals. A sufficient condition, expressed as a bilinear matrix inequality, is derived to ensure asymptotic stability of the filtering error system while satisfying a prescribed H_infinity performance index. Since the bilinear matrix inequality has coupled matrix variables of the Lyapunov matrix P and the decentralized controller coefficient matrix, it is generally difficult to solve. Assuming that the Lyapunov matrix P takes a block diagonal structure as diag{P_1, P_2, ... } corresponding to the interconnection structure of the system, the structure of each Lyapunov matrix P_i is analyzed in this paper, revealing that its block diagonal or non block diagonal forms impacts the filter design and the estimation performance. Simulation examples demonstrate the effectiveness of the proposed approaches.
|
| |
| FrAT2 |
Cattleya, 3F |
| Human-Robot Interaction |
Oral Session |
| |
| 09:00-09:15, Paper FrAT2.1 | |
| WaveSync: Constrained Wavefront Optimization for Synchronized Co-Speech Gestures in Humanoid Robots |
|
| Nguyen Canh, Thanh | Japan Advanced Institute of Science and Technology |
| Tran Viet, Thang | Vietnam National University, Hanoi, University of Engineering and Technology |
| Uong Gia, Huy | University of Engineering and Technology, Vietnam National University |
| Dinh Van, Phuc | VNU University of Engineering and Technology |
| Tuyen, Nguyen Tan Viet | University of Southampton |
| Hoang, Van Xiem | VNU - University of Engineering and Technology |
| Chong, Nak Young | Japan Advanced Institute of Science and Technology |
Keywords: Human-Robot Interaction
Abstract: Expressive co-speech gestures are crucial for natural human--robot interaction, yet generating them on physical humanoid robots remains challenging because, unlike virtual avatars, robots must synchronize gestures with speech under strict kinematic and actuator constraints. We present textbf{WaveSync}, a hybrid framework in which a Large Language Model decomposes dialogue responses into structured semantic schemas and assigns per-word importance weights, forming a continuous Semantic Importance Wave. Gesture trajectories are shaped through Dynamic Movement Primitives to ensure kinematic feasibility while enhancing expressiveness. A Wavefront Optimization stage aligns gesture stroke peaks with speech emphasis peaks and resolves residual temporal conflicts through gesture-duration compression and forward propagation. Experimental evaluation across five dialogue scenarios demonstrates effective gesture--speech alignment and favorable performance in both objective and subjective evaluations. The results further show that the key components of WaveSync contribute to producing gestures that are expressive, semantically grounded, and kinematically feasible. The code, resources, and videos are available at href{https://github.com/pairs-lab/WaveSync}{WaveSync}.
|
| |
| 09:15-09:30, Paper FrAT2.2 | |
| Predictive Intent-Risk Alignment for Cooperative Shared Steering in Obstacle Avoidance |
|
| Lee, Changuk | Pusan National University |
| Ahn, Changsun | Pusan National University |
Keywords: Human-Robot Interaction, Autonomous Vehicle Systems, Artificial Intelligence Systems
Abstract: This paper presents a cooperative shared-steering framework for obstacle avoidance based on dual-intent prediction and a newly defined reciprocal intent-risk alignment (RIRA) model. RIRA is introduced as an operational task level compatibility measure rather than a psychological trust estimate, so that authority can be allocated from measurable prediction outcomes. The driver intent is predicted as a 10-step future steering sequence using a Transformer teacher and a compact multilayer-perceptron student, while the automation intent is obtained by rolling out a proximal-policy optimization steering policy through a vehicle model. RIRA evaluates two complementary factors: intent compatibility from steering-sequence magnitude mismatch and sign conflict, and motion quality from predicted collision-risk and front slip stability. These factors generate a reference authority, and a discrete predictive optimizer selects the final blending coefficient by considering safety, alignment consistency, and authority smoothness. CarSim-Simulink co-simulation for a stationary-obstacle avoidance scenario at 80 km/h shows reductions in potential-field risk, front-tire slip, steering disagreement, and conflict time by 3.43%, 6.26%, 10.51%, and 7.18%, respectively, compared with a baseline allocator without predicted-intent conflict consideration.
|
| |
| 09:30-09:45, Paper FrAT2.3 | |
| Toward Trust-Aware Control in Wearable Gait Assistance: A Pilot Study Linking Controller Conditions and Audited Control Features to Trust-Related Reported Responses |
|
| Cho, Kwonseung | Gwangju Institute of Science and Technology |
| Hur, Pilwon | Gwangju Institute of Science and Technology |
Keywords: Human-Robot Interaction, Exoskeleton Robot, Rehabilitation Robot
Abstract: Trust-aware control in wearable gait assistance requires linking user trust responses to interpretable control signals, but most HRI frameworks treat trust as a single global score. This pilot study organizes 140 expressions from 35 pHRI and wearable robotics sources into trust-related semantic labels, compares subjective ratings across controller conditions, and examines associations between audited control features and trust-related reported responses. Seven post-stroke participants walked with a wearable knee exoskeleton under Zero-Torque, PD, and Active Inference conditions, providing 72 trial-level observations. Participant random-intercept mixed models identified five exploratory associations after BH correction: peak motor power with Wearability experience, and motor-power variability with reported Trust, perceived Ability, Comfort response, and perceived Safety. These exploratory results provide candidate feature–response relationships for future trust-aware control studies rather than validated control mappings.
|
| |
| 09:45-10:00, Paper FrAT2.4 | |
| Design and Kinematic Analysis of a Five-Bar Closed-Chain Bipedal Leg Architecture |
|
| Ganjam Chandramohan, Mohnish | Kyoto University of Advanced Science |
| Nanayakkara, J.A Gajitha Ganganath | Kyoto University of Advanced Science |
| Meitani, Khalid Zeinelabden Yousef | Kyoto University of Advanced Science |
| Nisar, Sajid | Kyoto University of Advanced Science |
Keywords: Human-Robot Interaction, Robot Mechanism and Control, Robotic Applications
Abstract: Recent progress in humanoid robotics has intensified the demand for legs that locomote with human-like efficiency. Most bipedal robots, however, adopt a bent-knee stance to preserve controllability and avoid singularities, which forces the knee actuators to continuously bear gravitational and inertial loads and drives up power consumption.This paper presents the design and kinematic analysis of an adjustable five-bar closed-chain leg that achieves a straight- leg stance with the use of closed loop design. The mechanism couples a biomimetic femur–tibia pair through an actuated cross-link whose length adapts to the leg configuration, and is driven by two actuators: a rotary actuator at the hip and a mode-switching linear–rotary mechanism integrated into the knee link. A design benchmark derived from the femur-to-tibia ratio and the five-bar Grashof criteria sizes the links for any target scale. Kinematic analysis using a decoupled 3+2 link model and the resulting Jacobian shows that the leg reproduces serial 2R-like trajectories and evaluates serial singularities for the proposed closed-chain mechanism.
|
| |
| 10:00-10:15, Paper FrAT2.5 | |
| Decoupled Object-Centric Video Understanding for Generating Robotic Manipulation Commands |
|
| Nguyen Canh, Thanh | Japan Advanced Institute of Science and Technology |
| Tran, Thanh-Tuan | Faculty of Electrical and Telecommunications, University of Engineering and Technology - VietNam National University |
| Zhang, Haolan | Japan Advanced Institute of Science and Technology |
| Gao, Ziyan | Japan Advanced Institute of Science and Technology |
| Hoang, Van Xiem | VNU - University of Engineering and Technology |
| Chong, Nak Young | Japan Advanced Institute of Science and Technology |
Keywords: Human-Robot Interaction, Robot Vision, Artificial Intelligence Systems
Abstract: Translating video demonstrations into executable robot commands remains challenging because existing methods often fail to identify which objects are functionally involved in the demonstrated action. As a result, they may generate commands that are linguistically plausible but operationally ambiguous. We propose an object-centric video understanding framework that decouples action recognition from object identification to generate precise, grammar-free manipulation commands. Our approach integrates Temporal Shift Modules (TSM) for efficient spatio-temporal action classification with a novel textbf{Object Selection} algorithm that identifies task-relevant objects through trajectory-based role classification, blur detection, and overlap minimization. The selected objects are then processed by Vision-Language Models (VLMs) for robust category recognition and zero-shot generalization. Evaluated on a modified Something-Something V2 dataset, our method achieves 86.79% action classification accuracy and BLEU-4 scores of 0.337 on standard objects and 0.261 on novel objects. These results improve over the strongest task-specific baseline by 80.2% and 143.9%, respectively. Larger gains are observed in METEOR and CIDEr, reaching 157.9% and 171.7% on novel objects. Across all semantic metrics, our approach consistently outperforms task-specific methods and remains competitive with, or surpasses, large general-purpose VLMs while retaining a modular, object-centric design.
|
| |
| FrAT3 |
Azalea, 3F |
| Process Control Systems |
Oral Session |
| |
| 09:00-09:15, Paper FrAT3.1 | |
| Data-Driven Predictive Pitch Control of a Wind Turbine above the Rated Wind Speed |
|
| Kim, Daehan | Kwangwoon University |
| Kim, Hyungsuk | Kwangwoon University |
| Back, Juhoon | Kwangwoon University |
Keywords: Control Theory and Applications, Industrial Applications of Control, Process Control Systems
Abstract: This paper presents a data-driven predictive pitch control method for wind turbines operating in the above-rated wind speed region. Under high-wind conditions, wind-speed fluctuations can lead to excessive aerodynamic torque and generator-speed excursions, particularly because the blade-pitch actuator has limited bandwidth and rate constraints. To address this problem, a collective blade-pitch control strategy is developed based on input-output data, while the conventional generator-torque control loop is retained. The proposed controller constructs a non-parametric prediction model from measured data and computes the optimal pitch command in a receding-horizon manner. The control objective is to improve generator-speed regulation and mitigate overspeed while satisfying practical pitch-angle and pitch-rate constraints. Wind-speed information is incorporated into the prediction model as an exogenous signal to enhance the predictive capability under varying wind conditions. The proposed method is evaluated using a nonlinear simulation model of a utility-scale wind turbine in above-rated wind scenarios. The simulation results demonstrate that the proposed data-driven predictive controller generates proactive pitch actions in response to wind-speed variations and maintains stable generator-speed and power regulation under high-wind operation.
|
| |
| 09:15-09:30, Paper FrAT3.2 | |
| Optimal Electric Bus Depot Charging: Cost Savings, Grid Limits, and Robustness Trade-Offs |
|
| Widmer, Fabio | ETH Zürich |
| Pinter, Luca | ETH Zürich |
| Moradi, Mohammad Hossein | ETH Zurich |
| Onder, Christopher | ETH Zürich |
Keywords: Control Theory and Applications, Industrial Applications of Control, Process Control Systems
Abstract: Depot charging of electric bus fleets must minimize electricity costs, respect grid limits, and remain feasible despite uncertain trip energy demand. While cost-optimal charging is well studied, its value under different electricity prices and grid connection capacities, as well as the economic cost of robustness, remain poorly quantified. We address these gaps with a convex robust formulation in which bounded demand uncertainty is enforced through worst-case state-of-energy constraints. The formulation is evaluated against charge-on-arrival using realistic service schedules for four Swiss depots containing 7-35 buses. In the depots studied, smaller depots achieve the greatest relative benefit from optimization, with total electricity cost reductions exceeding 50%, because optimization mitigates charging peaks that strongly affect their costs. For all depot sizes, the savings from optimization increase with electricity price volatility. Optimization can also lower the grid capacity required for feasible operation by over 40%, although tight limits reduce peak shaving potential. Protection against energy-demand deviations of 10% increases total electricity cost by less than 0.1%. The resulting charging power profiles exhibit interpretable price-threshold and peak-shaping behavior, providing practical guidance for real-world implementations.
|
| |
| 09:30-09:45, Paper FrAT3.3 | |
| Techno-Economic Evaluation of the CIRCE Process Via an Equation-Oriented Model |
|
| Hyeon, Jisu | Incheon National University, Incheon |
| Kim, Jong Woo | Incheon National University |
Keywords: Process Control Systems, Industrial Applications of Control, Control Theory and Applications
Abstract: An equation-oriented (EO) techno-economic optimization framework is developed for a CIRCE-based heavy-water production system with five LPCE columns fed by alkaline water electrolysis (AWE) and/or steam methane reforming (SMR). A stack-level AWE model with isotope-resolved deuterium selectivity is coupled with the CIRCE material balance into a single-stage NLP that simultaneously optimizes stack temperature/current, column height, and operating pressure under an LCOH objective. At zero carbon tax, the smr+smr scenario attains the lowest LCOH of 2.24 USD/kg, about 16% below ele+ele at 2.67 USD/kg. Above approximately 21.7 USD/tCO2, the fully electrified configuration becomes the most economical; this study quantifies the crossover at the integrated process level.
|
| |
| 09:45-10:00, Paper FrAT3.4 | |
| High-Speed Position Sensorless Control of SRM Using Commutation-Current Response |
|
| Kim, Gyuwon | Kumoh National Institute of Technology |
| Park, Seung Yun | Kumoh National University of Technology |
| Ban, Jaepil | Kumoh National Institute of Technology |
Keywords: Sensors and Signal Processing, Process Control Systems, Control Devices and Instruments
Abstract: This paper presents a high-speed position sensorless control method for switched reluctance motor (SRM) drives using the early commutation current response of the incoming phase. The proposed method constructs a phase-dependent lookup table between rotor position and the phase current sampled at a fixed short delay after phase turn-on. Since this feature is based on the short-duration current response after voltage excitation, it can be naturally obtained during low-speed model-free sensorless operation, enabling a simple hybrid transition from low-speed operation to medium-to-high-speed sensorless control. During operation, the measured early current response is inversely mapped to rotor-position candidates through the lookup table. Since the same current response can correspond to multiple rotor-position candidates, current-slope variation is used as a branch discriminator rather than as an exact position marker. The selected discrete position information is then propagated by an event-based predictor to obtain a continuous rotor-position estimate for commutation timing and current control.The proposed framework does not require premeasured magnetic characteristic data, motor models, additional hardware, or computationally intensive estimation algorithms, thereby reducing implementation complexity. The effectiveness of the proposed approach is validated through experiments.
|
| |
| FrAT4 |
Lilac, 3F |
| Robot Mechanism and Control 2 |
Oral Session |
| |
| 09:00-09:15, Paper FrAT4.1 | |
| Momentum-Observer-Based State Augmentation for VLA-Based Mass Comparison |
|
| Kim, Chanwoo | Kookmin University |
| Cho, Baek-Kyu | Kookmin University |
Keywords: Robot Mechanism and Control, Artificial Intelligence Systems, Robotic Applications
Abstract: Vision-Language-Action (VLA) models have shown strong generalization capabilities in robotic manipulation. However, most VLA models rely heavily on visual observations, which limits their ability to infer physical interaction properties such as mass differences between visually indistinguishable objects. In this paper, we present a vertical- disturbance-based state input augmentation method for VLA-based robotic mass comparison using momentum-observer- based disturbance estimation. The proposed method estimates external disturbances from the robot’s proprioceptive state without additional force/torque or tactile sensors, and transforms the estimated joint-space disturbance into end-effector disturbance components. The vertical disturbance estimates of the left and right end-effectors are then selected and appended to the state input of the VLA policy, providing physical interaction cues during object lifting. The proposed method is evaluated on a dual-arm mass comparison task in which the robot lifts two visually identical cubes with different masses and places the heavier cube into a tray. Across 300 paired evaluation trials, adding the vertical disturbance estimates of the two end-effectors improved the task success rate from 60.3% for the baseline policy without disturbance information to 77.3%. These results demonstrate that momentum-observer-based vertical disturbance estimates can serve as sensorless physical cues for improving the mass comparison performance of VLA policies.
|
| |
| 09:15-09:30, Paper FrAT4.2 | |
| Radial Basis Function Neural Network-Aided Data-Driven Unscented Kalman Filter for Nonlinear Systems |
|
| Matsubara, Isaki | Kyoto Institute of Technology |
| Sawada, Yuichi | Kyoto Institute of Technology |
Keywords: Robot Mechanism and Control, Artificial Intelligence Systems, Sensors and Signal Processing
Abstract: State estimation is an important problem for the control of partially observed and noisy nonlinear dynamical systems. The Kalman filter (KF) and its nonlinear extensions are widely used for this purpose, but their performance strongly depends on the accuracy of the dynamical model. Therefore, data-driven state estimation is useful when detailed modeling or system identification is difficult. This paper proposes an RBFNN-aided data-driven unscented Kalman filter (UKF) for improving state-estimation accuracy from a small collected dataset without detailed modeling. The proposed method introduces a radial basis function neural network (RBFNN) as a compensator for the approximate dynamics used in the UKF. The compensator is trained offline from observed input-output data by minimizing a criterion based on the innovation likelihood and the prior error covariance, without using the true state as supervised data. The proposed method is evaluated through Acrobot simulations with two different types of approximate models. The results show that the proposed compensator reduces the normalized mean squared error (NMSE) on both the training and test datasets, and that the modeling error of the approximate model affects the performance on the test dataset.
|
| |
| 09:30-09:45, Paper FrAT4.3 | |
| Series-Based Finite Prediction Horizon Model Predictive Control and Its Applications in UAV Attitude Control |
|
| Sun, Siyuan | National University of Singapore |
| Huang, Sunan | National University of Singapore |
Keywords: Robot Mechanism and Control, Autonomous Vehicle Systems, Control Devices and Instruments
Abstract: This paper proposes a series-based model predictive control (MPC) method that reduces computational complexity and load. By utilizing series expansion techniques, the proposed method transforms the traditional MPC problem into an algebraic computational problem, making it more suitable for high-frequency real-time applications, for example, we apply it to the control issue in UAV rotational dynamics.
|
| |
| 09:45-10:00, Paper FrAT4.4 | |
| Quantification of Actuator Placement Effects on Proximal-Joint Current Via Serial and Parallel Wrist Modules |
|
| Jeon, Minhyuk | Sungkyunkwan University |
| Kim, Juhyun | Seoul National University |
| Sung, Eunho | Seoul National University |
| You, Seungbin | Korea Institute of Science and Technology |
| Park, Jaeheung | Seoul National University |
Keywords: Robot Mechanism and Control, Control Devices and Instruments
Abstract: A parallel wrist generally differs from its serial counterpart in workspace, singularity behaviour, achievable wrist angular velocity, payload capability, and the spatial distribution of the wrist actuator mass at the same time, so a direct comparison cannot isolate which property is responsible for an observed effect. This study quantifies the extent to which the actuator current at the shoulder and elbow is attributable specifically to wrist actuator placement, by comparing a serial wrist and a parallel wrist that are equalized in the other properties. A parallel wrist with two degrees of freedom (DOF) is presented for a humanoid manipulator with 7-DOF; roll and pitch motion of the wrist is generated by two Prismatic-Spherical-Spherical branches sharing a central universal joint, with each branch driven by a lead screw that converts the rotation of its actuator into a linear slider displacement. The geometric parameters of the wrist are determined by a constrained genetic algorithm optimization that equalizes the parallel wrist to its serial counterpart in workspace, range free of singularities, and achievable wrist angular velocity, while module mass is equalized by construction to within 1%. The two equalized modules are exchanged on the same arm and compared under a common cyclic trajectory, and the resulting reduction in the actuator current at the shoulder and elbow is reported as an estimate of the placement effect.
|
| |
| 10:00-10:15, Paper FrAT4.5 | |
| Global Corridor-Guided MPPI for Robotic Manipulator Pose Control in Obstacle-Constrained Environments |
|
| Lee, Munhaeng | Pukyong National University |
| Suh, Jinho | Pukyong National University |
Keywords: Robot Mechanism and Control, Control Theory and Applications
Abstract: This paper proposes a global corridor-guided model predictive path integral control framework (GC-MPPI) for manipulator pose control. The proposed framework first generates a collision-free global path from the current tool center point (TCP) pose to the target pose. Using path nodes as centers, we construct a global corridor, with its width adaptively scaled according to the clearance from surrounding obstacles. This method preserves the local sampling flexibility of MPPI and mitigates the dead-end behavior caused by the constrained environments. During online control, we compute a Corridor Guiding Vector (CGV) by combining a tangential component along the corridor centerline and an inward component activated near the corridor boundary. The corridor and CGV costs are incorporated into the MPPI cost function, where the corridor cost encourages the TCP to remain within the feasible motion region and the CGV cost provides directional guidance along the corridor. To validate the performance of the proposed framework, we perform a simulation using a 6-DoF manipulator in a detour scenario. The simulation results show that GC-MPPI avoids the stagnation observed in baseline MPPI and reaches the target TCP pose. In addition, compared with GP-MPPI w/ 0.05, GC-MPPI reduces the reaching and the settling times by 20.2% and 19.3%, respectively.
|
| |
| 10:15-10:30, Paper FrAT4.6 | |
| On Comparison of Discrete Controllers for 6 Degree of Freedom Manipulator Robot |
|
| Jansri, Anurak | Bangkok Thonburi University |
| Rianpreecha, Chompoonut | Bangkok Thonburi University |
| Yoneyama, Keito | Chulalongkorn University |
Keywords: Robot Mechanism and Control, Control Theory and Applications, Industrial Applications of Control
Abstract: A Six Degree of Freedom (6DoF) manipulator robot is widely used in various applications such as line production in manufacturing and station maintenance in space. From many applications, the challenges faced in control engineering is precise control system for robotic movement and robotic position. This paper presents kinematics robot modelling via Denavit-Hartenberg (DH) method, applied forward and inverse kinematics, and the well-known controller in order to overcome the challenges. Both discrete PID and discrete Fuzzy-PI controllers are proposed. MATLAB/Simulink is mainly used to simulate the robot and controllers to observe robot’s behavior in continuous and position movements. The simulation results are carried out and showed satisfactory performance using three standard criteria including Mean Absolute Error, Mean Square Error, and Root Mean Square Error.
|
| |
| FrAT5 |
Crown, 3F |
| Sensors and Signal Processing |
Oral Session |
| |
| 09:00-09:15, Paper FrAT5.1 | |
| Evaluation of Subject Composition Dependence and Individual Calibration Potential in Lifted-Weight Regression Using a Single IMU |
|
| Tomura, Kento | Meiji University |
| Hayasaka, Haruki | Meiji University |
| Itami, Taku | Meiji University |
Keywords: Sensors and Signal Processing, Artificial Intelligence Systems, Biomedical Instruments and Systems
Abstract: Lifted object mass is one objective factor de scribing manual lifting conditions and may provide supplementary information for workload assessment. Previous studies have mainly investigated posture or risk assessment using multiple inertial measurement units (IMUs), or weight classification using a single IMU; however, continuous lifted weight regression using a chest-mounted single IMU and its generalization performance for unseen subjects have not been sufficiently clarified. In this study, we constructed lifted-weight regression models using a chest-mounted single IMU and consistently evaluated five-subject nested leave one-subject-out (LOSO), all ten three-subject combinations, unseen-subject evaluation, and few-shot Ridge calibration using a small amount of labeled personal data under controlled feature construction and selection conditions. In the five-subject nested LOSO evaluation, the best Ridge model achieved MAE = 3.122 kg and R2 = 0.105, confirming that room for improvement remains in average subject-independent regression. In contrast, across all ten three-subject combinations, within-group MAE varied from 2.315 to 4.062 kg, showing that performance differences attributable to subject composition exceeded those among baseline models. Moreover, within-group performance did not simply correspond to mean generalization performance for unseen subjects, indicating that high within-group performance does not necessarily guarantee generalization to unseen subjects. Few-shot Ridge calibration yielded large error reductions for some subjects, suggesting the potential for calibration using a small amount of labeled personal data, whereas its effect was subject-dependent. These results show that, in lifted-weight regression using a single IMU, average performance alone cannot sufficiently characterize generalization behavior; subject composition dependence, generalization performance for unseen subjects, and individual calibration potential using a small amount of labeled personal data should be evaluated separately.
|
| |
| 09:15-09:30, Paper FrAT5.2 | |
| A Multi-Stream Hierarchical Head for Automatic Pronunciation Assessment |
|
| Jeankor, Benjamaporn | School of Information Technology, King Mongkut’s Institute of Technology Ladkrabang |
| Ploysuwan, Tuchsanai | King Mongkut's Institute of Technology Ladkrabang |
Keywords: Sensors and Signal Processing, Artificial Intelligence Systems, Multimedia Systems
Abstract: Automatic pronunciation assessment (APA) gives second-language learners fine-grained, multi-aspect feedback at the phone, word, and utterance levels. State-of-the-art systems increasingly fine-tune large self-supervised speech encoders. Updating hundreds of millions of parameters, however, is computationally expensive and tends to over-fit the few thousand expert-scored utterances typical of APA datasets. We show that once the encoder is frozen, the dominant lever is not trunk size but complementary per-phone feature engineering and training objectives. Our model, Qwen3-APA, keeps a pre-trained Qwen3-ASR encoder frozen and trains only a compact 1.26M-parameter hierarchical head, comparable in trainable size to the HMamba baseline. A learned gated fusion combines eight complementary per-phone streams: four goodness-of-pronunciation (GOP) variants, namely canonical, class-normalized, second-recognizer, and SSL-template, together with articulatory, prosodic, and two self-supervised views. The head then refines the phone-, word-, and utterance-level representations with a dilated temporal-convolution branch and a residual graph-convolution module, all in a single forward pass. Qwen3-APA ranks first among the compared systems on eight of the ten SpeechOcean762 metrics, including a sentence-level total PCC of 0.839. It also reaches 0.746 on completeness, well above the 0.16–0.51 of prior systems.
|
| |
| 09:30-09:45, Paper FrAT5.3 | |
| SiL-DT: A Simulation-In-The-Loop Digital Twin Framework for Resilient UAV Swarms in Dynamic Urban Environments |
|
| Tomiwa Kunle, Oluwadare | University of Staffordshire |
| Lutta, Pantaleon | University of Staffordshire |
| Oluwasegun, Habeeb Alli | University of Staffordshire |
Keywords: Sensors and Signal Processing, Navigation, Guidance and Control, Autonomous Vehicle Systems
Abstract: Unmanned Aerial Vehicle (UAV) swarms have advanced significantly, with multi-UAV systems offering advantages such as redundancy, cost-effectiveness, and reduced dependency on infrastructure. These systems benefit applications like remote surveillance, search and rescue, and environmental monitoring. Ideally, UAVs should operate distributively and make autonomous decisions. However, developing distributed algorithms for multi-UAV systems poses challenges, particularly in real-world testing, where hardware constraints, environmental factors, and limited analytics hinder progress. Simulations have emerged as a key solution, enabling rapid iteration and detailed analytics. Yet, current simulators often lack the ability to accurately model both physical dynamics of UAVs and complexity of their communication networks. This paper introduces SiL-DT, a Simulation-in-the-Loop Digital Twin framework designed to bridge this gap. Specifically, we combine Gazebo for physics simulation, PX4 for flight control, network emulation principles, and empirical models for energy and computation. Experimental results demonstrate improved mission success under increasing jamming intensity, with success rates of 95%, 88%, and 75% for SiL-DT compared with 85%, 65%, and 40% for the baseline at jamming intensities of 0.2, 0.5, and 0.8, respectively.
|
| |
| 09:45-10:00, Paper FrAT5.4 | |
| Unified Grid Evaluation of IMU-Based Gait Event Detection: Algorithms, Sensor Locations, and Robustness to Perturbations |
|
| Park, Jihwan | Sejong University |
| Yang, Hyungseung | Sejong University |
| Kim, Gahyun | Sejong University |
| Lim, Taesoo | Sejong University |
| Kang, Brian Byunghyun | Sejong University |
Keywords: Sensors and Signal Processing, Rehabilitation Robot, Exoskeleton Robot
Abstract: Practical deployment of inertial measurement unit (IMU)-based gait event detection requires choices in sensor placement, signal channel, algorithm, and robustness to perturbations, yet existing studies address these dimensions individually. This study compares 5 sensor locations × 9 channels × 4 algorithms (TB, prominence-adaptive TB, BPF, template matching) on 4 normal-gait subjects, and further analyzes 5 single perturbations, 4 compound scenarios, and 9-channel 1D-CNN and 1D-LSTM via Leave-One-Subject-Out (LOSO) 4-fold cross-validation. In the grid evaluation, adaptive prominence thresholding produced statistically significant differences from fixed thresholding in 30 of 90 cells (Wilcoxon p < 0.05), most prominently on non-standard channels (Acc, Euler). BPF was the most robust to noise and DC bias. The product of single-perturbation degradation ratios matched the measured compound degradation within 0.03, suggesting a multiplicative relationship within the tested range. Both 9-channel deep models exceeded single-channel simple algorithms by +0.09 to +0.17 F1, with LSTM more stable in out-of-sample generalization than CNN. The results yield accuracy-first, robustness-first, and simplicity-first recommendations for IMU-based gait detection system design.
|
| |
| 10:00-10:15, Paper FrAT5.5 | |
| Development of Robot Control System Using Computer Vision Based Finger Gesture Recognition |
|
| Goto, Kohei | Kitami Institute of Technology |
| Ravankar, Abhijeet | Kitami Institute of Technology |
Keywords: Sensors and Signal Processing, Robot Mechanism and Control, Artificial Intelligence Systems
Abstract: Approximately 1% of the global population requires mobility support, yet conventional joystick-controlled wheelchairs remain difficult to be operated for individuals with severe neuromuscular disorders. Although alternative interfaces such as EEG, eye-gaze, EMG, and wearable sensors have been investigated, they frequently suffer from latency, high training overhead, unintended command triggers, or physical discomfort. Similarly, state-of-the-art deep-learning based hand gesture recognition often demands substantial computational resources. To address these limitations, this study proposes a lightweight, non-contact robot control system based on finger gesture recognition utilizing residual physical functions. The system extracts spatial features by combining physical color markers with a custom computer vision pipeline consisting of hand segmentation, centroid detection, and the Hough Transform. To avoid relying solely on the classification accuracy of the neural networks, a frame differencing algorithm is integrated into the decision-making process, thereby ensuring robust control even during occasional prediction errors. Operating in indoor environments with stable lighting, this hybrid framework eliminates the need for complex illumination compensation or heavy neural networks, providing a computationally efficient, stable, and intuitive solution to enhance user independence and overall Quality of Life.
|
| |
| 10:15-10:30, Paper FrAT5.6 | |
| Real-Time Simultaneous Estimation of Direction and Distance of Warning Sounds Using Two Microphones for Hearing Assistance |
|
| Shiratori, Shodai | Meiji University |
| Izumi, Takuto | Meiji University |
| Itami, Taku | Meiji University |
Keywords: Sensors and Signal Processing, Biomedical Instruments and Systems, Artificial Intelligence Systems
Abstract: In recent years, the immediate recognition of warning and environmental sounds occurring in traffic environments and daily life spaces has become an important issue for ensuring the safety of individuals with hearing loss. Therefore, this study proposes a real-time sound source localization model capable of simultaneously estimating the direction and distance of warning sounds to assist people with hearing loss. This model uses only two microphones as a minimal configuration. By integrating time difference of arrival (TDOA) and interaural level difference (ILD) within a CRNN-based framework, it simultaneously estimates the direction and distance of the sound source. Specifically, it calculates the sound source direction through time difference of arrival (TDOA) estimation based on the generalized cross-correlation phase transform method, and combines this with a distance estimation method based on the interaural level difference (ILD) between the micro phones. As a result, the proposed method showed stable direction estimation performance under real environmental conditions and achieved a distance estimation error of 0.15 0.81m under real-environment conditions. Furthermore, it was confirmed to be compatible with real-time visual presentation. The contribution of this study lies in the construction of a sound source localization model capable of simultaneously estimating the direction and distance of warning sounds with a minimal microphone configuration and computational load.
|
| |
| FrAT6 |
Symphony A, 4F |
| Robotic Applications 3 |
Oral Session |
| |
| 09:00-09:15, Paper FrAT6.1 | |
| Process and Alignment-Aware Pick-And-Place Automation with Two Humanoid Dual-Arm Mobile Manipulators |
|
| Park, Woong-hee | Kyungpook National University |
| Joe, Hyun-Min | Kyungpook National University |
Keywords: Robotic Applications, Industrial Applications of Control
Abstract: This paper proposes a process- and alignment-aware pick-and-place automation system for practical manufacturing environments using two RB-Y1 humanoid upper-body dual-arm mobile manipulators. The system implements an inbound/outbound workflow in which the inbound robot transfers products from a rack to a conveyor, and, after drilling and vision inspection, the outbound robot inserts conveyor products into target rack slots. A ZED RGB-D camera and YOLO-based object detection estimate the product class and 3-D center position, while AprilTag-based conveyor/rack localization and camera-to-base transformation generate grasp and place target poses. A process-state-based task decision logic determines the grasping arm, target rack, target slot, approach direction, and placement offset from the product class, OK/NG inspection state, lateral position on the conveyor, and rack-slot occupancy. To handle mobile-base yaw misalignment, relative yaw estimated from conveyor/rack markers is reflected in wrist/end-effector orientation commands as marker-based alignment compensation. Experiments with two RB-Y1 robots achieved success rates of 20/20 for inbound tasks and 19/20 for outbound tasks. In conveyor grasping, stable grasps were achieved under relative yaw errors up to 35° when compensation was applied. These results verify the applicability of the proposed system to manufacturing-process automation.
|
| |
| 09:15-09:30, Paper FrAT6.2 | |
| Digital-Twin Fine-Tuned RL Planning with BEV Traversability Cost Maps for Quadruped Navigation in Unstructured Outdoor Environments |
|
| Sohn, Minsung | Kyungpook National University |
| Joe, Hangil | Kyungpook National University |
Keywords: Robotic Applications, Navigation, Guidance and Control, Autonomous Vehicle Systems
Abstract: This paper presents a hierarchical navigation framework for a quadruped robot in unstructured outdoor environments. Depth-derived geometric features and SegFormer-based semantic costs are fused into a unified Bird's-Eye-View (BEV) traversability cost map. On top of this map, a discrete trajectory-bank planner is trained with Proximal Policy Optimization (PPO) to select velocity commands. To balance general navigation capability with site-specific adaptation, the high-level policy is first pre-trained in a procedurally generated outdoor environment and then fine-tuned in a photogrammetry-based digital twin of the deployment site to adapt it to the target domain. A frozen low-level locomotion policy executes the selected command, decoupling navigation from motion control. On unseen digital-twin sites, the framework outperformed both baselines on two unstructured environments, achieving up to an 82.5% improvement in success rate and a 59.9% improvement in path efficiency over the pre-trained-only policy. On a physical quadruped at the target site, fine-tuning in the digital twin improves goal-reaching over the pre-trained-only policy, giving preliminary evidence that digital-twin fine-tuning benefits real-world navigation.
|
| |
| 09:30-09:45, Paper FrAT6.3 | |
| Experimentally Validated Dynamic FEM-Based Modeling of a Pneumatic Morphing Soft Quadrotor |
|
| Sumathy, Vidya | Luleå University of Technology |
| Labra Caso, Fernando | Luleå University of Technology |
| Ferrentino, Pasquale | Vrije Universiteit Brussels |
| Haluska, Jakub | Luleå University of Technology |
| Vanderborght, Bram | Vrije Universiteit Brussel |
| Nikolakopoulos, George | Luleå University of Technology |
Keywords: Robotic Applications, Robot Mechanism and Control, Control Devices and Instruments
Abstract: This work presents an experimental evaluation of a finite-element model of the Pneumatic Morphing Soft Quadrotor (PMSQ) arm components and validates the model against real-time experiments under dynamic conditions. The PMSQ has four soft pneumatically actuated arms that deflect under differential chamber pressures to enable in-flight morphing. The PMSQ has four soft pneumatically actuated arms that deflect under differential chamber pressures to enable in-flight morphing. The pneumatic arm is modeled in SOFA and consists of an hyperelastic actuator and a linear-elastic spine. To validate the model, real-time experiments are conducted on the physical soft arm, which is part of the PMSQ, actuating one arm at a time and recording the arm's base and tip poses using Vicon under commanded pressure cycles and varying propeller thrust. Similar experiments are also conducted using the soft arm of PMSQ SOFA model in the simulation. Experimental and simulated data are then processed to compute 3D arm deflection and orientation, and the results are compared to assess the SOFA model’s accuracy under pneumatic actuation and aerodynamic loading. Comparison results show strong alignment in deformation and deflection across axes, confirming the model’s accuracy for configuration-dependent response under pneumatic actuation and aerodynamic loading. Thus, the validated finite-element model provides a basis for quantifying configuration-dependent aero-structural coupling under aerodynamic disturbances for future morphing-flight studies.
|
| |
| 09:45-10:00, Paper FrAT6.4 | |
| Bipedal Locomotion Control of a Snake-Like Robot Via Velocity-Constrained Deep Reinforcement Learning |
|
| Kimoto, Tsuyoshi | Osaka Metropolitan University |
| Yamano, Akio | Osaka Metropolitan University |
| Iwasa, Takashi | Osaka Metropolitan University |
Keywords: Robotic Applications, Robot Mechanism and Control, Control Theory and Applications
Abstract: Snake-like robots exhibit high adaptability through various locomotion modes, but conventional methods face physical limitations in complex three-dimensional environments due to continuous ground contact. Serpentine and rolling motions require substantial clearance, making it difficult to traverse narrow spaces or overcome high vertical steps. To address these limitations, this paper proposes a novel bipedal walking mode for a slender, multi-link snake-like robot. This mode minimizes the robot’s footprint and enhances its step-over capability without requiring dedicated parts or altering the hardware configurations needed for other gaits. To control this highly nonlinear system under unknown environment variations, we employ deep reinforcement learning via Proximal Policy Optimization within the Genesis physics engine. A robust locomotion policy is developed using a two-stage training approach that transitions from standing to walking, integrated with domain randomization on ground bumps. Simulation results demonstrate that incorporating a soft-penalty pre-training phase, followed by strict joint constraints, successfully yields a stable walking policy capable of enduring rough terrains.
|
| |
| 10:00-10:15, Paper FrAT6.5 | |
| State-Based Posture Adjustment Control for Continuous Transverse Ledge Brachiation Robots |
|
| Lin, Yen-Jui | National Taiwan University of Science and Technology |
| Pangestu, Reno | National Taiwan University of Science and Technology |
| Lin, Chi-Ying | National Taiwan University of Science and Technology |
Keywords: Robotic Applications, Sensors and Signal Processing, Control Theory and Applications
Abstract: This paper presents a state-based motion planning and posture adjustment strategy for a transverse ledge brachiation robot that moves horizontally across a wall ledge. Unlike conventional bar brachiation, ledge brachiation relies on partial force closure, which allows small gripper slips and body yaw errors to accumulate over repeated cycles, eventually resulting in a collision or fall. This proposed study consists of a compact symmetric robot model, a four-phase finite-state gait, phase-dependent trajectory generation, wrist-based body posture compensation, and a posture adjustment phase for grasping depth compensation. This additional contribution takes the form of a depth/yaw release quality parameter that is specifically established by the two grippers, triggering corrective action before the next hand release. The controller determines swing timing, pre-impact preparation, grasp completion, and release quality by combining joint feedback, gripper hard-contact dynamic information, inertial measurement data, and distance measurements. The proposed system is validated through simulations using PyBullet physics engine. The simulation results demonstrate that without depth compensation control, the gripper gradually drifts toward the wall, resulting in a bottom-side ledge collision. The proposed posture compensation strategy assists the robot to maintain a stable grasping depth within the commanded range and complete 108 continuous brachiation cycles over a 17-meter distance in 150 seconds. The cumulative COT decreases from the start-up peak of 24.7 to 2.37 over long-distance locomotion under the simulation.
|
| |
| 10:15-10:30, Paper FrAT6.6 | |
| Impact of Design Modernization on Seismic Signal Recording Quality in Autonomous Seismic Acquisition Device (ASAD) |
|
| Kartashova, Mariia | Aramco Research Center - Moscow |
| Timoshenko, Artem | Aramco Research Center - Moscow |
| Yashin, Grigoriy | Aramco Research Center - Moscow |
| Hamadov, Rustam | Aramco Research Center - Moscow |
| Danilin, Maksim | Aramco Research Center – Moscow |
| Popkov, David | Aramco Research Center – Moscow |
| Golikov, Pavel | EXPEC Advanced Research Center |
Keywords: Sensors and Signal Processing, Robotic Applications
Abstract: Conventional seismic data acquisition involves several labor-intensive stages, including delivery, placement of sensors, data recording, and sensor retrieval, which are accompanied by significant time and labor costs, as well as operational challenges in difficult terrain and extreme climatic conditions. For this reason, there is a need to automate field operations. One of the developing areas is the use of UAVs for sensor delivery, deployment, data acquisition, and return to base. This paper presents experimental results obtained during the modernization of these UAVs for seismic data acquisition by switching to recording seismic data using accelerometers only. Two experiments were conducted: the first was aimed at validating the performance of the new UAV configuration under field conditions, while the second focused on comparing the quality of the recorded seismic signal among different versions of Autonomous Seismic Acquisition Devices (ASADs). The results showed that the updated version of the robot is superior to the previous generation in terms of signal quality, while also providing greater reliability and ease of use.
|
| |
| FrAT7 |
Symphony B, 4F |
| Artificial Intelligence Systems 5 |
Oral Session |
| |
| 09:00-09:15, Paper FrAT7.1 | |
| Acoustic-Based Bogie Health Monitoring: A Review for Edge AI Onboard Anomaly Detection |
|
| Nicodeme, Claire | Alstom |
Keywords: Artificial Intelligence Systems, Sensors and Signal Processing, Control Devices and Instruments
Abstract: Railway bogies are safety-critical subsystems whose condition directly affects vehicle stability, ride comfort, reliability, and maintenance cost. While vibration-based monitoring remains the industrial baseline, acoustic approaches are gaining interest for earlier detection of weak anomalies and easier integration into onboard architectures. This paper presents a deployment-oriented review of bogie health monitoring methods with a focus on acoustic anomaly detection under onboard Edge Artificial Intelligence (Edge AI) constraints. The paper first links major bogie failure mechanisms to their observable dynamic and acoustic signatures, then compares the main sensing modalities - vibration, airborne acoustics, and acoustic emission - from the viewpoint of field deployability. It subsequently reviews anomaly detection methods compatible with embedded execution, including compact time-frequency representations, lightweight convolutional neural networks, autoencoders, generative models, and classical outlier-detection algorithms. Beyond a narrative synthesis, the paper proposes (i) a deployment-oriented taxonomy of sensing and inference options, (ii) a comparative framework for Edge AI-compatible approaches, and (iii) practical design guidelines for balancing detection challenges. Compact acoustic representations combined with lightweight models currently provide the most realistic path toward scalable, real-time onboard bogie monitoring.
|
| |
| 09:15-09:30, Paper FrAT7.2 | |
| Comparative Study of Neural Network and Physics-Informed Models for Leak Localization in Main Pipelines |
|
| Satybaldina, Dana | Department of Systems Analysis and Control, L.N. Gumilyov Eurasian National University |
| Shmitov, Nurbol | Department of Systems Analysis and Control, L.N. Gumilyov Eurasian National University |
| Teshebayev, Nurdaulet | Department of Artificial Intelligence and Big Data, Al-Farabi Kazakh National University |
| Kulniyazova, Korlan | Department of Systems Analysis and Control, L.N. Gumilyov Eurasian National University |
| Zakarina, Aina | Department of Systems Analysis and Control, L.N. Gumilyov Eurasian National University |
| Kissikova, Nurgul | Department of Systems Analysis and Control, L.N. Gumilyov Eurasian National University |
Keywords: Artificial Intelligence Systems, Sensors and Signal Processing, Process Control Systems
Abstract: Main oil and gas pipelines are critical infrastructure facilities whose reliability and operational safety depend on effective monitoring technologies. With the digital transformation of the oil and gas industry, rapid leak detection and accurate leak localization are increasingly important for reducing economic losses, improving industrial and environmental safety, and minimizing emergency risks. This study investigates leak location determination in a main pipeline using hydraulic parameters measured at control points along the pipeline. A dynamic simulation model of the pipeline system was developed in MATLAB Simulink using the Simscape Fluids library to generate the experimental dataset. Leak scenarios were simulated at five fixed locations: 44, 60, 70, 90, and 99 km. Three models were evaluated: a multilayer perceptron (MLP), a cascade-forward artificial neural network (Cascade-forward ANN), and a physics-informed neural network (PINN). Pressure and flow rate values from multiple pipeline sections, together with inlet and outlet boundary conditions, were used as input features. Comparative analysis based on RMSE and MAE showed that the PINN model achieved the highest localization accuracy for the initially held-out 70 km location with the lowest prediction error (RMSE = 2.1009 km, MAE = 1.8460 km). Additional leave-one-location-out validation and ablation experiments were conducted to assess spatial generalization and the contribution of the physics-based constraint. The ablation results showed that incorporating the mass-balance constraint reduced the localization error in the primary evaluation setting. The overall results demonstrate the effectiveness of combining data-driven learning with physics-based regularization for pipeline leak localization under transient hydraulic conditions.
|
| |
| 09:30-09:45, Paper FrAT7.3 | |
| Optimal-Control-Based Training of Neural ODEs with Turnpike Regularization and Adaptive Gradient Descent |
|
| Sanngai, Niyata | Chiang Mai University |
| Warwicker, John Alasdair | Lancaster University Leipzig |
| Wongkaew, Suttida | Department of Mathematics, Faculty of Science, Chiang Mai University |
Keywords: Artificial Intelligence Systems, Control Theory and Applications, Information and Networking
Abstract: Neural Ordinary Differential Equations (Neural ODEs) provide a continuous-depth framework for learning nonlinear dynamical transformations, but their training can be hindered by nonconvex loss landscapes and unstable adjoint gradients. We formulate Neural ODE training as a finite-horizon optimal control problem and introduce a Turnpike-inspired trajectory-tracking term that guides the continuous-depth flow toward prescribed target states. For the resulting Pontryagin-based gradient system, we propose Adaptive Gradient Descent (AdapGD), which uses a one-time bisection initialization followed by curvature-based step-size updates without repeated line searches. The method is evaluated on Two Rings, Two Moons, and Two Spirals classification problems and compared with Adam and vanilla gradient descent over five random seeds. AdapGD generally achieves lower terminal loss, faster attainment of the target accuracy, and more reliable geometric separation. On Two Spirals, it achieves an 80% success rate for accuracy above 85%, compared with 60% for Adam and 0% for vanilla gradient descent.
|
| |
| 09:45-10:00, Paper FrAT7.4 | |
| An Online Scheduling Method for Car Sharing Service by Vehicle Dispatch and Customer Assignment Optimization |
|
| Kobayashi, Hayato | Osaka University |
| Sakurama, Kazunori | The University of Osaka |
Keywords: Autonomous Vehicle Systems, Civil and Urban Control Systems, Artificial Intelligence Systems
Abstract: Existing studies on vehicle dispatch optimization for car-sharing systems generally assume that reservation information for all users is known in advance. Consequently, online dispatching decisions for newly arriving requests during operation—such as whether to accept a request and which vehicle or station to assign—have not been sufficiently investigated. To address this issue, this study proposes an online vehicle assignment method by extending an existing optimization model. The proposed method fixes the assignments of users with advance reservations and sequentially processes newly arriving requests in an online manner. Specifically, variables representing request rejection and the associated rejection costs are newly introduced, enabling simultaneous optimization of both acceptance/rejection decisions and vehicle assignments.The results demonstrate that the vehicle dispatch optimization problem in an online environment can be effectively solved without significantly modifying the framework of the existing model.
|
| |
| 10:00-10:15, Paper FrAT7.5 | |
| GRACE: Lightweight Symbolic Verification and Local Repair for LLM Task Planning in Embodied Agent |
|
| Bui, Trung Minh | Korea Electronics Technology Institute |
| Moon, JongSul | Korea Electronics Technology Institute |
| Kim, YoungOuk | Korea Electronics Technology Institute |
| Shin, Dongin | Korea Electronics Technology Institute |
Keywords: Artificial Intelligence Systems, Robotic Applications, Human-Robot Interaction
Abstract: : Large language models are now the default cognitive core of embodied household agents, yet high taskcompletion rates are typically sustained by execution-time retry loops that paper over plans failing symbolic precondition checks. We present GRACE (Grounded Repair-and-Critique Embodied planner), a single-LLM planner that gates each sub-plan through a rule-based symbolic precondition verifier (∼250 lines of Python, zero tokens) and, on rejection, regenerates only the suffix of the failed sub-goal. We evaluate on the 38-task AI2-THOR benchmark across two open-weight Ollama models under a leak-free leave-one-out retrieval protocol, necessary because the seed memory is built from the same file used for evaluation: letting it return a test task’s own reference plan inflates completeness by 10–14 pp. Under the corrected protocol the verifier’s contribution is plan validity rather than coverage: strict-precondition satisfaction rises +16 pp over the one-shot baseline on llama3.2 (0.76 → 0.92) and +13 pp on qwen2.5:7b (0.83 → 0.96), while completeness only matches the strongest baselines, at 7–8× the token cost. We also diagnose two benchmark-protocol failure modes — a refine-overwrite bug and a cold-start emission instability — and give minimal mitigations.
|
| |
| FrAT8 |
Symphony C, 4F |
| Robust Perception and Adaptive Control Systems for Industrial Intelligence |
Oral Session |
| Organizer: Yun, Jong Pil | Korea Institute of Industrial Technology |
| Organizer: Kim, Min Su | KITECH |
| |
| 09:00-09:15, Paper FrAT8.1 | |
| Automated Adaptive Threshold Optimization for Robust Particle Detection in Wafer Inspection (I) |
|
| Park, Jaehyun | Korea Institute of Industrial Technology |
| SeungTaek, Kim | KITECH |
| Yoo, Young-Jun | KITECH |
Keywords: Industrial Applications of Control, Sensors and Signal Processing, Artificial Intelligence Systems
Abstract: Particle detection on semiconductor wafers is essential for ensuring product quality and yield in advanced manufacturing processes; however, conventional threshold-based segmentation methods depend on manually tuned parameters and are highly sensitive to illumination variations and process conditions. While deep learning and generalpurpose segmentation models, such as the Segment Anything Model (SAM), have shown strong performance in generic vision tasks, their reliance on large labeled datasets and limited robustness in detecting small-scale defects restrict their applicability in industrial environments. This study proposes an automated adaptive threshold optimization method for wafer particle and scratch detection that systematically determines optimal threshold parameters by jointly exploiting local and global image characteristics. The proposed approach eliminates the need for manual tuning, pre-trained models, and prompt-based inputs, while maintaining low computational complexity suitable for real-time inspection. Experimental results on wafer defect datasets demonstrate that the proposed method consistently outperforms conventional thresholding techniques and SAM in detecting small particles and elongated defects under varying illumination and manufacturing conditions, providing a practical training-free solution for semiconductor wafer inspection.
|
| |
| 09:15-09:30, Paper FrAT8.2 | |
| EML Symbolic Policies for Embedded Robotic Control: A Compact, Interpretable Alternative to Deep RL for Slip Prevention (I) |
|
| Song, Jisu | Kyungpook National University |
| Lee, Sangmoon | Kyungpook National University |
Keywords: Artificial Intelligence Systems, Robotic Applications, Control Theory and Applications
Abstract: We present a compact, interpretable alternative to deep reinforcement learning for embedded robotic control. Our policies are symbolic expressions evolved as directed-acyclic graphs of the Exp-Minus-Log (EML) operator, eml(x, y) = exp(x) − ln(y), which is universal enough to represent all elementary functions but parsimonious enough to encode a useful control law in a few hundred bytes. On two control tasks — CartPole-v1 and a four-finger slip-prevention task in a MuJoCo simulator — an EML policy matches or strictly dominates a PPO baseline trained under an identical wall-clock budget while being about 8× faster at raw inference and 80× smaller as deployable bytes. On the slip-prevention task an EML policy of 301 bytes achieves 43.8% success across an 8 × 4 mass–friction grid versus 28.1% for a PPO MLP of 24 KB, winning five scenarios and losing none. The learned policy is a single one-line formula that uses exactly four vibration features — one per finger chain — making the inductive bias of the task explicit in a way the PPO MLP is not.
|
| |
| 09:30-09:45, Paper FrAT8.3 | |
| OLIVE: Lightweight Latent Residual Reinforcement Learning for Vision-Language-Action Models Via Action Basis Projection (I) |
|
| Kim, Gibeom | Kyungpook National University |
| Lee, Sangmoon | Kyungpook National University |
Keywords: Artificial Intelligence Systems, Robotic Applications, Robot Mechanism and Control
Abstract: Vision-Language-Action (VLA) models possess powerful pre-trained representations for robotic manipulation; however, online reinforcement learning (RL) adaptation to new environments is computationally prohibitive, as it requires updating billions of parameters. In this paper, we propose OLIVE (Online Latent In-situ VLA Enhancement), a framework that performs PPO-based adaptation by injecting latent residuals into the hidden state of a 7B-scale OpenVLA backbone, while keeping the backbone entirely frozen. Specifically, a 256-dim latent code z is projected through a basis matrix constructed from the action token rows of the frozen LM head, generating a 4096-dim Delta h that is added to the backbone hidden state. This design constrains all perturbations to lie within the semantic action subspace already encoded by the language model. OLIVE trains only 300K parameters (0.004% of 7.3B total), and achieves 91.9% success rate (SR) on LIBERO-Spatial (+31.9pp over baseline), 71.4% on LIBERO-Goal (+1.4pp), and 50.0% on LIBERO-10 (+10.0pp). This work presents a novel direction for lightweight online VLA adaptation by repurposing the pre-trained structural knowledge of frozen LM heads for RL head design.
|
| |
| 09:45-10:00, Paper FrAT8.4 | |
| Beyond Condition-Mixed Evaluation: Label-Free Bearing Fault Diagnosis Via Synthetic Prototype Inference (I) |
|
| Kim, Dongju | Korea Institute of Industrial Technology |
| Yoo, Joonhyuk | Daegu University |
| Won, Hong-In | Korea Institute of Industrial Technology |
| Yun, Jun-Seok | Korea Institute of Industrial Technology |
| Yun, Jong Pil | Korea Institute of Industrial Technology |
| Kim, Min Su | KITECH |
Keywords: Artificial Intelligence Systems, Sensors and Signal Processing
Abstract: Acquiring labeled fault data across diverse operating conditions is costly and often impractical in industrial settings, as fault annotation requires specialized expertise. Against this background, two problems compound each other: models are typically evaluated under condition-mixed splits that share operating conditions between train and test, systematically overestimating generalization; and inference pipelines assume the availability of real labeled samples at deployment. We address both problems. Under a strict leave-one-condition-out (LOCO) protocol with multi-seed validation (N = 9 trials) on Paderborn, simple baselines scoring 0.81–0.82 under mixed evaluation drop by 9–13 points, and DANN-style adversarial training collapses to near-chance. Holding the learned representation fixed and varying only the classifier, we show that a label-free synthetic prototype — averaged over CA-Synth samples generated across all training conditions — significantly surpasses the labeled real-prototype classifier (Δ = +0.066, p = 0.0005; 9/9 trials in favor), matches RF/KNN/LogReg on average (p > 0.13), and achieves the largest advantage on the hardest unseen condition (+0.067 over RF). We attribute this to a training-condition bias-reduction effect: in cross-condition regimes, the structure of supervision matters more than its quantity. In practice, these results suggest that fault diagnosis can generalize to newly encountered operating conditions without requiring condition-specific data collection or labeling, thereby reducing the re-annotation burden as condition diversity increases.
|
| |
| 10:00-10:15, Paper FrAT8.5 | |
| Solar Panel Pose Estimation Using Stereo Vision and Time-Of-Flight Sensors (I) |
|
| Lee, MyeongWoo | Pukyong National Universitiy |
| Jeong, Tae Young | Pukyong National University |
| Joo, Moon Gab | Pknu |
Keywords: Robot Vision, Sensors and Signal Processing, Robotic Applications
Abstract: This paper presents a robotic vision framework for estimating the five-degree-of-freedom (5-DOF) pose of photovoltaic solar panels for automated installation and grasping tasks. The proposed system uses an Intel RealSense D435 stereo depth camera and a Synexens CS30 time-of-flight (ToF) depth camera as two independent depth sources, each of which fails under different conditions caused by specular reflection from the panel glass surface and by outdoor illumination. A YOLO-based instance segmentation network first localizes the individual photovoltaic cells in the RGB image, and the corresponding point cloud region of each depth sensor is passed to RANSAC plane fitting to obtain two independent 5-DOF pose estimates, one from each camera, together with their associated uncertainty. These two depth-based estimates are then combined using three sensor-fusion strategies, which are implemented and compared under the same test conditions: a random-walk Kalman filter that exploits temporal continuity across frames, a Covariance Intersection (CI) fusion that remains consistent when the cross-correlation between the two sensors errors is unknown, and a closed-form inverse-variance weighted fusion that assumes the two sensors errors are independent. The framework is intended to serve as the perception front end of an excavator-boom-mounted manipulator for unattended solar panel installation.
|
| |
| 10:15-10:30, Paper FrAT8.6 | |
| Onboard Inference-Budget Analysis of Deep Reinforcement Learning for Catheter Cross-Section Alignment (I) |
|
| Hwang, Jae Jin | Pukyong National University (PKNU) |
| Kim, Yejin | Korea Institute of Industrial Technology |
| Yun, Jong Pil | Korea Institute of Industrial Technology |
| Joo, Moon Gab | Pknu |
Keywords: Artificial Intelligence Systems, Biomedical Instruments and Systems
Abstract: Automated inspection of multi-lumen catheter cross-sections requires a canonical orientation, so rotational alignment is a necessary front-end step. On a constrained on-device system its cost depends not only on the latency of one evaluation but also on how many the policy spends before committing. We derive an oracle minimum action count by reverse breadth-first search on the residual-angle grid, then score six agents that pair DQN, DDQN, and D3QN with ε-greedy or NoisyNet exploration over 480 rollouts. Even the lowest-budget agent spends 55% more inferences than this minimum, so inference count is a deployment metric alongside model size.
|
| |
| 10:15-10:30, Paper FrAT8.7 | |
| Development of a Vision-Based Recognition and Tracking System for Moving Yard Tractor Identification in Port Environments (I) |
|
| Choi, Eunji | Pukyong National University |
| Joo, Moon Gab | Pknu |
Keywords: Robot Vision, Artificial Intelligence Systems, Sensors and Signal Processing
Abstract: In this paper, we propose a video-based recognition and tracking framework to reliably identify the identification numbers of yard tractors in a port environment. Conventional license plate recognition systems target standardized plates, presenting limitations when applied to unstandardized yard tractor numbers. To address this, the proposed method maintains temporal information using object detection and multi-object tracking, applying OCR, candidate extraction, and time-axis Frame Voting. Furthermore, Track Merging using OCR information and temporal connectivity is introduced to mitigate single-frame OCR errors and ID Switching. Experimental results using actual port video data demonstrated that 129 out of 140 yard tractors were correctly recognized, achieving an accuracy of 92.14%. The proposed method is confirmed to be effective for stable yard tractor identification in moving and low-resolution environments.
|
| |
| FrAT9 |
Symphony D, 4F |
| Autonomous Vehicle Systems 3 |
Oral Session |
| |
| 09:00-09:15, Paper FrAT9.1 | |
| Development of an Autonomous Mobile Robot for Both Indoor and Outdoor Environments |
|
| Fushimi, Yuki | Shonan Institute of Technology |
| Yuzawa, Satoshi | Shonan Institute of Technology |
| Inoue, Fumihiro | Shonan Institute of Technology |
| Ohno, Hidetaka | Shonan Institute of Technology |
Keywords: Autonomous Vehicle Systems, Navigation, Guidance and Control, Robotic Applications
Abstract: For the purpose of autonomous navigation using LiDAR SLAM, a robot system was designed and built for operating in both indoor and outdoor environments using a common algorithm by applying light-resistant LiDAR. By performing 3D range-finding and adding 9-axis IMU measurements, the accuracy of displacement measurement was improved, demonstrating that map construction and autonomous navigation can be performed with sufficient accuracy. As a result, the realization of autonomous navigation was verified on sloped surfaces, slightly uneven paved surfaces, tiled/blocked surfaces, and pathways with steps, demonstrating its effectiveness particularly outdoors. With this robot system, maps of continuous routes can be generated across indoors and outdoors. As a demonstration of continuous autonomous navigation, it is possible to travel along a route that passes through building exits. For travel through building exits, it is necessary to detect transparent glass doors and slopes to avoid steps as adjustments, and a solution of adding this information to the map data was presented.
|
| |
| 09:15-09:30, Paper FrAT9.2 | |
| A Two-Loop Finite-State-Machine Controller with Bounded-Retry Verification for Resource-Constrained Autonomous Marine Debris Investigation |
|
| Rajagopalan, Advaith | BASIS Independent Silicon Valley |
Keywords: Autonomous Vehicle Systems, Navigation, Guidance and Control, Robotic Applications
Abstract: Persistent marine debris monitoring requires au- tonomous surface vehicles (ASVs) that detect, investigate, and validate floating targets autonomously. ASV autonomy often takes it for granted that perception operates reliably; however, a resource-constrained ASV that stages a thermal proposal before an expensive vision-based confirmation must deal with the possibility of ambiguous detections. We introduce a two-loop finite-state-machine (FSM) controller—an exploration loop to cover an area and an investigation loop to validate a detection— where the key idea is to bound the number of rotational retries in verification. Such a verification stage allows for distinguishing be- tween confirmation (DONE) and verification rejection (ABORT) while giving an ambiguous detection a chance of retry. Rather than fixing design parameters arbitrarily, we explore their effect through an experimental voting threshold sweep, an analysis of the failure mode of ABORT outcomes from 46 trials of the FSM controller in a dry run, and a coverage-efficiency framework. The dry run yields 35 DONE (76%) and 11 ABORT (24%) outcomes, demonstrating simulation-level correctness of the FSM logic; we show that every one of the 11 observed ABORTs stems from a verification rejection rather than a navigation or state-transition failure. On-water mission validation is the immediate next step of the work.
|
| |
| 09:30-09:45, Paper FrAT9.3 | |
| Direction-Decomposed Force and Sound Evaluation of Side Bumpers for Door Operation by a Powered Wheelchair: A Low-Cost Rigid Bumper versus a Roller and a Tensioned Belt |
|
| Ida, Yosuke | University of Tsukuba |
| Date, Hisashi | University of Tsukuba |
Keywords: Autonomous Vehicle Systems, Rehabilitation Robot, Robotic Applications
Abstract: Driving a powered wheelchair through a doorway often brings the side of the chair into contact with the door or frame. Such contact produces two distinct force components: a normal (pressing) force that governs the crushing and pinching risk, and a tangential (backward) force that opposes forward motion. Previous rolling-element accessories were designed mainly to reduce the tangential component. In this paper we evaluate four side bumpers --- a 3D-printed roller representing the conventional rolling element, a low-cost rigid bumper made of a sill-slide seal on a paulownia bar, a flat belt on cam followers, and the same belt under tension --- by decomposing the contact force measured with a four-axis load cell into normal and tangential components, and by also recording the contact sound. Four door types (outward, heavy outward, inward, and sliding) were tested on a WHILL CR2 at 0.1--0.5~m/s. The rigid wood bumper produced the lowest normal force in every condition (median peak about 1.3~kgf, roughly one third of the roller) and was acoustically quiet, while the roller did not reduce the normal force and was the loudest. The tensioned belt was clearly better than the roller, but its prototype showed larger trial-to-trial scatter. The results indicate that a simple, low-cost rigid bumper is an effective alternative to the conventional roller for reducing the safety-relevant contact force.
|
| |
| 09:45-10:00, Paper FrAT9.4 | |
| Nonlinear Opinion Dynamics for Fast Collective Decision-Making in Multi-Robot Systems |
|
| Sung, Eunwoo | KAIST |
| Kim, Sohyun | Korea Advanced Institute of Science and Technology (KAIST) |
| Shin, Hyo-Sang | KAIST |
Keywords: Control Theory and Applications, Robotic Applications, Autonomous Vehicle Systems
Abstract: In this paper, we study fast collective decision making in multi-robot systems, where a team must quickly commit to a common intent among a small number of discrete maneuver choices. Examples include a swarm deciding which side of an obstacle to pass and vehicles deciding who yields and who proceeds at a shared junction. In such settings, a decision layer that is slow, weak, or prone to flip-flopping can undermine the mission regardless of the quality of the low-level controller. We present a decentralized decision layer based on nonlinear opinion dynamics (NOD), in which cooperative inter-agent coupling shares local evidence and an attention parameter controls when weak preferences are amplified into a committed collective choice. The main contrast is with well-established linear consensus; it can make agents agree, but when the shared evidence is weak it converges to a low-magnitude opinion that remains below a decision threshold. NOD, by contrast, uses an attention-controlled bifurcation to turn the same weak bias into a fast threshold-crossing commitment. A deterministic simulation shows that cooperative NOD reaches a common committed decision under weak initial evidence, while linear consensus reaches agreement without commitment. An attention sweep further shows that the decision time decreases as attention increases and follows the trend predicted by the linearized dynamics.
|
| |
| 10:00-10:15, Paper FrAT9.5 | |
| Real-Time Unsupervised Probabilistic Magnetic Disturbance Detection for Magnetometer-Based Heading Estimation |
|
| Park, Jiwon | Chungnam National University, Department of Autonomous Vehicle Systems Engineering |
| Jung, Jongdae | Chungnam National University |
Keywords: Artificial Intelligence Systems, Autonomous Vehicle Systems
Abstract: Unmanned ground vehicles (UGVs) often use magnetometers as an absolute heading reference. However, magnetometer measurements are vulnerable to corruption by magnetic disturbances from nearby ferromagnetic structures and onboard electromechanical systems. Conventional threshold-based magnetic anomaly detectors such as the three-parameter AND-gate method of Lee et al. require platform-specific manual calibration. This paper proposes an unsupervised probabilistic magnetic disturbance detector based on a Temporal Convolutional Network with a Mixture Density Network (TCN-MDN) output head. The detector is trained only on undisturbed data, allowing the model to learn the temporal probability density of a four-dimensional feature sequence consisting of the latency-compensated cross-sensor angular velocity residual, horizontal magnetic field magnitude, lagged gyroscope angular velocity, and horizontal field magnitude derivative. During inference, the negative log-likelihood (NLL) under the learned Gaussian mixture is used as a continuous disturbance score. When the NLL exceeds a data-driven threshold, magnetometer-derived heading updates are rejected and heading is propagated by gyroscope integration. The system is deployed with TensorRT on a Jetson Orin Nano and operates in real-time at 100 Hz on a Clearpath Husky A300 UGV. Experimental results show reduced heading RMSE and peak heading error compared with the Lee et al. baseline.
|
| |
| FrBT1 |
Room T1 |
| Control Theory and Applications 4 |
Oral Session |
| |
| 15:40-15:55, Paper FrBT1.1 | |
| Finite-Time Spectral Observer |
|
| Murakami, Madoka | Tokyo University of Science |
| Kitamura, Tomoya | Tokyo University of Science |
| Nakamura, Hisakazu | Tokyo University of Science |
Keywords: Control Theory and Applications, Sensors and Signal Processing
Abstract: Estimation of sinusoidal signals is fundamental in engineering, and finite-time observer-based methods have attracted attention owing to their fast convergence and robustness. However, homogeneous finite-time spectral observers have mainly been developed for DC and single-frequency components, while their extension to multiple sinusoidal components remains insufficient. This paper proposes a finite-time spectral observer (FTSpO) for estimating the amplitudes and phases of multi-sinusoidal signals with known frequencies. An explicit nonsingular similarity transformation is derived to convert the original state-space model into an observable canonical-form realization with a left-companion matrix, enabling direct application of a simple finite-time observer structure. Numerical simulations show that, even for a quasi-periodic signal, the proposed method estimates the output, amplitudes, and phases in finite time. In particular, convergence is achieved within 300 ms for quasi-periodic signals with a minimum frequency of approximately 1.41 Hz, the square root of 2.
|
| |
| 15:55-16:10, Paper FrBT1.2 | |
| Real-Time Distance Estimation Method of Sudden Sound Source with Two Microphones |
|
| Okubo, Shiori | Meiji University |
| Endo, Riku | Aoyama Gakuin University |
| Itami, Taku | Meiji University |
Keywords: Multimedia Systems, Biomedical Instruments and Systems, Sensors and Signal Processing
Abstract: Globally, the prevalence of hearing loss is increasing, and ensuring that individuals with hearing impairment can detect sudden warning sounds―such as warning sirens of emergency vehicles or bicycle bells―is critical for their daily safety. Conventional methods for estimating sound source distance typically rely on large microphone arrays or deep learning techniques. However, these approaches suffer from substantial computational delays, a lack of real-time responsiveness, and excessive hardware resource requirements, making them impractical for wearable devices like smart glasses. To address these challenges, this paper proposes a novel, real-time distance estimation method for sudden sound sources using only two microphones without requiring machine learning. The proposed method utilizes the differences in decibel (dB) values between the two microphones to calculate the distance based on geometric acoustics. Furthermore, a three-step filtering algorithm is introduced to remove decibel overshoots and environmental noises caused by sudden acoustic signals. Experimental validations demonstrate that the proposed method significantly reduces estimation errors, achieving a 71% success rate within an error margin of 0.3 meters while maintaining real-time processing performance. These results confirm the effectiveness and feasibility of the method for low-resource smart glass applications.
|
| |
| 16:10-16:25, Paper FrBT1.3 | |
| Pressure-Dependent Degradation Trend Modeling and RUL Estimation for Hinge Systems |
|
| Chiou, Huei-Cyuan | National Taipei University of Technology |
| Wu, Hsiu-Ming | Department of Intelligent Automation Engineering, National Taipei University of Technology |
Keywords: Artificial Intelligence Systems
Abstract: Machine learning and deep learning techniques have been widely applied in prognostics and health management (PHM) for degradation trend prediction (DTP) and remaining useful life (RUL) estimation. In addition to general machinery, accurate prediction is also critical for mechanical components such as laptop hinges. Hinge degradation is typically evaluated based on torque reduction relative to a predefined health indicator (HI) threshold. During the design-for-manufacturing (DFM) stage, conventional approaches for evaluating hinge torque adequacy rely heavily on empirical experience, which may lead to inaccurate assessments and unexpected failure when degradation exceeds the HI threshold. To address this issue, this study proposes a data-driven framework termed pressure value-marked bidirectional gated recurrent unit (PM-BiGRU). The proposed method incorporates the pressure value per unit friction surface area (pressure value) to characterize degradation under various loading conditions. A sequence-to-one (Seq2One) BiGRU architecture is adopted to model temporal dependencies and predict single-step degradation, providing stable and interpretable predictions. Experimental results based on hinge torque degradation datasets demonstrate that the proposed method effectively captures degradation trends and accurately estimates the corresponding RUL under varying pressure conditions, providing a robust evaluation tool for the DFM stage.
|
| |
| 16:25-16:40, Paper FrBT1.4 | |
| LocalVLM-TPR: Local Vision-Language Model-Based Task Planning and Re-Planning for Robotic Manipulation |
|
| Wang, Shih-Ting | National Chung Hsing University |
| Chao, Yin-Fan | National Chung Hsing University |
| Chen, Yi-Fan | National Chung Hsing University |
| Li, I-Hsum | National Chung Hsing University |
Keywords: Human-Robot Interaction, Robotic Applications
Abstract: Tabletop robotic tasks often require visual understanding of changing object locations, spatial relationships, and execution states, while cloud-based large models may introduce API cost, network dependence, and privacy concerns. This paper presents LocalVLM-TPR, a locally deployed vision-language model framework for robotic task planning and re-planning. The system combines eye-to-hand and eye-in-hand vision to support scene perception, task planning, execution validation, and failure recovery. A logic audit is used to check plan feasibility, while pre-action, post-action, and final-goal validation form a closed-loop execution process. When a local failure occurs, plug-in re-planning generates a corrective subtask. Experiments on three tabletop tasks show that LocalVLM-TPR achieves a 93.33% average success rate, close to the cloud-based ReplanVLM baseline with 96.67%, while avoiding cloud-based inference.
|
| |
| 16:40-16:55, Paper FrBT1.5 | |
| Toward Extending Dynamical Network Biomarker-Based Pre-Disease State Detection Using a Temporal Filter Gate |
|
| Nakao, Yoshiki | Kyushu Institute of Technology |
| Nakakuki, Takashi | Kyushu Institute of Technology |
|
|
| |
| 16:55-17:10, Paper FrBT1.6 | |
| Abkowitz-Based Gray-Box Identification and Prediction Analysis Using an IEC 62065-Based Container Ship Simulator |
|
| Kyung Bae, Lee | Samsung Heavy Industries Co., Ltd |
| Seongyeop, Choi | Samsung Heavy Industries Co., Ltd |
| Seongjun, Kim | Samsung Heavy Industries Co., Ltd |
Keywords: Control Theory and Applications
Abstract: Maneuvering prediction of container ships is important for the development and evaluation of model-based navigation and collision-avoidance algorithms. However, constructing an accurate maneuvering model is difficult when detailed hydrodynamic coefficients and maneuvering data are limited. This study investigates the applicability of an Abkowitz-based gray-box identification approach to container ship maneuvering prediction using limited maneuvering data. The model parameters are identified using acceleration/deceleration, turning-circle, and zig-zag maneuvering data generated from an IEC 62065-based container ship simulation model under noise-free and disturbance-free conditions. The trajectory-prediction performance of the identified model is examined through maneuvering simulations, including additional test cases not used in identification. In addition, the prediction range of the identified model is discussed by analyzing the accumulation of trajectory errors under repeated rudder actions and heading changes.
|
| |
| FrBT2 |
Cattleya, 3F |
| Industrial Applications of Control |
Oral Session |
| |
| 15:40-15:55, Paper FrBT2.1 | |
| Grid-Voltage-Sensorless State and Parameter Estimation for PFC Converters Using an Adaptive ESO |
|
| Park, Keunhoon | Hanyang University |
| Lee, Youngwoo | Hanyang University ERICA |
Keywords: Industrial Applications of Control, Autonomous Vehicle Systems
Abstract: Reliable state estimation and control performance in power factor correction (PFC) converter systems are frequently degraded by grid voltage fluctuations and parametric uncertainties. This article develops an extended state observer (ESO)-based adaptive framework capable of jointly reconstructing physical states and identifying uncertain hardware parameters without direct grid voltage sensing. The proposed ESO-based adaptive observer consists of an ESO that treats the grid voltage as an extended state variable and an adaptive update law that tracks parametric uncertainties. Simulations implemented in MATLAB/Simulink are conducted for validation. Simulation results show the proposed ESO-based adaptive observer significantly improves the accuracy of state and parameter estimation.
|
| |
| 15:55-16:10, Paper FrBT2.2 | |
| Sensorless Battery Temperature Estimation and Predictive Thermal Management Using EKF and MPC |
|
| Baek, Hyeonmin | Korea Advanced Institute of Science and Technology(KAIST) |
| Jang, Byeonggwan | KAIST |
| Lee, Sieun | KAIST |
| Kim, Kyung-Soo | KAIST(Korea Advanced Institute of Science and Technology) |
Keywords: Industrial Applications of Control, Control Theory and Applications
Abstract: Lithium-ion batteries are highly sensitive to temperature variations, and excessive temperature rise can significantly degrade battery performance, lifetime, and safety. Although temperature sensors are commonly used in battery thermal management systems, direct temperature measurement is difficult in densely packed battery modules due to cost, installation limitations, and reliability concerns. To address this issue, this paper proposes a sensorless battery thermal management framework combining Extended Kalman Filter (EKF)-based temperature estimation and Model Predictive Control (MPC)-based cooling control. A lumped thermal model and a first-order equivalent circuit model were integrated to estimate battery temperature using only current and voltage measurements. In addition, observability analysis was performed to verify the feasibility of temperature estimation using measurable electrical signals. Based on the estimated temperature, MPC was employed to predict future thermal behavior and determine the optimal cooling intensity while minimizing unnecessary cooling energy consumption. The proposed framework was implemented in MATLAB/Simulink and validated under multiple discharge conditions. Simulation results demonstrated that the EKF could accurately estimate battery temperature despite model uncertainty and measurement noise. Furthermore, the MPC-based cooling strategy effectively reduced cooling energy consumption while maintaining acceptable thermal performance.
|
| |
| 16:10-16:25, Paper FrBT2.3 | |
| Scenario-Based Q-Learning Model Predictive Control for Green Ammonia Production |
|
| Park, Hyun Min | Ulsan National Institute of Science and Technology |
| Oh, Tae Hoon | UNIST |
Keywords: Industrial Applications of Control, Process Control Systems, Artificial Intelligence Systems
Abstract: Green ammonia production systems powered by intermittent renewable energy must meet periodic demand under tight unit and storage constraints. To control such a system effectively, we propose scenario-based Q-learning model predictive control (scQMPC), which augments a scenario-based stochastic model predictive control with a pre-trained Q-function as the terminal cost. The stochastic model predictive control component enforces hard state bounds within the prediction horizon and rejects short-term disturbances through a scenario fan sampled from a transition probability matrix, while the Q-function captures long-term stochasticity that extends beyond the prediction horizon. Simulation results on a 12-day rollout show that the proposed method outperforms nonlinear model predictive control, double deep Q-network, and deterministic Q-learning-based model predictive control baselines, achieving the lowest total cost and eliminating both ammonia tank overflow and periodic demand shortfall.
|
| |
| 16:25-16:40, Paper FrBT2.4 | |
| Data-Driven Parameter Tuning Framework for Enhancing Accuracy of Refrigeration System Model |
|
| Cho, Seung Hyun | Seoul National University |
| Byun, Jisung | Seoul National University |
| Kim, Chan | Seoul National University |
| Lee, Jongyeop | Seoul National University |
| Lee, Jong Min | Seoul National University |
Keywords: Process Control Systems, Industrial Applications of Control
Abstract: This study investigates data-driven parameter tuning for improving the accuracy of physics-based simulation models of refrigeration systems. A physics-based model of the overall refrigeration system is developed in Dymola to describe the refrigeration process under varying operating conditions. To enhance the reliability and prediction accuracy of the model, parameter tuning strategies are explored by combining operational data with a simulation-based parameter estimation process. The tuning process focuses on identifying key parameters that significantly affect system performance and adjusting them to reduce discrepancies between simulation results and real operation data. In this study, pressure-drop parameters in the cooling-water network are selected and estimated using a reduced flow-network model. Candidate parameters are reduced through an identifiability analysis, and the selected parameters are estimated using differential evolution. Fixed and online parameter update strategies are compared using steady-state operation data. The results show that the online update strategy improves the prediction accuracy of the total cooling-water flow rate under time-varying operating conditions. The proposed framework provides a reliable basis for data-assisted calibration of physics-based refrigeration system models.
|
| |
| 16:40-16:55, Paper FrBT2.5 | |
| Physics-Informed Relative Temperature Estimation for Battery Management Systems Via Lightweight GRU |
|
| Kim, Youngseoung | Hyundai MOBIS |
| Lee, Hongjun | Hyundai MOBIS |
| Oh, Seungyeob | Hyundai MOBIS |
| Yang, Siwon | Hyundai MOBIS |
| Yoo, Younguk | Hyundai MOBIS |
| Byun, Jinsu | Hyundai MOBIS |
| Kim, Dongrak | Hyundai MOBIS |
Keywords: Industrial Applications of Control, Control Theory and Applications
Abstract: Accurate cell-level thermal monitoring is essential for battery safety, power capability, and lifetime, but full sensor deployment is impractical in production battery packs. This paper proposes a sparse-sensing cell temperature estimator that anchors a compact GRU to a production-available reference temperature sensor. Instead of directly regressing absolute cell temperatures, the GRU estimates a reference-sensor-based relative-temperature vector, and the absolute cell temperatures are reconstructed using the measured reference temperature. A lumped thermal analysis shows that the ideal relative-temperature target corresponds to a deviation-subspace coordinate after removing the dominant common-mode component. The practical measurement anchoring is represented as a bounded additive term, and the resulting reduced deviation dynamics provide an input-to-state stability condition under bounded perturbations. The proposed method is validated on an automotive PHEV battery pack using vehicle driving profiles under active-cooling conditions at 10/10◦C, 25/25◦C, and 35/35◦C ambient/coolant temperatures, including a cooling ON-to-OFF transient and cases with intramodule temperature deviations exceeding 8◦C. Compared with an identical-capacity absolute-temperature baseline, the proposed estimator achieves sub-degree RMSE and reduces the maximum error from 4.17◦C to 1.82◦C. The trained model is also deployed on an Infineon AURIXTM TC387D automotive microcontroller, where inference completes within 2 ms per module and the onboard output matches the offline FP32 reference within 0.07◦C MAE.
|
| |
| FrBT3 |
Azalea, 3F |
| Robot Vision 2 |
Oral Session |
| |
| 15:40-15:55, Paper FrBT3.1 | |
| Vision-Based 2D Scan Generation for Obstacle Avoidance Using Floor Segmentation |
|
| Lee, Hyunseo | Handong Global University |
| Yoo, Gunmin | Handong Global University |
| Gu, Hyunwoo | Handong Global University |
| Hwang, Sung Soo | Handong Global University |
Keywords: Robot Vision, Navigation, Guidance and Control, Robotic Applications
Abstract: This extended abstract presents a monocular virtual scan generation method for indoor autonomous mobile robots. The proposed system segments floor regions from RGB camera images and converts non-floor pixels into approximate 2D scan observations through camera-calibrated lookup tables. The generated scan is published in a ROS-compatible range format and integrated with the Nav2 navigation stack for local obstacle avoidance. Experiments in a structured indoor corridor show that the method provides useful, but limited, obstacle cues without a dedicated LiDAR sensor. In controlled range-estimation trials, the absolute error remained within 0.261 m under static conditions. In a 10 m static frontal obstacle avoidance experiment, the robot completed 39 of 50 trials. These results suggest that monocular scan generation can be a lightweight sensing option for constrained indoor environments, while robustness under reflective floors, illumination changes, and broader obstacle conditions remains future work.
|
| |
| 15:55-16:10, Paper FrBT3.2 | |
| YOLO-Based Chilli Pepper Detection and Classification in a Cloud-Assisted Robotic Harvesting System |
|
| Shaybo, Ahmed Sharif | Southeast University |
| Liang, Han | Southeast University |
| Baig, Mohammad Abbas | Southeast University, China |
| SalahEldeen Mohammed Hassan, Almogdad | Southeast University |
| Sibtain, Muhammad | Southeast University, China |
| Qasem, Arfan Ali Mohammed | Southeast University |
| Jibreel, Alnoman Bashir Abdalla | Southeast University |
| Sann, Channvechhika | Southeast University |
Keywords: Robot Vision, Robotic Applications
Abstract: The growing global population has increased demand for automated agricultural harvesting systems to support higher food production. Although robotic harvesting offers a promising solution, many existing systems remain costly, computationally demanding, and operationally inefficient because a single robot must often perform plant-data acquisition, processing, detection, localization, and harvesting. Moreover, real-time onboard processing requires high-performance hardware, further increasing system cost and complexity. This paper proposes a cloud-assisted multi-robot system for chilli pepper harvesting that enables information sharing between a Data-Collection Robot (DCR) and Harvesting Robots (HARs). A cloud server processes, stores, and manages the collected data while coordinating inter-robot operations. The detection model was trained and evaluated using chilli pepper images captured under varied conditions, achieving reliable detection, classification, and localization performance. The system was validated in a simulated tabletop chilli pepper laboratory environment. During the experiment, the DCR captured plant data using a depth camera and transmitted it to the cloud server for processing and storage within a few seconds. The results demonstrate the feasibility, scalability, efficiency, and cost-effectiveness of the proposed cloud-assisted harvesting system.
|
| |
| 16:10-16:25, Paper FrBT3.3 | |
| OSDAG: Online Scheduling for Efficient Multi-Robot Collaboration |
|
| Nguyen Canh, Thanh | Japan Advanced Institute of Science and Technology |
| Tran Viet, Thang | Vietnam National University, Hanoi, University of Engineering and Technology |
| Dinh Van, Phuc | VNU University of Engineering and Technology |
| Hoang, Van Xiem | VNU - University of Engineering and Technology |
| Chong, Nak Young | Japan Advanced Institute of Science and Technology |
Keywords: Robot Vision, Robotic Applications, Artificial Intelligence Systems
Abstract: Coordinating heterogeneous multi-robot systems (MRS) for complex, long-horizon tasks requires both flexible high-level reasoning and efficient execution-time scheduling. Existing LLM-based approaches struggle to balance reasoning efficiency and execution flexibility. Flat sequential plans are efficient to generate but overlook parallel execution opportunities, while repeated LLM reasoning introduces high latency, and offline schedules may unnecessarily keep robots idle due to fixed execution orders. This paper presents OSDAG, a novel framework that resolves this trade-off by employing a Directed Acyclic Graph (DAG) as the central representation for multi-robot coordination, coupled with constraint-aware online scheduling. The LLM is typically invoked once as a semantic parser that decomposes a natural-language instruction into a dependency-annotated task graph encoding precedence relations, together with robot capability and resource-feasibility constraints. A lightweight online scheduler then dynamically dispatches dependency-ready tasks to their assigned robots as soon as they become idle, exposing available parallelism while preserving correctness. Experiments across five benchmark scenarios demonstrate that OSDAG achieves 5-15times faster reasoning time than dialogue-based methods, reduces makespan by up to 38% over sequential baselines, and maintains competitive success rates. Both simulation and real-world experiments on human-robot collaboration tasks validate the effectiveness and practicality of the proposed approach for efficient multi-robot coordination. The website and resources are available at url{http://thanhnguyencanh.github.io/LLM_DAG4MultiRobot}
|
| |
| 16:25-16:40, Paper FrBT3.4 | |
| Impact of Dataset Composition on Embedded Real-Time UAV Wildfire Detection Using Compact YOLO Models |
|
| de los Santos, Eduardo | Technological University of Uruguay |
| Kelbouscas, André | FURG |
| Grando, Ricardo | Federal University of Rio Grande |
| Guterres, Bruna de Vargas | Universidad Technologica Del Uruguay |
Keywords: Robot Vision, Robotic Applications, Autonomous Vehicle Systems
Abstract: The development of vision-based wildfire detection systems for unmanned aerial vehicles is constrained by the limited availability of diverse real-world training images. This paper investigates the impact of dataset composition on embedded real-time UAV wildfire detection using compact YOLO models as a controlled validation family. Four training configurations were evaluated: real non-augmented, real augmented, hybrid non-augmented, and hybrid augmented, where the hybrid sets combine real wildfire images with AI-generated samples. The objective is to determine whether synthetic data mixing and image augmentation improve practical detection performance under resource-constrained deployment conditions. Experimental results show that the best overall operating point was obtained with the real non-augmented dataset, which achieved the strongest balance between recall and mean average precision for UAV-based wildfire detection. The results show that hybridization with synthetic data did not improve the final deployment choice, while conventional augmentation also yielded lower aggregate performance under the evaluated test conditions. These findings suggest that domain alignment between training and deployment imagery is an important factor for embedded UAV wildfire detection, while the benefits of augmentation may depend on the specific operating conditions being represented.
|
| |
| 16:40-16:55, Paper FrBT3.5 | |
| Self Supervised Learning from Automatically Generated Demonstrations for Visual Robotic Manipulation |
|
| Rivas, Andres | UTEC |
| Cukla, Anselmo | Federal University of Santa Maria |
| da Silva Guerra, Rodrigo | Universidade Federal Do Rio Grande (FURG) |
| Guterres, Bruna de Vargas | Universidad Technologica Del Uruguay |
| Grando, Ricardo | Federal University of Rio Grande |
Keywords: Robot Vision, Robotic Applications, Industrial Applications of Control
Abstract: Robotic manipulation often requires object specific programming, manual data annotation, or calibrated perception pipelines, which limits rapid deployment in practical settings. Learning from demonstration offers a more direct alternative, but collecting demonstrations can still demand human teleoperation or kinesthetic teaching. This paper presents a self supervised visual manipulation method in which a robot automatically generates demonstrations around a target pose and learns relative pose corrections directly from wrist mounted RGB images. The proposed pipeline uses ROS2 and Isaac Sim to collect labeled image-pose pairs without requiring explicit camera to robot extrinsic calibration. Separate datasets are generated for planar refinement and coarse three dimensional approach, and a convolutional network is trained to regress relative translation and rotation from single frame RGB observations. During execution, a coarse to fine controller first approaches the object using models trained with height variation and then refines the final alignment using planar data. The method is evaluated both in simulation and on a real UR5e collaborative robot equipped with a gripper and a monocular camera. In simulation, the refinement stage reduces the final planar dispersion from 9.69 mm to 5.38 mm. In real world experiments, the system performs end to end grasp attempts on three physical objects and reaches success rates of 66.6% and 63.6% for two objects without object rotation, while still maintaining partial robustness under rotated conditions. These results show that automatically generated demonstrations can support practical visual manipulation with limited setup effort, while also exposing remaining challenges in depth prediction and object dependent generalization.
|
| |
| 16:55-17:10, Paper FrBT3.6 | |
| Lightweight Multi-Scale Knowledge Distillation for 3D Object Detection in Service Robotics |
|
| Park, Jongkuk | LG Electronics |
| Lee, Jaemin | LG Electronics |
| Lee, Byunghoon | LG Electronics |
| Noh, DongKi | LG Electronics Inc |
Keywords: Artificial Intelligence Systems, Robot Vision, Robotic Applications
Abstract: Embedded service robots must perform 3D perception in real time under strict memory and compute budgets. A common way to reduce latency is to downsample the input point cloud, but naive downsampling can remove geometric details that are important for indoor object detection. We propose PoinTeacher, a resource-aware knowledge distillation framework for accelerating PointNet++/VoteNet-style 3D detectors. A multi-resolution teacher equipped with an inputdependent Multi-Scale Feature Fusion (MSFF) module transfers coordinate-aligned features and proposal-level responses to a low-resolution student. During final fine-tuning, a separate teacher with the same architecture as the student is updated by exponential moving average (EMA) and supplies a temporally smoothed consistency target. The EMA teacher is used only during training, so the deployed student incurs no additional branch. On SUN RGB-D, the 5,000-point student reaches 60.1±0.26 mAP@25 and 34.5±0.30 mAP@50 over three seeds while processing 38.1 scenes/s on an NVIDIA Jetson AGX Orin, compared with 15.8 scenes/s for the 20,000-point VoteNet baseline. These results support a practical 2.4× throughput improvement with comparable detection accuracy.
|
| |
| FrBT4 |
Lilac, 3F |
| Robot Mechanism and Control 3 |
Oral Session |
| |
| 15:40-15:55, Paper FrBT4.1 | |
| Robust Position Tracking Control of Robotic Manipulators Using Deep Koopman-Based Equivalent Input Disturbance Estimation |
|
| Nam, Kyung Ho | Kumoh National Institute of Technology |
| Ban, Jaepil | Kumoh National Institute of Technology |
| Koo, Gyogwon | Daegu Gyeongbuk Institute of Science and Technology |
Keywords: Robot Mechanism and Control, Control Theory and Applications, Robotic Applications
Abstract: This paper presents a Koopman-based robust trajectory tracking control of a robotic manipulator. Specifically, a deep Koopman model is identified to approximate the nonlinear robot dynamics as a linear system in a latent observable space. Since the finite-dimensional Koopman model inevitably contains approximation errors and model mismatch, the discrepancy between the trained deep Koopman model and the actual robot dynamics exists. We consider the model mismatch as an equivalent input disturbance (EID), and an extended state observer is then designed by augmenting the Koopman state with the EID. The estimated EID is used to compensate for model mismatch, integrated with a linear quadratic (LQ) optimal state feedback control. The simulation results demonstrate that the proposed observer improves tracking accuracy compared with the conventional Koopman-LQ optimal controller and effectively compensates for both internal model mismatch and external end-effector force disturbances.
|
| |
| 15:55-16:10, Paper FrBT4.2 | |
| Sensorless Load Position Estimation and Balancing Method Based on Proprioception for Ball-On-Plate Bipedal Robot Systems |
|
| Ye, Chengzheng | Zhejiang University |
| Qian, Tianwei | Guilin University of Electronic Technology |
| Dong, Fanrong | Zhejiang University |
| Liu, Shaoxun | Zhejiang University |
| Jing, Hui | Guilin University of Electronic Technology |
| Wang, Rongrong | Zhejiang University |
Keywords: Robot Mechanism and Control, Control Theory and Applications, Sensors and Signal Processing
Abstract: This paper presents a sensorless framework for dynamic load observation and stabilization on bipedal robot system. The proposed framework reconstructs load-related states directly from onboard joint and inertial measurements, eliminating the need for external vision systems or force sensors. A cascaded observer combining a nonlinear disturbance observer and a tracking differentiator is developed to achieve real-time proprioceptive reconstruction of load contact states, position, and velocity. Based on the reconstructed load states, a hybrid control architecture is further designed, where an impedance-based module provides compliant impact buffering and a backstepping sliding mode controller (B-SMC) enables load regulation through posture-based adjustment. The framework is evaluated using a highly dynamic free-moving spherical payload as a representative for unstable loads. Simulation results show that the proposed method achieves reliable contact perception, smooth load-state reconstruction, and stable sensorless balancing under impact disturbances. For a 2kg payload dropped from 2m height, contact is detected within 1ms after impact, and the reconstructed load state converges within 0.05s, validating the effectiveness of the proposed framework.
|
| |
| 16:10-16:25, Paper FrBT4.3 | |
| Multi-Manipulator Coordination for Compressor Manufacturing with Welding State Feedback |
|
| P S, Viamrsh | SASTRA Deemed University |
| S, Shanjay Sundhar | SASTRA Deemed University |
| Krishank, Sri Kamal | SASTRA Deemed University |
| Chakraborty, Arindam | Indian Institute of Technology Kanpur |
| Kumar, Abhishek | Kirloskar Pneumatic Company Limited |
Keywords: Robot Mechanism and Control, Industrial Applications of Control, Robotic Applications
Abstract: This paper presents a multi-robot coordination framework for compressor welding that integrates task allocation, motion planning, and welding-state feedback to improve accuracy, throughput, and process stability. The proposed system uses multiple industrial robots positioned around a rotating compressor workpiece, where each robot is assigned a welding segment based on geometry, reachability, and collision-free workspace constraints. A central coordination layer synchronizes robot motion with turntable rotation, ensuring that the end-effectors maintain proper torch angle, standoff distance, and weld path continuity throughout the process. Welding-state feedback is incorporated through real-time sensing and fusion of arc current, torch position, orientation, and vibration, joint alignment, and seam-tracking signals, enabling adaptive correction during welding. When deviations such as misalignment, gap variation, or arc instability are detected, the controller updates robot trajectories and process parameters to maintain bead quality. The framework also includes safety monitoring, redundancy resolution, and shared workspace management to prevent robot-to-robot interference and protect equipment. Simulation results indicate that welding state feedback-driven coordination reduces path error, improves weld consistency, and shortens cycle time compared with open-loop operation.
|
| |
| 16:25-16:40, Paper FrBT4.4 | |
| DART: Action-Rate-Regularized DAgger Distillation for Recovery-Aware Locomotion of Bipedal Wheeled Robots |
|
| Dong, Fanrong | Zhejiang University |
| Zhou, Shiyu | Shanghai Jiao Tong University |
| Qian, Tianwei | Guilin University of Electronic Technology |
| Ye, Chengzheng | Zhejiang University |
| Liu, Zhengsheng | Zhejiang University |
| Liu, Shaoxun | Zhejiang University |
| Wang, Rongrong | Zhejiang University |
Keywords: Robot Mechanism and Control, Robotic Applications, Artificial Intelligence Systems
Abstract: Bipedal wheeled robots combine efficient locomotion with arbitrary-pose self-recovery, but these skills induce conflicting action distributions and discontinuities in multi-controller systems. This paper presents DART (DAgger-based Action-Rate-Regularized Distillation), a framework that integrates locomotion and recovery into a single proprioceptive policy with action rate regularization. Separate experts are trained for rough-terrain locomotion and recovery, and their labels are aggregated through student-induced DAgger rollouts. The student is optimized with behavioral cloning and an action rate loss that regularizes consecutive predictions near recovery-to-locomotion transitions. In a flat-ground transition evaluation, DART reduced action-change peak, action-change RMS, and torque-rate peak by 4.83%, 2.73%, and 6.22% relative to finite-state expert switching, and improved over DAgger without action rate loss. MuJoCo validation and TRON 1W deployment show that the same student policy executes recovery-aware locomotion without expert querying or runtime switching.
|
| |
| 16:40-16:55, Paper FrBT4.5 | |
| Toward Terrain-Aware Latent Representations for Proprioceptive Quadrupedal Locomotion |
|
| Kim, Taemin | Kookmin University |
| Cho, Baek-Kyu | Kookmin University |
Keywords: Robot Mechanism and Control, Robotic Applications, Artificial Intelligence Systems
Abstract: Robust blind quadrupedal locomotion requires an implicit representation of hidden environment properties, but reconstruction-oriented latents are not guaranteed to be organized by terrain type. This paper investigates whether a DreamWaQ-style latent can be encouraged to become terrain-aware without adding a separate perception module. We compare a baseline implicit latent, full-latent terrain-conditioned prior regularization, and terrain–residual latent separation. We further conduct a matched component ablation to examine the effects of the terrain-conditioned prior, classification loss, property prediction, and latent split on both representation learning and locomotion behavior. Across the evaluated settings, terrain-aware objectives alter the latent distribution but do not consistently produce clear terrain-wise clustering or improved locomotion performance. These results suggest that explicitly enforcing semantic terrain separation may not be sufficient to obtain a terrain-discriminative and reliably policy-useful representation from limited proprioceptive information.
|
| |
| 16:55-17:10, Paper FrBT4.6 | |
| Design and Predictive Extension Control of Extendable Robotic Arms with Time-To-Alignment Prediction |
|
| Kwon, Daehyun | DGIST |
| Song, Yujin | DGIST |
| Park, Junhyun | DGIST |
| Hwang, Minho | DGIST |
Keywords: Robot Mechanism and Control, Robotic Applications, Human-Robot Interaction
Abstract: Wearable robotic systems and supernumerary robotic limbs have been actively studied to augment human motor capability and support assistive manipulation tasks. However, relatively limited research has addressed how to extend the user’s reachable workspace while maintaining a low mechanical burden and natural interaction. This paper presents an extendable robotic arm system designed as a supernumerary assistive limb for human users. The proposed system adopts a scissor-based extension mechanism with a cable-driven structure that relocates heavy actuation components toward the rear of the arm, thereby reducing the user-side gravitational torque predicted by the analytical model. In addition, we introduce a perception-driven extension strategy using a forearm-mounted depth camera. By analyzing temporal changes in the target position in image space, the system predicts the remaining time until the user’s arm and the target become aligned, and initiates arm extension before full alignment is reached. Experimental results show that the proposed method reduces task completion time by 23.1% compared to a conventional alignment-based approach, with grasp success rates of 78% (39/50) and 84% (42/50) for the proposed and conventional methods, respectively. These results demonstrate the feasibility of combining a rear-mounted actuation design for reducing predicted user-side torque with vision-based time-to-alignment prediction for efficient human-assistive robotic interaction.
|
| |
| FrBT5 |
Crown, 3F |
| Exoskeleton Robot |
Oral Session |
| |
| 15:40-15:55, Paper FrBT5.1 | |
| ExoSync: FPGA-Accelerated Terrain-Adaptive Impedance Control for Lower-Limb Exoskeletons with Online Parameter Estimation |
|
| Elsayed, Saher | University of Pennsylvania |
Keywords: Exoskeleton Robot, Control Theory and Applications, Rehabilitation Robot
Abstract: Lower-limb exoskeletons for stroke rehabilitation and augmentation require impedance controllers that adapt joint stiffness, damping, and inertia to terrain in real time, yet current embedded implementations either fix these parameters offline or update them on ARM processors with 40–80 ms latency, exceeding the 15 ms terrain-transition window that prevents stumbling. We present ExoSync, an FPGA co-processor on Zynq-7000 SoC implementing a Terrain-Adaptive Impedance Module (TAIM) that classifies terrain from a lightweight IMU+FSR feature vector in 2.8 ms, updates impedance parameters {K, B, M} via an online gradient-descent estimator in 0.15 ms, and enforces joint-torque safety bounds in combinational logic at 0.04 ms. Impedance parameters converge to terrain-optimal values within 8 strides. Evaluated across 10 subjects (6 stroke survivors, 4 able-bodied) on five terrain conditions (flat, 5°/10° ramp, stair ascent/descent, 1,800+ strides), ExoSync achieves a mean RMS joint-angle tracking error of 3.8° vs. 13.4° for fixed-impedance and 19.1° for pure-PD baselines, at 2.4 W total system power with zero safety violations.
|
| |
| 15:55-16:10, Paper FrBT5.2 | |
| Real-Time Gait Phase Estimation and Ground Reaction Force Prediction Via Knowledge-Distilled Temporal Convolutional Network for Lower-Limb Exoskeleton Control |
|
| Moon, Sunwoong | Gwangju Institute of Science and Technology |
| Hur, Pilwon | Gwangju Institute of Science and Technology |
Keywords: Exoskeleton Robot, Sensors and Signal Processing, Artificial Intelligence Systems
Abstract: Accurate estimation of gait biomechanical variables from limited wearable sensors is essential for adaptive lower-limb exoskeleton control. Simultaneous estimation of continuous gait phase variable and discrete gait events is critical because the stance-to-swing transition timing varies across individuals and even within the same person, and detecting such events typically requires additional contact sensors that compromise wearability. This paper presents a two-stage framework that estimates bilateral ground reaction forces and continuous gait phase using only sagittal-plane joint angles. A non-causal Diffusion-Transformer teacher reconstructs missing states via conditional diffusion inpainting with a bidirectional gated recurrent unit refinement head. A causal knowledge-distilled temporal convolutional network student then reproduces the teacher outputs in a streaming, frame-by-frame manner with 8-bit quantization-aware training. Gait events are detected by applying a force threshold to the estimated vertical ground reaction force, eliminating the need for contact sensors. Evaluated on the Camargo treadmill dataset with five unseen subjects across 35 trials at 0.5--1.85~m/s, the student model achieves a phase RMSE of 2.05%, ground reaction force RMSE of 0.421~N/kg, and gait event MAE of 9--24~ms at 275~Hz on a single CPU thread. The model is further fine-tuned and validated on a custom lower-limb exoskeleton during overground walking at normal and slow speeds, achieving phase RMSE of 3.96--4.87% and heel-strike detection MAE of 34--52~ms, demonstrating real-time feasibility on an embedded platform (median onboard latency of 3.274~ms).
|
| |
| 16:10-16:25, Paper FrBT5.3 | |
| Design and Development of a Sensorized Tendon-Driven Robotic Glove for Grasp Assistance in Stroke Rehabilitation |
|
| Sibtain, Muhammad | Southeast University, China |
| Liang, Han | Southeast University |
| Danaish, Mr | Southeast University, China |
| Al-shameri, Mohammed Ail Abdurahman | Southeast University |
| Qasem, Arfan Ali Mohammed | Southeast University |
| Baig, Mohammad Abbas | Southeast University, China |
| Shaybo, Ahmed Sharif | Southeast University |
Keywords: Robotic Applications, Rehabilitation Robot, Exoskeleton Robot
Abstract: Stroke often affects hand motor function, impacting the patient’s ability to grasp, hold and release objects, among other activities of daily living (ADLs). Robotic hand exoskeletons and assistive gloves have been identified as potential rehabilitation tools that can provide repetitive, controlled and task-oriented movement training. The design and development of a wearable robotic rehabilitation glove is presented in this paper, for grasping rehabilitation with real-time feedback. The feasibility of the proposed glove’s tendon-driven actuation mechanism for supporting finger motion and lightweight and wearable design is explored. The system includes a multimodal sensing approach with ten sensors, six of which are located on the palm side to measure forces exerted by the fingertips, and four flex sensors on the dorsal side to measure fingers bending and motion. This sensor setup allows for the continuous monitoring of hand posture, grasping performance and user interaction with objects during ADLs. The portable and sensorized assistive glove can provide a maximum assisted grasping force of 27 N to users. The designed glove aims at increasing the effectiveness of the rehabilitation by using assistive actuation and quantitative sensing for performance assessment and feedback.
|
| |
| 16:25-16:40, Paper FrBT5.4 | |
| Anticipatory Textile-Based Tactile Sensing: Unified Pre-Contact and Contact Response for Embodied Robotic Manipulation |
|
| Shim, Edward | Brighter Signals |
| Stellingwerff, Ludo | Brighter Signals |
| Klein, Andrew | Brighter Signals |
| Fraser, Christine | Brighter Signals B.V |
Keywords: Sensors and Signal Processing, Industrial Applications of Control, Exoskeleton Robot
Abstract: Humans use pre-contact cues together with tactile feedback to anticipate and adapt to physical interaction, whereas many robotic tactile systems primarily respond after contact. This paper presents a compliant textile sensing architecture integrating projected capacitive proximity sensing with distributed resistive contact sensing within a single sensing apparatus. Controlled testing demonstrated a distance-dependent capacitive response from 150 mm through direct contact, with signal separation increasing as the target approached the sensing surface. Resistive characterization demonstrated graded response under applied masses up to 12 kg, while testing the same physical specimen before and after approximately 90° folding produced a folded-to-flat mean-response ratio of 105.3%, indicating preservation of sensing functionality under geometric deformation. Controlled climate-chamber testing additionally identified systematic temperature dependence of the capacitive response. Dense temporal sampling was used to characterize within-condition signal stability and separation rather than as independent experimental replication. Qualitative robotic integration with rigid, compliant, and fragile objects further demonstrates the applicability of the architecture to embodied sensing surfaces. The results support a unified textile interface capable of extending robotic perception from pre-contact approach into physical contact while retaining mechanical conformability.
|
| |
| FrBT6 |
Symphony A, 4F |
| Robotic Applications 4 |
Oral Session |
| |
| 15:40-15:55, Paper FrBT6.1 | |
| L1 Adaptive-Augmented Computed Torque Control with Constrained APO-Based Gain Tuning for the UR10e Collaborative Manipulator |
|
| Noaman, Mohanad | School of Electrical and Electronic Engineering, University of Sheffield, Sheffield, S10 2TN, United Kingdom |
| Mahfouf, Mahdi | University of Sheffield |
| Aitken, Jonathan Maxwell | University of Sheffield |
Keywords: Control Theory and Applications, Robotic Applications
Abstract: Collaborative robots such as the UR10e are central to Industry 5.0, where manipulators must operate safely alongside humans while maintaining accurate tracking and bounded transients in the presence of uncertainty and external disturbances. Fixed-gain controllers can degrade under model mismatch, whereas conventional adaptive methods may suffer from slow adaptation, high-gain sensitivity, and transient growth. This paper presents a new integrated framework, combining Computed Torque Control (CTC), L₁ adaptive augmentation, and constrained APO-based offline gain optimisation. The established CTC layer provides nominal inverse-dynamics compensation, while the L₁ layer estimates matched uncertainties using a state predictor, projection-based adaptation, and a low-pass filter that decouples estimation from robustness shaping. The CTC gains are optimised offline using a constrained modified Artificial Protozoa Optimiser (APO), parameterised through interpretable natural-frequency (fₙ) and damping-ratio (ζ) specifications with tracking-error and actuator-torque constraints. The contribution lies in coordinating these established elements for robust adaptive control of a coupled 6-DoF manipulator. Validation uses a UR10e Simscape model following a helical trajectory under nominal and combined-disturbance scenarios. APO reduces the maximum Cartesian error from 39.2 mm to 4.52 mm. Under combined payload variation, impulse contact, and viscous friction, CTC-L₁ achieves an RMS Cartesian error of 1.72 mm, improving over CTC and CTC-IMRAC by 93.2% and 86.0%, respectively, while keeping joint torques within actuator limits.
|
| |
| 15:55-16:10, Paper FrBT6.2 | |
| A Decentralized Dual-Processor Architecture for Autonomous Hazardous Inspection Rovers Using ROS 2 and Edge AI |
|
| Muthukumar, Srihariharan | PES University |
| Kesavraj, Prasanna | PES University |
| Chandrashekhar, Chandrashekhar | PES University |
| M, Ananda | PES University |
| Awati, Mahesh | PES University |
Keywords: Autonomous Vehicle Systems, Robotic Applications, Robot Mechanism and Control
Abstract: Autonomous robotic inspection systems are increasingly being adopted in Industry 4.0 environments to reduce human exposure to hazardous conditions such as toxic gases, elevated temperatures, confined spaces, and industrial accidents. However, modern inspection robots face a fundamental systems-engineering challenge: computationally intensive perception algorithms require Linux-based computing platforms that exhibit non-deterministic scheduling behavior, while safe navigation requires deterministic real-time control. This conflict is referred to in this work as the latency paradox. This paper presents an Autonomous Multi-Modal Rover (AMR) implementing a decentralized dual-processor control hierarchy that separates perception and control into independent computational domains. A Raspberry Pi 5 running ROS 2 Jazzy Jalisco performs high-level perception, navigation, hazard detection, and multi-modal sensor fusion, while an ESP32 microcontroller operating under FreeRTOS executes deterministic motor control and safety-critical sensing. Experimental evaluation demonstrates a worst-case control latency of approximately 10 ms independent of AI workload. The vision subsystem achieves a precision of 0.932 with an average inference latency of approximately 470 ms using YOLOv8n. The proposed architecture provides a scalable and cost-effective framework for hazardous industrial inspection while preserving deterministic control performance in the presence of computationally intensive edge-AI workloads.
|
| |
| 16:10-16:25, Paper FrBT6.3 | |
| Robust Contact Torque Estimation of a Ledge-Climbing Robot Using Adaptive STA Momentum Observer with RBFNN Compensation |
|
| Pangestu, Reno | National Taiwan University of Science and Technology |
| Lin, Chi-Ying | National Taiwan University of Science and Technology |
Keywords: Robotic Applications, Sensors and Signal Processing, Control Theory and Applications
Abstract: Precise external torque estimation is essential for contact-based robot navigation, particularly when contact force is utilized to determine appropriate motion during wall-ledge climbing using admittance control. External torque estimation using momentum-observer methods does not require direct force sensing, but torque-sensor noise/disturbance and dynamic model uncertainty still affect their accuracy. In this study, the main challenge is not just high-frequency noise but also mid-frequency disturbances. A high level of filtering may minimize this disturbance but delays the impact response time, whereas low filtering preserves response speed but results in large estimation errors. To address this trade-off, we propose an adaptive momentum observer that integrates generalized momentum observation, a super-twisting algorithm (STA), a radial basis function neural network (RBFNN), impact-adaptive gain control, and unscented Kalman filtering. The RBFNNpredicts unmodeled disturbance online, while the adaptive gain can be controlled to increase rapidly after impact or release before decaying to prevent noise amplification. The simulation results indicate that the proposed method outperforms several conventional filtering methods in eliminating disturbances and smoothing the output signal, achieving a 26% reduction in estimation error compared to unfiltered SOMO, while only increasing the average contact latency by 0.052 seconds. An experiment validation using the proposed method confirms a smoother, more accurate, and faster response of estimation results compared with the previous method.
|
| |
| 16:25-16:40, Paper FrBT6.4 | |
| Laying the Foundations for Bioinspired Active Acoustic Sensing in Gas Leak Localization |
|
| Dehghan Niri, Ehsan | Arizona State University |
| Mohammad, Nasr | Arizona State University |
| Nemati, Hamidreza | Arizona State University |
| Masurkar, Nihar | Arizona State University |
| Haghshenas-Jaryani, Mahdi | New Mexico State University |
Keywords: Sensors and Signal Processing, Robotic Applications, Robot Mechanism and Control
Abstract: Gas leaks in pressurized systems are known to emit high-frequency acoustic signals, typically ranging from 10 kHz to 100 kHz, with a significant concentration of energy in the ultrasound domain (20–40 kHz). Interestingly, several biological systems—such as cats and dogs—have evolved to detect ultrasonic frequencies within this range, making their auditory morphology a compelling model for bioinspired design. This paper presents a bioinspired framework for gas leakage localization, leveraging cat head-related transfer function (HRTF) modeling to inform head-positioning strategies in a motion-oriented, binaural sensing system for active acoustic detection and localization. In the first stage, a behavioral study of domestic cats provided insights into characteristic head movements in response to acoustic cues. In the second stage, simulations were performed to investigate the underlying causes of these motions. Results suggest that increased signal-to-noise ratio (SNR) in the HRTFs can serve as a control objective for formulating head-motion control laws. Finally, we present initial steps toward a novel bioinspired HRTF-based control algorithm that integrates pressure-field directly into the control law.
|
| |
| 16:40-16:55, Paper FrBT6.5 | |
| Hybrid Attention Estimation Pipeline for Adaptive HRI Using an Expressive Robotic Head |
|
| Moraes, Pablo | Universidad Tecnologica Del Uruguay |
| RodrÍguez Espinosa, Mónica Alexandra | Universidad Tecnológica Del Uruguay |
| Peters, Christopher | Ostfalia University Wolfenbüttel |
| Jacobs Sodre Pereira, Hiago | Technological University of Uruguay |
| Doernbach, Tobias | Ostfalia University of Applied Sciences |
| Guterres, Bruna de Vargas | Universidad Technologica Del Uruguay |
| Grando, Ricardo | Federal University of Rio Grande |
Keywords: Human-Robot Interaction, Robotic Applications, Robot Vision
Abstract: This paper presents an applied case study on visual attention estimation for human-robot interaction using an expressive robotic head based on the InMoov ecosystem. The proposed pipeline combines a fast geometric perception layer with an independent semantic layer based on a vision-language model. The geometric layer provides high-frequency face and head-pose information and drives the finite state machine for real-time interaction regulation, including activation, waiting, resumption, and return to rest. In parallel, the semantic layer receives raw egocentric camera frames and produces contextual labels related to attention toward the robot, phone use, or attention elsewhere. The semantic layer is not used as a direct control signal; it is logged as an independent observer for post-hoc comparison and ambiguity analysis. The system was evaluated with 10 participants across 40 trials. Results show reliable interaction start across all trials, consistent pause behavior in the adaptive distraction condition using the geometric control layer, and non-redundant semantic labels for analyzing ambiguous interaction moments.
|
| |
| 16:55-17:10, Paper FrBT6.6 | |
| Distributed Task-Space Estimation under Time-Varying Communication Delays: Stability Analysis and Explicit Error Bounds |
|
| Yeh, Hua-Hsuan | National Cheng Kung University |
| Liu, Yen-Chen | National Cheng Kung University |
Keywords: Control Theory and Applications, Robotic Applications
Abstract: In this study, a distributed task-space estimation is investigated for heterogeneous agents communicating over a network subject to time-varying delays. Using only delayed information received from neighboring agents, each agent estimates the task-space positions of all other agents. The estimation errors are uniformly ultimately bounded, with explicit ultimate bounds characterized by the target velocities, communication-delay bounds, and network topology. When the target velocity converges to zero, both the estimation errors and observer velocities asymptotically converge to zero. Simulation results verify the predicted delay- and topology-dependent behavior.
|
| |
| FrBT7 |
Symphony B, 4F |
| ICT Convergence in Intelligent Robotics and AI Development |
Oral Session |
| Organizer: Kwon, Wookyong | ETRI |
| |
| 15:40-15:55, Paper FrBT7.1 | |
| KaDDP: Kalman-Inspired Density Diffusion Policy for Mode-Consistent Trajectory Selection (I) |
|
| Sung Jin, Kim | Keimyung University |
| Sangwon, Kim | Electronics and Telecommunications Research Institute (ETRI) |
| Kyoungoh, Lee | Electronics and Telecommunications Research Institute (ETRI) |
| Ko, Byoung Chul | Keimyung University |
| Kim, Kwang-Ju | Electronics and Telecommunications Research Institute (ETRI) |
Keywords: Robot Mechanism and Control, Artificial Intelligence Systems, Robot Vision
Abstract: Diffusion Policy has emerged as a powerful framework for robot behavior cloning, enabling multimodal action distributions that capture diverse motion strategies. However, this multimodality introduces a critical failure mode we term Decision Wavering: at each control step, the policy independently samples a trajectory, risking abrupt mode switches between consecutive steps. This causes the robot to oscillate between conflicting modes, converging toward unintended intermediate behaviors rather than committing to a single coherent strategy. We propose KaDDP, an inference-time framework that resolves Decision Wavering without retraining. Given N candidate trajectories sampled from Diffusion Policy, we estimate their densities via Kernel Density Estimation with a Gaussian kernel. A Kalman-inspired prior then incorporates the previous step's mode, computing a posterior score that favors trajectories consistent with the currently committed mode. The trajectory with the highest posterior score is selected for execution. We evaluate KaDDP on the Push-T benchmark against Diffusion Policy, demonstrating reduced steps to task completion and improved success rate over the Diffusion Policy baseline.
|
| |
| 15:55-16:10, Paper FrBT7.2 | |
| Construction and Performance Comparison of a Real-Time Action Recognition Dataset for Patrol Robots in Daytime and Nighttime Outdoor Environments (I) |
|
| Kang, Miseon | Electronics and Telecommunications Research Institute |
| Kim, Jinwoo | Kyungpook National University |
| Lee, Janghoon | Kyungpook National University |
| Lee, Sanghyeon | Kyungpook National University |
| Lee, Jong Taek | Kyungpook National University |
| Kim, Kwang-Ju | Electronics and Telecommunications Research Institute (ETRI) |
Keywords: Robot Vision, Robotic Applications, Artificial Intelligence Systems
Abstract: This paper presents an action recognition dataset and a comparative evaluation of action recognition models for patrol robots operating in daytime and nighttime outdoor environments. Unlike fixed surveillance camera environments, patrol robots operate as mobile platforms; therefore, viewpoint changes, perspective variations, camera shake, and dynamic background changes continuously occur. Robust recognition performance is also required under diverse illumination conditions. In this study, a daytime and nighttime outdoor action recognition dataset reflecting real-world patrol robot operating scenarios was constructed, and the performance of video-based and skeleton-based action recognition models was compared. The experimental results show that the skeleton-based models generally outperformed the video-based model. In particular, the model that converts skeleton information into two-dimensional heatmap representations and learns spatio-temporal features achieved the best performance, with an accuracy of 0.95, an F1 score of 0.77, and a mean average precision (mAP) of 0.84. These results demonstrate that skeleton-based approaches can be effectively applied to real-time action recognition for patrol robots in daytime and nighttime outdoor environments.
|
| |
| 16:10-16:25, Paper FrBT7.3 | |
| Rotation-Aware Continuous-Time Conflict-Based Search for Multi-Agent Path Planning (I) |
|
| Nam, Seung-Woo | Electronics and Telecommunications Research Institute |
| Kwon, Wookyong | ETRI |
| Chung, Yun Su | ETRI |
| Li, Song | Electronics and Telecommunications Research Institute |
| Park, Soon-Yong | Kyungpook National University |
| Han, Dong Seog | Kyungpook National University |
Keywords: Robotic Applications
Abstract: This paper presents a rotation-aware extension of Continuous-Time Conflict-Based Search (CCBS) for multi-agent path planning in graph-based workspaces with heterogeneous robots. Conventional CCBS and its improved variants model each transition as a single linear edge traversal and detect conflicts solely between linear motion segments. This abstraction underestimates arrival times and misses collisions during in-place turns, which is critical when robots have non-negligible angular inertia or low angular speed. The proposed method decomposes each graph-edge transition into (i) an in-place rotation phase and (ii) a linear translation phase, and integrates both phases into the SIPP low-level planner and the CBS high-level conflict resolution. The rotation time depends on the heading difference between consecutive edges and on the agent-specific angular velocity w_(k ); the linear time depends on the edge weight and the agent-specific linear velocity v_k. Conflict detection is extended to cover move–move, rotate–move, move–rotate, and rotate–rotate pairs, ensuring safety during turning. Because v_k and w_(k )are per-agent parameters, the framework naturally accommodates fleets of robots with heterogeneous kinematic capabilities. Experiments on graph-based roadmaps with ten agents show that the rotation-aware formulation reduces the number of high-level CT expansions by 47.2% and low-level searches by 54.7% compared with the original CCBS, at the cost of a 3.7% increase in flowtime that reflects the physical realities of robot turning.
|
| |
| 16:25-16:40, Paper FrBT7.4 | |
| Reinforcement Learning-Based Precision Control for Robot Manipulators in Repetitive Tasks (I) |
|
| Kwon, Wookyong | ETRI |
| Song, Bongsub | Electronics and Telecommunications Research Institute |
| Jin, Yongsik | Daegu Gyeongbuk Institute of Science and Technology |
Keywords: Control Theory and Applications, Robotic Applications, Artificial Intelligence Systems
Abstract: Repetitive industrial manipulation tasks such as pick-and-place, polishing, and assembly demand high precision under stochastic disturbances. Classical iterative learning control (ILC) exploits trial-to-trial repetition but is brittle under non-repetitive disturbances, while reinforcement-learning (RL) residual policies often ignore the trial-domain error structure that is the dominant signal in repetitive operation. This paper proposes Iteration-Domain Residual ILC (IDR-ILC), a Soft Actor-Critic (SAC) policy conditioned on the previous trial's error trajectory at the same time index, generating a per-step adaptive update to a stored ILC feedforward. A saturated accumulator and an error-multiplicative action structure make the residual vanish at convergence. A emph{contraction filter} (IDR-ILC-CF) projects each trial-domain accumulator update onto the set {Deltadelta : |Deltadelta|_F le rho |e_k|_F}, giving a sufficient condition for monotonic iteration-domain contraction. On a 7-DOF Franka Panda simulation under random payload, IDR-ILC reaches final-trial joint-space RMSE of 0.0202~rad (sinusoidal) and 0.0206~rad (Lissajous), strictly dominating PD and a same-budget Residual-PPO (paired t-test p<0.05) and matching PD-ILC and RLILC in mean. IDR-ILC-CF trades a 10% mean increase for a 3times reduction in cross-seed standard deviation. At a matched 8000-step budget, memoryless Residual-PPO is worse than PD, identifying previous-trial error conditioning as the decisive signal at this data scale.
|
| |
| 16:40-16:55, Paper FrBT7.5 | |
| Toward Omnidirectional Snake Robot Locomotion Via Direction-Decomposed Reinforcement Learning (I) |
|
| Song, Bongsub | Electronics and Telecommunications Research Institute |
| Lee, Seongmun | DGIST |
| Kwon, Wookyong | ETRI |
| Yun, Dongwon | Daegu Gyeongbuk Institute of Science and Technology (DGIST) |
Keywords: Artificial Intelligence Systems, Robotic Applications, Control Devices and Instruments
Abstract: Snake robots achieve locomotion through continuous, anisotropic frictional contact distributed over the whole body, which makes a single omnidirectional reinforcement learning (RL) policy difficult to train. In this paper, we take a first step toward omnidirectional snake robot control by decomposing the control problem along the direction of travel. We first calibrate the contact parameters of a MuJoCo simulation against drop-test experiments on five real surfaces, so that both the coefficient of restitution and the peak impact force are reproduced. On this calibrated model, we randomly sample compound-serpenoid gait parameters and analyze the resulting motion in terms of forward, lateral, and yaw velocity, identifying representative parameter sets that produce effective motion in each of eight principal directions. A torque-based Motion Matrix derived from each parameter set is then used as a morphological prior that constrains the action space of a direction-specific PPO policy. Simulation results show that the learned policies reliably generate motion in their assigned directions and outperform a conventionally tuned open-loop serpenoid controller in both path straightness and average forward speed.
|
| |
| FrBT9 |
Symphony D, 4F |
| Autonomous Vehicle Systems 4 |
Oral Session |
| |
| 15:40-15:55, Paper FrBT9.1 | |
| Leveraging Planar Regularity for Gaussian Splatting SLAM |
|
| Seo, Dong-Uk | Korea Advanced Institute of Science and Technology |
| Park, Jiwon | KAIST |
| Kong, Jei | KAIST |
| Jeon, Jinwoo | KAIST |
| Myung, Hyun | KAIST (Korea Advanced Institute of Science and Technology) |
Keywords: Robot Vision, Navigation, Guidance and Control
Abstract: Feature-based visual SLAM has been incorporated into Gaussian splatting SLAM pipelines, providing camera poses, keyframes, and sparse point landmarks for Gaussian map initialization or optimization. However, the geometric information transferred from the SLAM to the Gaussian map is still largely limited to point-level structure, while other structural regularities available in man-made environments remain underexplored. In this paper, we propose a plane-regularized Gaussian splatting SLAM framework that exploits planar regularity from a feature-based visual SLAM frontend to improve the geometric quality of Gaussian maps. Our method builds the Gaussian mapping backend with 2D Gaussian splatting, representing the scene with surface-aligned Gaussian primitives that provide an interface for planar regularization. We employ a feature-based plane tracking module to recover planar structures from sparse SLAM landmarks, and incorporate the resulting planar cues into Gaussian map optimization. The resulting planar cues provide a structural bridge between the sparse SLAM map and the dense Gaussian representation by guiding initialization, supervising planar image regions, and regularizing Gaussian primitive positions and normals. Experimental results demonstrate that the proposed method improves reconstruction accuracy while preserving photometric rendering quality and maintaining real-time performance.
|
| |
| 15:55-16:10, Paper FrBT9.2 | |
| Study on Alleviating Data Acquisition Constraints in In-Cabin Monitoring Using Synthetic-To-Real Generalization |
|
| Kim, Suk Hyun | Mobase Electronics |
| Kim, In Ji | Mobase Electronics |
| Park, Sam Min | Mobase Electronics |
| Choi, Jun Sam | Mobase Electronics |
Keywords: Artificial Intelligence Systems, Autonomous Vehicle Systems
Abstract: The In-Cabin Monitoring System (ICMS) is a key vehicular safety system designed to detect the states and behaviors of drivers and passengers. However, real-world image-based data collection is time-consuming, costly, and inherently limited in capturing sufficiently diverse and complex scenarios. To address this issue, this paper proposes a data training framework based on a synthetic-to-real generalization strategy. The proposed framework generates synthetic images through a diffusion model by jointly incorporating semantic information extracted from a Visual Semantic Encoder, appearance information encoded by a Variational Autoencoder (VAE), and conditioning information derived from text prompts. Experimental results show that the model trained with a mixed dataset of real and synthetic images achieved detection performance close to that of real-only training on real vehicle images. These findings suggest that the proposed framework can serve as a practical complementary approach for mitigating the limitations of real-world data collection in vehicle environments.
|
| |
| 16:10-16:25, Paper FrBT9.3 | |
| Adaptive RBF Control, and Constrained Optimal Allocation for Autonomous USVs: Real-Sea Validation |
|
| Yang, Cheng-Teng | National Cheng Kung University |
| Chen, Yung Yu | National Cheng Kung University |
| Yu, Wen-Sheng | National Cheng Kung University |
| Lan, Chun-Han | National Cheng Kung University |
| Liu, Cheng-Wei | National Cheng Kung University |
Keywords: Autonomous Vehicle Systems, Navigation, Guidance and Control, Control Theory and Applications
Abstract: This paper presents a fully integrated guidance and control framework for autonomous unmanned surface vessels (USVs) that achieves high-precision trajectory tracking in real maritime environments. The framework integrates two tightly coupled modules: (1) a Lyapunov-based adaptive RBF nonlinear controller with online parameter estimation, whose stability is rigorously established via Lyapunov stability analysis together with Barbalat's Lemma; and (2) a Lagrange multiplier-based constrained optimal control allocation scheme for rotatable waterjet propulsors, augmented with dedicated RBF networks for accurate actuator modeling. The proposed architecture has been successfully implemented and validated on the Lab611 USV. Extensive real-sea experiments demonstrate exceptional performance, achieving sub-meter tracking accuracy with average tracking errors of 0.26 m over an approximately 2 km ring-port trajectory, 0.14 m on straight-line-plus-curve missions in constrained narrow channels, and 0.42 m on serpentine cruising paths, all while strictly respecting actuator saturation constraints. These results confirm the practical feasibility, robustness, and superior tracking accuracy of the integrated approach in realistic sea states (0-3).
|
| |
| 16:25-16:40, Paper FrBT9.4 | |
| Temporal Model Smoothing for Sample-Limited Local Learning MPC |
|
| Pranayanuntana, Pawornrat | Department of Control Systems and Instrumentation Engineering, KMUTT |
| Panomruttanarug, Benjamas | King Mongkut's University of Technology Thonburi |
| Boonto, Sudchai | KMUTT |
Keywords: Autonomous Vehicle Systems, Navigation, Guidance and Control, Control Theory and Applications
Abstract: This paper studies temporal inconsistency in local regression-based Learning Model Predictive Control (Local LMPC) under sample-limited model identification. Using a large local regression neighborhood can reduce model variation, but may increase data processing and make the learned model less local. Conversely, using fewer samples makes the identified affine model more sensitive to neighbor-set switching and regression conditioning. To diagnose this effect, we use the model-variation metric Δmodel and the normalized neighbor-switching score sneigh. We then apply a fixed smoothing-parameter exponential update to the locally identified model matrices before they enter the MPC quadratic program, without changing the horizon, constraints, local regression rule, or QP dimensions. Closed-loop nonlinear kinematic-bicycle simulations on an L-shaped track show a consistency–responsiveness trade-off: large α gives faster but more chattering behavior, whereas small α reduces model and steering variation at the cost of slower response.
|
| |
| FrPo6P |
Symphony D, 4F |
| Poster Session 6 |
Poster Session |
| |
| 09:30-10:30, Paper FrPo6P.1 | |
| A Feasibility Study on In-Fingertip Null-Space Control for Vial Cap Engagement and Screwing |
|
| Heo, Jeong | Hanyang University |
| Seo, TaeWon | Hanyang University |
| Hwang, Donghyun | Korea Institute of Science and Technology |
Keywords: Robot Mechanism and Control
Abstract: Robotic automation in laboratory environments can involve object handling in confined workspaces, where surrounding structures restrict large wrist rotations and arm motions. This paper presents the feasibility of the alignment and insertion preceding vial cap engagement and screwing by formulating the problem as an object-handling task and solving it with a task-priority based control framework that performs manipulation through finger-level motion of a dexterous hand, without relying on arm-centric movement. Grasp geometry that is represented by the three pairwise fingertip distances is maintained as the primary task while object translation and rotation are generated within its null-space as the secondary task. The object proxy pose is estimated from fingertip positions via forward kinematics, without any external pose sensor. The framework was validated on hardware with the arm wrist held fixed, so that all object motion was produced by finger-level manipulation. In single-axis experiments, we measured a sub-millimeter primary grasp deviation of 0.59 mm on average, and in task-level demonstration the robotic hand aligned and inserted the object while preserving the grasp geometry. These results support the feasibility of finger-level in-hand manipulation for object handling in workspace-constrained laboratory settings.
|
| |
| 09:30-10:30, Paper FrPo6P.2 | |
| Optimal Path Search for Robot-Arm-Based Magnetic Control of a Wireless Soil Sampler |
|
| Kang, SeonKyu | Chungbuk National University |
| Kim, Jayoung | Korea Institute of Medical Microrobotics |
Keywords: Robot Mechanism and Control, Robotic Applications, Control Devices and Instruments
Abstract: This study proposes a method for determining optimal positions and orientations of a robotic end-effector used in curved-path drilling of a wireless soil sampler. The system consists of an actuation magnet attached to the robotic end-effector and a permanent magnet embedded in the soil sampler. By controlling the end-effector pose, magnetic force and torque are generated to wirelessly propel and steer the sampler during drilling. Unlike conventional Edelman-auger-type soil samplers that are limited to straight-line penetration, the proposed method enables curved drilling trajectories for obstacle avoidance in underground environments. The drilling path is discretized into a sequence of steps along a refer-ence trajectory connecting the start and goal points. At each step, candidate end-effector poses are evaluated based on the magnetic force and torque required to propel and steer the sampler, which are estimated using a magnetic dipole model and an agarose (0.6 wt%) gel resistance model. Robot reachability is verified through inverse kinematics. A two-stage optimization identifies feasible poses satisfying both magnetic actuation and kinematic constraints, then selects the pose that best follows the reference path while minimizing joint motion. The proposed method enables curved-path drilling planning under coupled magnetic and robotic constraints for obstacle avoidance soil sampling.
|
| |
| 09:30-10:30, Paper FrPo6P.3 | |
| Preparation of a Digital Twin Model of Harmonic Drive Gear System Using Nonlinear Finite Element Analysis for Physical AI Applications |
|
| Park, Moon-Woo | Korea Institute for Robot Industry Advancement |
Keywords: Robot Mechanism and Control, Robotic Applications
Abstract: This paper presents a digital twin modeling approach for a harmonic drive reduction gear system based on nonlinear finite element analysis (FEA) for future Physical AI applications. Unlike conventional rigid-body transmission models, the proposed model explicitly represents the elastic deformation of the flexspline induced by the wave generator and the nonlinear tooth engagement between the flexspline and circular spline. A commercial harmonic drive reducer model was selected as the target system. A detailed finite element model was developed using HyperWorks/OptiStruct, including wave generator, flexspline, circular spline, and bearing assemblies. Nonlinear surface-to-surface contact formulations were employed to simulate deformation propagation and torque transmission mechanisms. Static structural analysis was performed under the average torque limit. The simulation results demonstrated realistic deformation patterns, tooth contact distributions, and stress concentrations. Maximum von-Mises stress of the flexspline was predicted as 549 MPa, remaining below the material yield strength of SCM440 steel. The developed model provides a physics-based foundation for digital twin construction and future reduced-order modeling techniques applicable to Physical AI systems. Future research will focus on integrating simulation data with machine learning frameworks to enable real-time state estimation, predictive maintenance, and intelligent motion planning of robotic systems utilizing harmonic drive reducers
|
| |
| 09:30-10:30, Paper FrPo6P.4 | |
| Simulation-Based Diffusion Policy Training Framework for Precision Task |
|
| Park, Sangyong | Kyungpook National University |
| Ko, Yeongmin | Kyungpook National University |
Keywords: Robot Mechanism and Control
Abstract: Precision manipulation is essential in modern robotic applications such as assembly, cutting, and inspection. Especially in cutting, simple position error can lead to project failure or damage of valuable component. Consequently, developing reliable manipulation policies for precision is required and remains a challenge. and this kind of task requires huge amount of data generated by physical robot manipulation. But this kind of data consumes time and resources. So this paper proposes a way to train precision manipulation task via simulation. The simulation can generate 100 trajectories in 511seconds and generated datasets contains three categories Video, joints, and trajectory. Each category contains as much related information as possible so it can be used for multiple model. In this paper trained model is diffuser. Result of the test shows diffuser can be trained with data generated by simulation and able to create trajectory. but It's only been tested in simulation environment, further research is required to test how it works with actual hardware. Contribution of paper is Propose a simulation for high-precision manipulation task training, Using open sourced manipulator(SO101) shows high-precision manipulation task can be performed with simple hardware and can be used in various systems, proves model can be trained with gathered data and able to generate trajectory.
|
| |
| 09:30-10:30, Paper FrPo6P.5 | |
| Real-Time Control of Tendon-Driven Continuum Robot Using MPPI and Residual Jacobian-Based Online Learning |
|
| Lee, Dongjun | Daegu Gyeongbuk Institute of Science and Technology(DGIST) |
| Kim, DongWook | Daegu Gyeongbuk Institute of Science and Technology (DGIST) |
Keywords: Robot Mechanism and Control, Control Theory and Applications, Artificial Intelligence Systems
Abstract: Tendon-driven continuum robots (TDCRs) are widely used in confined operating systems due to their thin shape, flexibility. Several modeling and control methods have been used for TDCRs, such as Cosserat rod model and model predictive control (MPC). However, Cosserat rod models and MPC have limitations in real-time control and handling the nonlinear behavior of TDCRs due to computational complexity and dynamics linearization, respectively. In this paper, to address the two problems, we employ a PCC model based residual radial basis function network (RBFN). Although the PCC model is computationally efficient due to its simple approximation, it cannot fully explain modeling errors; therefore, we augment with a residual RBFN to compensate for the errors that the PCC model cannot account for. Moreover, we employ a model predictive path integral (MPPI) controller that uses a Jacobian-based PCC–residual RBFN model as the rollout model, and the residual RBFN model is updated online using buffer data. We validate the proposed method on a handmade one-segment TDCR hardware consisting of four tendons, with vision-based tracking using two webcams. The reference trajectory tracking test shows that the proposed residual RBFN method reduces the mean tracking error, RMS tracking error, and model prediction error by 47.0%, 41.6%, and approximately 50%, respectively.
|
| |
| 09:30-10:30, Paper FrPo6P.6 | |
| RCM-Consistent Admittance Control with Inverse-Dynamics QP for Hands-On Robotic Manipulation |
|
| Jeong, Jaehun | Korea Advanced Institute of Science and Technology (KAIST) |
| Park, Seongsu | KAIST |
| Lee, Sanghoon | KAIST |
| Kim, Min Jun | KAIST |
Keywords: Robot Mechanism and Control, Robotic Applications, Human-Robot Interaction
Abstract: Robot-assisted minimally invasive surgery (RAMIS) often requires a surgical tool to pivot about a trocar while allowing direct hands-on manipulation by the surgeon. During kinesthetic guidance, however, operator-applied forces and moments act as external inputs to the robot. Their interaction with nonlinear robot dynamics and physical limits can result in remote center of motion (RCM) constraint violation or infeasible torque commands. This paper proposes an RCM-constrained control framework for hands-on manipulation that combines an RCM-consistent admittance filter with inverse-dynamics quadratic programming (QP). The admittance filter generates a joint-space reference trajectory from the measured external wrench while enforcing the RCM constraint at the position, velocity, and acceleration levels. The inverse-dynamics QP then computes torque commands that track the reference trajectory while considering the joint position, velocity, and torque limits. Baumgarte stabilization is incorporated into the acceleration-level RCM formulation to reduce errors caused by numerical drift and saturation. Simulation studies using a 6-DoF manipulator compare the proposed method with two baseline controllers under insertion-roll interaction scenarios. Overall, the results show that the proposed controller reduces both the RCM error and its time derivative while accounting for the prescribed joint torque limits in the inverse-dynamics QP.
|
| |
| 09:30-10:30, Paper FrPo6P.7 | |
| Dual-Attachment Cable Suspension System for Zero-Gravity Emulation of Space Manipulators |
|
| Shin, Seungmin | Korea Advanced Institute of Science and Technology |
| Kim, Jiwon | KAIST |
| Kim, Min Jun | KAIST |
Keywords: Robot Mechanism and Control, Robotic Applications, Control Devices and Instruments
Abstract: Space manipulators are increasingly used in various on-orbit operations. Accordingly, reliable on-ground validation of such systems has become increasingly important. However, because space manipulators are designed for a microgravity environment, their actuators are not capable of supporting gravity loads, which makes on-ground testing challenging. This paper proposes a dual-attachment cable suspension system for zero-gravity emulation of a 7-DOF space manipulator. For each attachment point, the required external force for gravity compensation is computed and mapped to feasible cable tensions under pull-only and tension-limit constraints. By applying cable-generated external forces at two different links, the proposed system achieves gravity compensation for the six dominant gravitational joint torques. The approach is evaluated in simulation with the CAESAR manipulator through joint-space trajectory tracking and a pick-and-place task. The results show that the proposed system successfully compensates for gravitational torques and produces motion close to zero-gravity behavior, while maintaining feasible cable tensions throughout the task.
|
| |
| 09:30-10:30, Paper FrPo6P.8 | |
| Learning to Utilize Passive Toe Dynamics in Bipedal Walking Via Adversarial Motion Priors |
|
| Kim, Minseok | Gwangju Institute of Science and Technology |
| Hur, Pilwon | Gwangju Institute of Science and Technology |
Keywords: Robot Mechanism and Control, Artificial Intelligence Systems, Robotic Applications
Abstract: Passive toe joints can provide compliance during stance, but their unactuated motion is difficult to induce because it depends on robot dynamics and ground contact. This paper presents a motion imitation approach based on Adversarial Motion Priors (AMP) for inducing passive toe-joint motion in bipedal walking without toe torque commands or prescribed toe trajectories. Human walking motion is retargeted to a toe-jointed bipedal robot and used for policy training. The passive toe joints are excluded from the policy action space and are not assigned reference trajectories. Instead, their motion emerges through passive spring-damper dynamics, robot motion, and ground contact. To prevent passive toe motion from being penalized as an imitation mismatch, passive toe states are masked in the discriminator observation. We compare the proposed condition with a fixed-toe baseline and a no-masking ablation. Simulation results across three independently trained random seeds show clear stance-phase passive toe dorsiflexion and toe loading. The no-masking ablation produces smaller toe dorsiflexion, suggesting that discriminator masking promotes greater passive-toe utilization. These results suggest that passive toe utilization can emerge from AMP-based motion imitation without explicit toe commands or prescribed toe trajectories.
|
| |
| 09:30-10:30, Paper FrPo6P.9 | |
| Hall-Only Wheel Friction-Change Detection: A Feasibility Study |
|
| Lee, Chang-Hwan | Gyeongkuk National Univ |
Keywords: Robot Mechanism and Control, Autonomous Vehicle Systems, Robotic Applications
Abstract: Detecting changes in wheel–terrain interaction, such as traction loss, slip, or contact loss, is essential for the safe autonomous operation of mobile robots. Conventional approaches typically rely on additional sensors such as inertial measurement units, force/torque sensors, or high-resolution encoders, which increase cost and integration complexity. This paper presents a simulation-based feasibility study of detecting wheel friction changes using only the three-phase Hall sensors already embedded in commercial in-wheel brushless DC (BLDC) motors, without any additional sensing hardware. The proposed pipeline combines a phase-locked loop (PLL) speed estimator, which reconstructs wheel angular velocity from quantized Hall edge timings, with a generalized momentum observer (GMO) that estimates the disturbance torque acting on the wheel axis. Because the observer residual converges to the negative of the disturbance torque, a change in rolling friction appears directly as a measurable shift in the residual, enabling friction-change detection without explicit acceleration measurement. The pipeline is evaluated under open-loop constant-torque operation across five operating points, from 20% to 100% of rated torque. The residual converges to the true disturbance torque with a steady-state error below 0.4%, independently of the operating point. Three representative disturbance events—an external impulse, a partial traction loss (slip), and a complete contact loss (drop-off)—are shown to be distinguishable in the time-domain residual through their amplitude and recovery behavior. A comparison against a 1000-pulse-per-revolution encoder shows that the Hall-plus-PLL configuration matches or exceeds the encoder over 80% of the operating envelope, while requiring no additional sensors. These results indicate that Hall-only sensing is sufficient for reliable wheel friction-change detection over a wide operating range, providing a practical, deployable basis for proprioceptive terrain monitoring. Ongoing work extends the approach to four-wheel platforms with active step-torque probing for quantitative friction estimation.
|
| |
| 09:30-10:30, Paper FrPo6P.10 | |
| Force-Manipulability-Aware Whole-Body Optimization for Quadruped Manipulators |
|
| Choi, Kanghyeon | University of Seoul |
| Hwang, Myun Joong | University of Seoul |
Keywords: Robot Mechanism and Control, Robotic Applications
Abstract: Quadruped manipulators require whole-body postures that ensure both end-effector reachability and effective force transmission during contact tasks. This paper presents a posture optimization method based on directional force manipulability for pulling tasks. Given a target end-effector pose and pulling direction, the method searches over the object-relative body distance, body height, and body pitch. For each candidate, arm and leg inverse kinematics are solved, and candidates are filtered using end-effector pose accuracy and support feasibility constraints. The feasible posture with the maximum directional force manipulability is selected as the whole-body configuration. The method is evaluated in Isaac Sim using a Go1 quadruped equipped with a PiPER manipulator and randomly sampled target poses. The results show that feasible postures can be generated across diverse targets and that body placement substantially affects directional force manipulability even when the same target end-effector pose is satisfied.
|
| |
| 09:30-10:30, Paper FrPo6P.11 | |
| Predictive versus Reconstructive Terrain Representations for Quadrupedal Locomotion under Depth Degradation |
|
| Kim, Taehyeong | Kyungpook National University |
| Lee, Sangmoon | Kyungpook National University |
Keywords: Robot Mechanism and Control, Artificial Intelligence Systems, Robot Vision
Abstract: Terrain-aware legged locomotion increasingly depends on a single onboard depth camera, yet real depth degrades through holes, occlusion, and dropout. Reconstruction-based encoders learn a latent by reproducing a height map, forcing it to account for every observed value and tying its quality to the sensor. We ask whether predictive representations—which predict the future latent rather than reconstructing the terrain—degrade more gracefully. We present a locomotion backbone with a swappable representation head for a controlled comparison: a fixed PointNet encoder, a proprioceptive encoder that also estimates base velocity, a recurrent memory, and an implicit terrain-aware RL policy are held identical, and only the head is switched between reconstruction and a JEPA-style latent prediction. On the Unitree Go2 in Isaac Lab, both heads walk under clean depth, but as it is corrupted the reconstructive representation degrades sharply while the predictive one degrades gracefully and keeps walking—consistently across dropout, holes, and occlusion.
|
| |
| 09:30-10:30, Paper FrPo6P.12 | |
| Analysis of the Propulsion Performance of a Coaxial Counter-Rotating UAV with Propeller Blade Numbers |
|
| Kang, ChanHwi | Poongsan |
| Yoo, SungJoo | Poongsan |
| Song, Yi-Hwa | Poongsan |
| Jang, JaeHun | Pukyong National University |
Keywords: Robot Mechanism and Control, Robotic Applications, Navigation, Guidance and Control
Abstract: Coaxial counter-rotating propulsion systems provide high thrust density within a compact airframe, making them suitable for uncrewed aerial vehicles (UAVs). However, the aerodynamic interaction between the upper and lower rotors causes propulsion characteristics that differ from those observed in single-motor tests. This study investigates the effect of propeller blade number on the propulsion performance of a coaxial counter-rotating UAV. A preliminary experiment was conducted using a single motor to compare the performance of two-blade and four-blade propellers under identical operating conditions. The results showed that the two-blade propeller generated higher thrust than the four-blade propeller. To evaluate the actual flight performance, both propeller configurations were applied to a coaxial counterrotating UAV, and flight tests were performed under identical flight conditions. Flight efficiency was evaluated based on flight tests conducted from takeoff until the battery voltage decreased to 17V. Although the two-blade propeller exhibited superior thrust performance in the single-motor experiment, the four-blade propeller achieved a longer flight time during the flight test. These results indicate that single-motor performance alone is insufficient for evaluating the propulsion efficiency of coaxial counter-rotating UAVs, and that aerodynamic interactions between coaxial rotors should be considered when selecting propeller configurations.
|
| |
| 09:30-10:30, Paper FrPo6P.13 | |
| A Fuzzy Null-Space Controller for Human-Like Elbow Configuration in Dual-Arm 7-DOF Humanoid Manipulators Based on Palm Orientation |
|
| Lee, Hunjo | Korea University of Science and Technology, Korea Institute of Industrial Technology |
| Yang, Gi-Hun | KITECH |
Keywords: Robot Mechanism and Control, Robotic Applications, Control Theory and Applications
Abstract: For a 7-DOF anthropomorphic manipulator, the redundant degree of freedom left by the end-effector task must be resolved to obtain a human-like posture. Prior approaches rely on either a structure-specific analytic swivelangle model or a posture prior learned from human motion-capture data. This paper proposes a lightweight null-space controller that infers a preferred elbow direction directly from the real-time palm orientation using a rule-based fuzzy model, combined with manipulability maximization. Since all variables are defined purely from forward kinematics and the Jacobian, the method needs no structure-specific derivation and no training data, and is readily portable across 7-DOF arm designs. On a dual-arm UFactory xArm7 platform in MuJoCo, the elbow direction converges to the fuzzy-inferred target as the palm orientation is swept, while the end-effector position error remains negligible.
|
| |
| 09:30-10:30, Paper FrPo6P.14 | |
| Revisiting Scene Change Detection with Pixel-Wise Temporal Fusion |
|
| Kim, Dong Yeop | KETI (Korea Electronics Technology Institute) |
| Kim, Keunhwan | Korea Electronics Technology Institute |
| Hwang, Jung-Hoon | Korea Eletronics Technology Institute |
| Kim, Euntai | Yonsei University |
Keywords: Robot Vision, Artificial Intelligence Systems, Robotic Applications
Abstract: Scene change detection aims to identify semantic changes between observations acquired at different times. Existing deep-learning-based approaches mainly process image pairs independently, although robotic perception systems naturally observe environments as temporally sequential image streams. Motivated by Bayesian fusion and the Markov assumption in probabilistic robotics, we revisit scene change detection from a temporal fusion perspective. We propose a pixel-wise temporal fusion framework that integrates current observations with temporally accumulated historical evidence. A frozen optical-flow model aligns the current-frame prediction with a fusion tensor propagated from previous frames, enabling temporally consistent pixel-wise correspondence. The proposed architecture provides two complementary pathways: one captures instantaneous semantic differences from the current image pair, while the other progressively aggregates temporal evidence across sequential observations. Experiments on the ChangeSim dataset demonstrate that the proposed approach effectively improves temporal consistency and achieves 29.6 mIoU in multi-class scene change detection.
|
| |
| 09:30-10:30, Paper FrPo6P.15 | |
| DUET-SLAM: Dual-Flow Consistency for Dense Monocular SLAM in Dynamic Environments with Geometry-Aware Foundation Model |
|
| Jeon, Jinwoo | KAIST |
| Seo, Dong-Uk | Korea Advanced Institute of Science and Technology |
| Myung, Hyun | KAIST (Korea Advanced Institute of Science and Technology) |
Keywords: Robot Vision, Robotic Applications, Artificial Intelligence Systems
Abstract: Geometry-aware foundation models have recently enabled monocular visual simultaneous localization and mapping (SLAM) to perform camera tracking and dense reconstruction without known camera intrinsics. In SLAM systems that leverage geometry-aware foundation models, estimating dense pointmap correspondences is essential for pose optimization. However, such correspondences are typically estimated under a static-scene assumption, and moving objects can corrupt dense correspondences and degrade pose estimation in dynamic environments. To tackle this problem, we propose DUET-SLAM, a training-free dense monocular SLAM framework that improves the robustness of SLAM systems based on geometry-aware foundation models in dynamic environments. Our key insight is that dense correspondences can be interpreted as an implicit optical flow. We introduce a training-free flow-consistency mechanism that compares the implicit flow with optical flow estimated using an off-the-shelf network. The resulting flow discrepancy is utilized to identify motion-inconsistent correspondences associated with dynamic objects. Experiments on the dynamic sequences of the TUM RGB-D dataset demonstrate that DUET-SLAM improves camera pose estimation accuracy compared with the baseline.
|
| |
| 09:30-10:30, Paper FrPo6P.16 | |
| Depth-Supervised Event Gaussian Splatting for Accurate Surface Reconstruction |
|
| Kim, Yunsoo | KAIST |
| Lee, Taeji | Korea Advanced Institute of Science and Technology |
| Myung, Hyun | KAIST (Korea Advanced Institute of Science and Technology) |
Keywords: Robot Vision, Sensors and Signal Processing, Robotic Applications
Abstract: Event cameras have emerged as a compelling sensing modality for novel view synthesis, offering high temporal resolution, low latency, and high dynamic range. However, existing event-based 3D Gaussian splatting methods suffer from inaccurate surface geometry, as flat surface regions with uniform appearance generate sparse or no events, leaving interior surface areas severely under-supervised. In this paper, we propose depth-supervised event Gaussian splatting (DSEGS), a framework that integrates two complementary geometric supervision signals into event-based 3D Gaussian splatting. First, for each event frame, we estimate a dense depth map using a pretrained monocular depth estimation model and align it with the rendered depth via scale-shift fitting. Second, we enforce multi-frame normal consistency by computing world-space surface normals from rendered depth maps and penalizing angular differences at depth-warped corresponding points across frames, leveraging the constraint that the same physical surface must exhibit consistent normals from any viewpoint. Experiments on the Replica dataset demonstrate that DSEGS significantly improves surface accuracy over existing event-based methods.
|
| |
| 09:30-10:30, Paper FrPo6P.17 | |
| Reliability-Weighted Active Gaussian Reconstruction under Camera Pose Uncertainty |
|
| Oh, Sangcheol | Kwangwoon University |
| Oh, Junghyun | Kwangwoon University |
Keywords: Robot Vision, Navigation, Guidance and Control, Artificial Intelligence Systems
Abstract: Active Gaussian reconstruction aims to efficiently reconstruct a scene by selecting informative viewpoints under a limited observation budget. Recent Gaussian Splatting-based active reconstruction methods guide viewpoint se- lection using the current map state, but they commonly assume reliable camera poses during map updating. In practical robotic scenarios, camera poses can contain translational and rotational errors, and observations acquired with inaccurate poses may be incorrectly accumulated into the Gaussian map. Since the updated map state is reused for subsequent view- point selection, such unreliable accumulation can distort the confidence map and degrade reconstruction performance. In this paper, we propose a reliability-weighted map update scheme for active Gaussian reconstruction under camera pose uncertainty. The proposed method keeps the viewpoint selection process of ActiveGS unchanged, while modifying how newly acquired RGB-D observations are accumulated into the Gaussian map. The input observation is first locally aligned with the current Gaussian map, and a frame-level pose reliability score is estimated from the residual depth discrepancy and geometric sensitivity. This reliability score, together with distance-dependent attenuation, adjusts the observation con- tribution of each Gaussian primitive during confidence update. Experiments on Replica indoor scenes under controlled pose-noise settings show that the proposed method consistently improves both peak signal-to-noise ratio (PSNR) and completion ratio compared with ActiveGS.
|
| |
| 09:30-10:30, Paper FrPo6P.18 | |
| 360DVIO: Deep Visual-Inertial Odometry Using a 360-Degree Camera |
|
| Lee, Seunghun | KAIST (Korea Advanced Institute of Science and Technology) |
| Nam, Jihun | KAIST (Korea Advanced Institute of Science and Technology) |
| Myung, Hyun | KAIST (Korea Advanced Institute of Science and Technology) |
Keywords: Robot Vision, Navigation, Guidance and Control, Sensors and Signal Processing
Abstract: 360-degree cameras keep the whole surroundings in view even under aggressive 6-DoF motion and are therefore well suited for state estimation on unmanned aerial vehicles (UAVs) and mobile robots. In such dynamic scenarios, metric scale recovery is critical, yet the scale ambiguity inherent to monocular vision remains, and the equirectangular projection (ERP) distorts scene geometry unevenly across latitude, further complicating depth estimation from visual cues. At the same time, the full-sphere field of view is geometrically complementary to inertial measurements and supports stable gravity estimation and scale initialization in visual-inertial odometry (VIO). Existing 360-degree VIO methods rely on handcrafted features that degrade under illumination changes and low-texture conditions, while learning-based 360-degree odometry remains visual-only and unable to recover metric scale. We propose a learning-based VIO framework for a single 360-degree camera. It integrates a pre-trained recurrent network for equirectangular patch correspondence into a tightly coupled factor graph, with the reprojection residual adapted to ERP spherical geometry. Experiments on public benchmark datasets show that the proposed framework recovers metric-scale trajectories and achieves the lowest trajectory error on four of the five evaluated sequences, including low-light and aggressive-motion settings.
|
| |
| 09:30-10:30, Paper FrPo6P.19 | |
| Geometry-Preserving Sim-To-Real Data Generation for Centroid-Stable Visual Servoing in EV Battery Disassembly |
|
| Park, Seong Eun | Hanyang University, THOTHInc |
| Lee, Sang Hyoung | Korea Institute of Industrial Technology |
| Cho, Nam Jun | Hanyang University |
| Kwon, Taesoo | Carnegie Mellon University |
Keywords: Robot Vision, Process Control Systems, Industrial Applications of Control
Abstract: End-of-life EV battery disassembly requires precise vision-guided bolt alignment under hazardous and non-standardized conditions. In Image-Based Visual Servoing (IBVS), the predicted bolt centroid directly drives robot motion; therefore, even a small centroid bias can exceed socket tolerances and cause alignment failure. However, accurately annotated centroid data are difficult to obtain because labeling requires expert knowledge and direct access to each battery pack. To address this bottleneck, we propose a geometry-preserving sim-to-real data generation method that separates bolt geometry from appearance variation. Pixel-accurate rendered geometry and ground-truth centroids are fixed first, and diffusion-based appearance variation is introduced only through constrained composition. A geometric consistency filter, ShapeCons, then removes samples with residual boundary or center-mark distortion that could shift centroid predictions. In real-robot experiments, the proposed method achieves 89.5% closed-loop IBVS alignment success at the ≤2 px criterion, yielding a 42.1 percentage-point improvement over an appearance-augmentation baseline without geometry filtering. These results show that geometric consistency in synthetic data is critical for reliable visual servoing, even when conventional detection accuracy remains high.
|
| |
| 09:30-10:30, Paper FrPo6P.20 | |
| UToM: Uncertainty-Aware Token Merging for Efficient Semantic Segmentatio |
|
| Kim, Wanhee | Korea Advanced Institute of Science and Technology |
| Sung, Chang Ki | KAIST |
| Myung, Hyun | KAIST (Korea Advanced Institute of Science and Technology) |
Keywords: Robot Vision, Artificial Intelligence Systems, Autonomous Vehicle Systems
Abstract: Vision Transformer-based semantic segmentation achieves strong performance, but its quadratic self-attention cost makes high-resolution dense prediction computationally expensive for resource-constrained systems. Token merging reduces this cost by combining redundant tokens, but existing merging criteria often rely primarily on feature similarity, which may not fully capture semantic reliability in dense prediction. To address this limitation, this paper proposes UToM, an uncertainty-aware token merging method for efficient semantic segmentation. UToM predicts token-level semantic probabilities from intermediate features using a lightweight auxiliary head and derives pair-level semantic confidence for candidate token pairs. This confidence is combined with feature similarity to compute an uncertainty-aware merge score. By suppressing semantically unreliable merges and encouraging merges between tokens that confidently support the same semantic class, UToM improves the accuracy-efficiency trade-off in semantic segmentation. Experimental results show that semantic confidence provides an effective guidance signal for reliable token merging and enables improved accuracy-efficiency trade-offs in dense prediction.
|
| |
| 09:30-10:30, Paper FrPo6P.21 | |
| VLM-Driven Characterization of Action Points in Unknown Environments for Autonomous Robots |
|
| Häuselmann, Ramona | Luleå University of Technology |
| Valdes Saucedo, Mario Alberto | Lulea University of Technology |
| Kanellakis, Christoforos | LTU |
| Nikolakopoulos, George | Luleå University of Technology |
Keywords: Robot Vision, Robotic Applications, Artificial Intelligence Systems
Abstract: Field robot deployments increasingly rely on scene representations that characterize actionable environment elements and their relationship to robot capabilities and tasks. This paper presents a hierarchical perception-reasoning architecture based on a VLM-driven scene representation that outputs the state and location of observed action points in the visited environment. A vision-language-based reasoning framework fuses RGB, depth, and voxelized occupancy representation into a spatial world model, extracting high-level action points with semantic context annotations. These VLM-proposed scene highlights are stored in a candidate database and actively refined using a best-view selection module that selects viewpoints to improve observation quality. A task-level reasoning module then operates over the action point representation to annotate them with additional information such as semantic properties relevant to robot missions. Real-world experiments in indoor environments demonstrate robust operation, improved viewpoint selection, and scene understanding for autonomous field robots.
|
| |
| 09:30-10:30, Paper FrPo6P.22 | |
| Chunk-Based Diffusion VLA with State-Free Inference for Dual-Arm Humanoid Robot Simulation |
|
| Sanaullah, Sanaullah | Korea Institute of Machinery and Materials |
| Kumar, Abhishek | Korea Institute of Machinery and Material |
| Abbasi, Saad Jamshed | Pusan National University |
| Kim, Jeong Yong | Korea Institute of Machinery and Materials |
| Han, Byung-Kil | Korea Institute of Machinery and Materials |
| Park, Dongil | Korea Institute of Machinery and Materials (KIMM) |
Keywords: Robot Vision, Artificial Intelligence Systems, Robotic Applications
Abstract: We evaluate StarVLA — a Qwen2.5-VL-3B vision-language backbone with a flow-matching diffusion action head — on the Robotis SG2 AI Worker in NVIDIA IsaacLab, comparing two separately fine-tuned variants: a state conditioned variant, trained and evaluated with proprioceptive joint state as an additional input, and a state-free variant, trained and evaluated on image-language pairs only, across 50 fully randomized trials per variant on a single manipulation task. The state-free variant achieves 66% task success versus 28% for the state-conditioned variant (z = 3.81, p < 0.001, two-proportion test), which we hypothesize is related to the added learning burden of jointly modeling proprioceptive state alongside vision and language during fine-tuning, though we do not establish this mechanism causally. Both models overfit beyond their optimal checkpoints (200k and 30k steps respectively) despite training to 400k, producing overconfident and unstable outputs — a practical finding for VLA checkpoint selection. Self-correction, where the robot visually re-assessed and redirected to the correct target, was observed in 14% of state-free trials exclusively. This study establishes a simulation baseline on this task and platform for VLA design choices prior to real-robot transfer.
|
| |
| 09:30-10:30, Paper FrPo6P.23 | |
| Mapless Indoor Room Finding Using Directional Signs and Room Number Recognition |
|
| Choi, Eun Hyuk | Department of Electronic Engineering, Kyung Hee University |
| Kim, Donghan | Kyung Hee University |
Keywords: Robot Vision, Navigation, Guidance and Control, Artificial Intelligence Systems
Abstract: Finding a target room in an indoor corridor environment is difficult to solve using only obstacle avoidance or local goal following. The robot must read directional signs at intersections, select the corridor branch that contains the target room, and then verify the room numbers beside the doors to reach the destination. In this paper, we propose a mapless semantic navigation system for indoor room finding without using a pre-built metric map. The proposed system uses directional signs and room numbers as semantic landmarks. Based on RGB camera perception, it determines both the corridor branch to follow and whether the target room has been reached. A YOLO-World-based detector first detects directional signs and door-plate regions. Then, a VLM and OCR read the room-number ranges on directional signs and the observed room numbers on door plates. The recognized semantic information is stored in a Scene Graph-based memory. The Rule Engine compares the target room number, sign ranges, and observed room-number sequence to determine the driving intent. Branch lock and target tracking conditions are also introduced to reduce incorrect branch selection and false approaches caused by intermediate room numbers. Experiments in an Isaac Sim indoor corridor environment show that the proposed system can find a target room using semantic landmarks without a prior map or predefined goal coordinates.
|
| |
| 09:30-10:30, Paper FrPo6P.24 | |
| Projective Monocular Depth Correction for 2D Occupancy Grid Mapping |
|
| Yoon, Chaehyun | Sookmyung Women's University |
| Lee, Alex | Sookmyung Women’s University |
Keywords: Robot Vision, Navigation, Guidance and Control, Sensors and Signal Processing
Abstract: 2D occupancy grid maps describe where a robot can travel and where obstacles are located, and are widely used for efficient robot navigation services. For cost-effective service robots, it is useful to update such maps from camera-based sensing rather than relying only on metric range sensors. Recent monocular depth models can convert RGB images into dense depth estimates, from which obstacle-height range measurements can be extracted as a camera-derived scan for 2D mapping. However, this scan is not a metric range scan. Frame-wise act either on the robot pose through SE(2) scan matching or on the whole scan through a single uniform scale, but these models cannot fully explain camera-induced range distortion. We propose a range-measurement correction method for camera-derived 2D occupancy mapping. A visual trajectory prior provides the initial map-frame pose, which is refined by a local SE(2) residual, while the camera-derived range measurements are corrected by a regularized determinant-normalized SL(3) measurement transform in scaled bearing and log inverse-range coordinates. The corrected ranges are fused by conservative ray casting along the original bearings without deforming the accumulated map. We evaluate the method against SE(2)-only and Sim(2)-scale baselines on indoor robot sequences, using LiDAR maps only as ground-truth references.
|
| |
| 09:30-10:30, Paper FrPo6P.25 | |
| Uncertainty-Aware Filtering for Robotic Grasping in Limited Observation |
|
| Choi, Youngtae | Pukyong National University |
| Lee, Munhaeng | Pukyong National University |
| Suh, Jinho | Pukyong National University |
Keywords: Robot Vision, Artificial Intelligence Systems, Robotic Applications
Abstract: Robust robotic grasping in limited observation environments remains challenging due to sensor occlusions and incomplete point clouds. Although deterministic shape completion networks can reconstruct missing geometry, they fail to quantify epistemic uncertainty, frequently leading to risky grasp proposals on hallucinated surface artifacts. To address this issue, we present a novel uncertainty-aware grasp filtering framework that directly bridges geometric perception and manipulation task spaces. We integrate Monte Carlo Dropout into a PoinTr backbone to estimate point-wise epistemic uncertainty based on the spatial variance of multiple stochastic forward passes. These point-wise uncertainties are then propagated into the grasp domain by aggregating them within the 3D bounding box of gripper candidates generated by GraspGen, formulating a density-independent grasp-space uncertainty metric. Leveraging a Top-K_s ranking strategy, our framework isolates unreliable surface artifacts and establishes a safety-verified grasp pool to select the optimal pose. Physics-based simulations in MuJoCo demonstrate that our framework substantially enhances both geometric reconstruction quality and downstream grasp success rates over conventional baselines.
|
| |
| 09:30-10:30, Paper FrPo6P.26 | |
| Occupancy-Grounded Room Segmentation for Hierarchical 3D Scene Graphs |
|
| Cueto Zumaya, Carlos Roberto | University of Turku |
| Catalano, Iacopo | University of Turku |
| Peña Queralta, Jorge | Zurich University of Applied Sciences |
| Bessa, Wallace M. | Federal University of Rio Grande Do Norte |
Keywords: Robot Vision, Artificial Intelligence Systems, Robotic Applications
Abstract: Hierarchical 3D scene graphs (3DSGs) for indoor robots organize geometric and semantic information across spatial scales, with a room layer that connects object-level perception to room-scale reasoning. Existing systems construct this layer from different spatial substrates (e.g., place clusters, wall planes, or segmentation outputs), and as a result, room nodes are not evaluated on a common geometric criterion. We present an occupancy-grounded 3DSG pipeline in which room nodes are anchored to tracked free-space regions derived from occupancy decomposition, giving each room an explicit polygonal footprint. We evaluate the pipeline on 12 Matterport3D scenes by matching predicted room polygons to annotated room instances and compare against Hydra, a representative state-of-the-art place-connectivity baseline. The results show that occupancy-grounded anchoring recovers substantially more room instances than place-connectivity construction, at the cost of lower precision, and that wall-accurate room boundaries remain an open problem for both methods. Code is available at https://github.com/crcz25/OccuSG.
|
| |
| 09:30-10:30, Paper FrPo6P.27 | |
| HEDGe-Net: Geometric Priors and a Design Pattern for Off-Road 3D LiDAR Semantic Segmentation |
|
| Ji, Youngseo | Gachon University |
| An, Jhonghyun | Gachon University |
Keywords: Robot Vision, Autonomous Vehicle Systems, Artificial Intelligence Systems
Abstract: Off-road 3D LiDAR semantic segmentation departs sharply from the urban regime that dominates the literature: vegetation is vertically continuous, the ground is class-bearing, and sparse distant safety-critical structures (poles, fences) dissolve into noise. We present a systematic study and a strong baseline on RELLIS-3D. (i) We characterize off-road failure as a four-mode taxonomy concentrating 96.4% of off-diagonal error mass, mapping each mode to a missing geometric prior and a specific inductive-bias mismatch of urban-tuned backbones. (ii) We report a new off-road state of the art - 49.24% float32 mIoU (the official, saturation-prone metric; 54.66% under a saturation-free int64 recount we introduce, for which no literature value yet exists) - from a single PT-v3 + 3D-SSL (Concerto) model on par with our prior five-member ensemble (49.04 / 54.45) - one model replacing five; a property of that recipe, not of the priors in (iii). A matched-recipe decomposition shows off-road performance is representation-bound: self-supervised initialization supplies nearly all of the gain, while post-hoc architectural priors give diminishing returns. (iii) We distill the discipline of safely augmenting a warm-init backbone into a reusable Design Pattern, instantiated by two transparently-reported geometric priors (HSE, DIG) whose headline effect is indistinguishable from zero at our seed budget - itself evidence for the representation-bound finding, not a separate performance claim.
|
| |
| 09:30-10:30, Paper FrPo6P.28 | |
| 3D Reconstruction of Tangled Deformable Linear Objects Using RGB-Guided Depth Fusion |
|
| Yoon, Youngbin | Chonnam National University |
| Hong, Ayoung | Chonnam National University |
Keywords: Robot Vision, Robotic Applications
Abstract: Accurate 3D shape reconstruction of Deformable Linear Objects (DLOs) such as ropes and cables is a fundamental prerequisite for robotic manipulation tasks, including knot-tying and cable routing. Existing methods rely either on color-based segmentation, which fails for monochrome objects, or provide only 2D centerline estimates without depth information, making it difficult to resolve crossing ambiguities under self-occlusion. We propose a training-free pipeline that reconstructs the full 3D shape of a monochrome DLO from RGB-D input. Boundary detection is performed by fusing RGB-estimated depth from Depth Anything V2 (DAV2) with sensor-measured depth from an Intel RealSense D435, where the two signals complement each other to capture boundaries that are otherwise difficult to detect. The detected boundaries partition the DLO mask into interior regions, from which centerline segments are connected via a direction-aware greedy matching algorithm. When the topology cannot be uniquely determined from a single observation, alternative candidates are maintained and refined with additional observations. Quantitative evaluation shows that our method achieves an F1 score of 97.8%, outperforming Canny, PiDiNet, and DiffusionEdge, and qualitative comparisons with mBEST and RT-DLO confirm its advantage for reconstructing monochrome tangled ropes.
|
| |
| 09:30-10:30, Paper FrPo6P.29 | |
| Comparison of SAM-Based Models for Tomato Stem Segmentation in Robotic Harvesting |
|
| Park, Jeongmo | Chonbuk National University |
| Seo, YongSeong | Jeonbuk National University |
| Park, Jaebyung | Jeonbuk National University |
Keywords: Robot Vision
Abstract: Accurate tomato stem segmentation is essential for determining cutting positions in robotic harvesting systems. Object detection can be used to localize target regions, whereas image segmentation provides pixel-level shape information for representing elongated stem structures. Several SAM-based models have recently been introduced for promptable segmentation, but their relative performance on tomato stems has not been sufficiently examined. This study compares four SAM-based models—MobileSAM, SAM, SAM2, and SAM3—using the same box prompts obtained from YOLOv8m-seg predictions. Quantitative and qualitative evaluations assess segmentation accuracy, boundary preservation, and inference time. SAM achieves the highest IoU and Dice scores, while SAM2 provides similar segmentation performance with a shorter inference time. MobileSAM shows greater boundary deviations in some cases, whereas SAM3 produces false-positive regions and misses substantial portions of the stem, resulting in lower segmentation performance. These results show that the performance of ROI-guided stem segmentation varies across SAM-based models and that model selection is important for robotic harvesting applications.
|
| |
| 09:30-10:30, Paper FrPo6P.30 | |
| Reflection-Robust Stereo Visual-Inertial Odometry Via Kalman-Filtered Ground-Height Estimation |
|
| Hwang, Uihyun | KAIST |
| Shin, Sungjae | Korea Advanced Institute of Science and Technology (KAIST) |
| Kim, Dongjae | KAIST |
| Myung, Hyun | KAIST (Korea Advanced Institute of Science and Technology) |
Keywords: Robot Vision, Autonomous Vehicle Systems, Artificial Intelligence Systems
Abstract: Visual-inertial odometry relies on visual features that correspond to 3D points with consistent projections across sequential frames. In reflective indoor-floor environments, however, specular floor patterns can be detected and triangulated as virtual features below the physical ground plane, producing unreliable feature associations and inconsistent visual residuals. This paper proposes a lightweight reflection-aware front-end for feature-based stereo VIO. Rather than suppressing all floor-region features, the proposed method estimates the ground height in real time and rejects only feature candidates whose inferred 3D positions are below the estimated ground height. A floor-region mask restricts the validation process to reflection-prone regions, while stereo height uncertainty and pose uncertainty are used to construct frame-level ground-height measurements. A scalar Kalman filter with Normalized Innovation Squared (NIS) gating then provides a stable ground-height estimate, and an uncertainty-aware safety margin is used for selective below-ground virtual feature rejection. Experiments on real reflective-floor indoor sequences show that the proposed front-end generally improves trajectory accuracy in reflection-dominant environments while maintaining accurate ground-height estimates and recovering from actual ground-height changes.
|
| |
| 09:30-10:30, Paper FrPo6P.31 | |
| Boundary-Aware Picking Center Generation for Foreign Object Removal in Continuous Dried-Red-Pepper Conveyor Frames |
|
| Shin, Sumin | Korea Institute of Industrial Technology (KITECH) |
| Lee, Yong Jun | Korea University |
| Inpyo, Lee | Korea Institute of Industrial Technology |
| Ahn, Woo Jin | Inha University |
| Myotaeg, Lim | Korea University |
| Lee, Kwang Hee | Korea Institute of Industrial Technology |
Keywords: Robot Vision, Robotic Applications, Industrial Applications of Control
Abstract: To bridge anomaly detection outputs with robotic manipulation, this work presents a decision-making layer tailored for spatial boundary resolution on a dried-red-pepper transport line. Standard frame-by-frame analysis often yields redundant picking actions because the same anomaly reappears across overlapping image boundaries. To mitigate this, our approach transforms candidate coordinates into localized targets via a smallest enclosing circle (SEC) formulation constrained by the effective suction radius Rend . Furthermore, the boundary region is partitioned into an overlap-body and a tail-zone, enabling dynamic assignment between immediate frame execution and cross-frame carry-over. In real conveyor test runs, the framework successfully streamlined target candidates from 29.68 to 12.58 target centers per frame while securing a finalized coverage rate of 1.000. Integrated physical testing demonstrated an overall extraction accuracy of 90.0% alongside a low latency of 97.32 ms for the vision-to-command pipeline.
|
| |
| 09:30-10:30, Paper FrPo6P.32 | |
| Bridging the Sim-To-Real Gap in Quadrupedal Visual Navigation Via 3D Gaussian Splatting |
|
| Lee, Minseok | Handong Global University |
| Han, Seongyu | Handong Global University |
| Mun, Huijae | University |
| Hwang, Sung Soo | Handong Global University |
Keywords: Robot Vision, Navigation, Guidance and Control, Robotic Applications
Abstract: Deploying visual navigation on physical quadrupedal robots is challenging due to the visual sim-to-real gap and locomotion-induced camera jitter. This jitter degrades feature tracking and depth consistency. This paper presents a pipeline that bridges these perceptual and dynamic gaps to enable virtual parameter optimization. In the simulation stage, we reconstruct a 3D Gaussian Splatting (3DGS) digital twin of a campus corridor and use a reinforcement learning (RL) locomotion policy to generate camera bobbing noise. Under these simulated visual and dynamic disturbances, the parameters of the RTAB-Map visual SLAM and Nav2 stacks are optimized. For deployment, the RL policy is bypassed, and velocity commands are routed directly to the physical robot’s built-in Sport API via a ROS 2 bridge. Real-world experiments demonstrate that parameters optimized under simulated jitter transfer directly to the physical platform, supporting visual mapping and qualitative obstacle-aware replanning without manual recalibration.
|
| |
| 09:30-10:30, Paper FrPo6P.33 | |
| Brick3D: Brick Detection and Localization Based on VLMs for Masonry Construction Drones |
|
| Valdes Saucedo, Mario Alberto | Lulea University of Technology |
| Stamatopoulos, Marios-Nektarios | Luleå University of Technology |
| Kanellakis, Christoforos | LTU |
| Nikolakopoulos, George | Luleå University of Technology |
Keywords: Robot Vision, Robotic Applications, Artificial Intelligence Systems
Abstract: This paper presents Brick3D, a monocular camera-based perception framework for UAV-assisted masonry construction that enables the detection, condition assessment, and metric localization of construction bricks using only onboard visual sensing. The proposed system leverages vision-language models (VLMs) to perform zero-shot semantic segmentation, allowing robust brick detection without task-specific retraining and enabling adaptation to varying brick appearances and construction environments. A coverage-based integrity assessment method evaluates the correspondence between segmented regions and their expected geometry, classifying bricks as undamaged, damaged, or uncertain under partial visibility conditions. For localization, the framework combines refined corner observations with a Perspective-n-Point (PnP) formulation to estimate the full 6-DoF pose of each brick relative to the onboard camera. To improve monocular localization accuracy, an offline nonlinear calibration model and Kalman filtering stage are introduced to compensate for systematic bias and reduce temporal noise. Experimental validation in both controlled laboratory and real-world indoor and outdoor environments demonstrates reliable brick detection, robust damage classification, and accurate localization.
|
| |
| 09:30-10:30, Paper FrPo6P.34 | |
| Real-Time Classification and Pose Estimation of Humans and Legged Humanoids |
|
| Kang, GyeongRim | JeonBuk National University |
| Kim, KangGeon | Jeonbuk National University |
| Jo, HyungGi | Jeonbuk National University |
Keywords: Robot Vision, Human-Robot Interaction, Artificial Intelligence Systems
Abstract: Humans and legged humanoids may be observed in the same workspace. The difficult cases considered in this study occur when their image trajectories cross and a temporary occlusion interrupts a track. After re-entry, a detection can then be associated with the wrong entity. We address this case with a single-RGB pipeline that performs 2D pose estimation, classification, tracking, and identity recovery together in near real-time. A single-stage multi-person pose estimator provides the keypoints, MobileNetV3-Small supplies track-level class evidence through temporal voting, and a dual-track Re-ID stage handles lost identities. On the recorded occlusion and re-entry sequences, the classifier reached 98.99% accuracy and a macro F1 score of 0.986. The complete pipeline used 181 MB of GPU VRAM on the target platform.
|
| |
| 09:30-10:30, Paper FrPo6P.35 | |
| Adaptive Object Localization with Depth Reliability-Aware Innovation-Gated Kalman Filter for Small Swarm Robot Systems |
|
| Kim, Seonggeon | Kumoh National Institute of Technology |
| Lee, Heoncheol | Kumoh National Institute of Technology |
Keywords: Robot Vision, Robotic Applications, Sensors and Signal Processing
Abstract: This paper addresses unstable RGB-D-based object localization on resource-constrained small robot platforms. Noisy depth measurements from individual robots can produce inconsistent object position estimates, which may subsequently degrade object-level map merging. RGB-D depth measurements are susceptible to noise, occlusion, and transient depth spikes, causing inconsistent object positions and duplicated landmarks in terms of map merging. To address this, we propose Depth Reliability-aware Innovation-Gated Adaptive Kalman filter (DRIG-AKF), which stabilizes object position estimates by quantifying depth reliability from valid depth ratio, depth variance, and object distance, and incorporating it into an adaptive measurement covariance. An innovation-gated mechanism suppresses transient depth spikes via Mahalanobis innovation distance gating, and the stabilized positions and their covariance estimates can be transmitted as inputs to a subsequent map-merging stage. The proposed depth reliability model and innovation gating are directly validated using a separate recording with raw per-pixel depth measurements. The overall framework is evaluated on three real-world datasets using bounding-box-derived reliability proxies due to the absence of per-pixel depth logs, demonstrating an average position standard deviation reduction of 66.9% over EMA smoothing and 27.5% over a KF+Hungarian baseline, with reductions of up to 79.5% and 54.1%, respectively, in datasets where the reliability indicators more accurately reflect depth measurement quality.
|
| |
| 09:30-10:30, Paper FrPo6P.36 | |
| Preserving Fine Details in Semantic-Map-Conditioned LiDAR Generation Via Pixel-Space Diffusion |
|
| Lee, Jaewon | Korea Institute of Industrial Technology |
| Kim, Jiwoong | Korea Institute of Industrial Technology(KITECH) |
Keywords: Robot Vision, Artificial Intelligence Systems, Autonomous Vehicle Systems
Abstract: LiDAR sensors provide precise 3D measurements of the surrounding environment and are widely used for perception in autonomous driving and robotics. However, collecting real driving data that covers diverse and rare situations is costly, which has motivated research on synthetic LiDAR scene generation using diffusion models. To reduce computation, most existing methods generate scenes in a compressed latent space and decode them back to the original resolution. This decoding is lossy, and the loss becomes critical when a semantic map is used as a condition, since the condition is also compressed and its detailed structure is not fully preserved. In this paper, we adapt PixelDiT, a pixel- space diffusion transformer, to conditional LiDAR range-image generation. Instead of encoding the range image and the semantic map into a compressed latent space, we inject the semantic map as tokens at native resolution. This helps the generated scan retain thin and small-scale structures that a latent-space model tends to blur. On the SemanticKITTI Semantic-Map-to-LiDAR task, the proposed method follows the semantic condition more faithfully than a latent-space baseline.
|
| |
| 09:30-10:30, Paper FrPo6P.37 | |
| Keyframe-Based LiDAR Scan-To-Map Registration in a Greenhouse Corridor |
|
| Park, Taewook | Ulsan National Institute of Science & Technology |
| Lee, Seunghoon | Korea Advanced Institute of Science and Technology |
| Yun, Won-Jae | METAFARMERS Inc |
| Lee, Kyu-Wha | METAFARMERS Inc |
| Oh, Hyondong | KAIST |
Keywords: Robot Vision, Autonomous Vehicle Systems
Abstract: Greenhouse robots require reliable localization despite unavailable or unreliable Global Navigation Satellite System signals and limited geometric features along crop corridors. We present a LiDAR-based localization method that registers multi-session scans to a unified prebuilt point cloud map using point-to-plane iterative closest point registration and LiDAR--inertial odometry priors. Evaluation in a 50 m-long commercial tomato greenhouse corridor demonstrates consistent multi-session localization in a common map frame, with position and orientation RMSEs below 3.5 cm and 4.8 degree, respectively.
|
| |
| 09:30-10:30, Paper FrPo6P.38 | |
| OR-MR: Structure-Aware Multi-Candidate Planning for Multi-Depot Grid Coverage |
|
| Seo, JangHo | Kyungpook National University |
| Lee, Joonwoo | Kyungpook National University |
Keywords: Robotic Applications, Navigation, Guidance and Control, Artificial Intelligence Systems
Abstract: Multi-depot multi-robot coverage path planning aims to cover all free cells while each robot starts from and returns to its own depot, making the longest closed route the mission-time bottleneck. Existing grid-based planners usually commit to a single decomposition or growth rule, and their makespan can change significantly with obstacle bottlenecks and depot placement. This paper presents OR-MR (Ownership-Routed Multi-Robot), a structure-aware bounded heuristic portfolio for 4-connected grid coverage. OR-MR constructs solutions on a common 2×2 supercell substrate, applies a rule-based Cut Rule using map-cut and depot-geometry features to activate a small set of synchronized balanced supercell growth (SBSG) and Peak-Balance candidates (observed K=1–13 per scenario), and selects the final plan by route-level closed makespan. On 18 corner-anchored scenarios over six 50×50 maps, OR-MR improves the closed makespan in most cases against three representative baselines, with the largest gains on the split-rich Map 4.
|
| |
| 09:30-10:30, Paper FrPo6P.39 | |
| ELite++: Efficient Lifelong LiDAR Mapping Via Window-Based Registrationand Voxel Ratio Maps |
|
| Lee, GeonHui | Seoul National University |
| Gil, Hyeonjae | SNU |
| Kim, Ayoung | Seoul National University |
| Jung, Minwoo | Seoul National University |
Keywords: Robotic Applications, Autonomous Vehicle Systems
Abstract: Long-term LiDAR mapping requires clean and consistent static maps across repeated sessions in dynamic environments. Existing lifelong mapping pipelines combine session alignment, dynamic filtering, and map update, but repeated dense point-wise updates and candidate-dependent alignment can limit scalability. To tackle these issues, we present ELite++, an efficient extension of ELite that combines automatic window-based registration, aligned-session dynamic object removal, and compact Voxel Ratio Map integration. ELite++ removes transient objects while preserving static structures through hit-exposure voxel stability, local ground protection, and structure-aware refinement. Experiments on diverse LiDAR datasets show improved dynamic object removal over existing DOR methods, preserved multi-session alignment quality, and runtime reduced to 11% of the original pipeline.
|
| |
| 09:30-10:30, Paper FrPo6P.40 | |
| Human-In-The-Loop Replanning for Instruction-Guided Object Retrieval in Fire Environments |
|
| Song, Yuri | Kwangwoon University |
| Kim, Yeonjin | Kwangwoon University |
| Oh, Junghyun | Kwangwoon University |
Keywords: Robotic Applications, Human-Robot Interaction, Navigation, Guidance and Control
Abstract: Disaster-response robots must operate under evolving hazards while respecting user instructions. In fire environments, predicted object risk can conflict with the retrieval priority implied by a natural-language instruction: a lower-priority target may become urgent, whereas automatic risk-based replanning may violate the user’s intent. This paper formulates this problem as instruction-conditioned object retrieval under dynamic fire risk. We propose a Priority- Aware Human-in-the-Loop (PA-HITL) replanning framework that extracts target objects and a preferred order, estimates predicted object risk, and presents reorder suggestions when risk-based urgency conflicts with instruction-implied priority. The final execution target is determined through accept/reject decisions instead of automatic reordering. Experiments in the HAZARD fire environment show that, compared with fully automatic risk-based replanning, PA-HITL improves SR from 0.55 to 0.60 in Kitchen and from 0.65 to 0.71 in Craftroom, while maintaining comparable Safe-SR (0.35 vs. 0.36 in Kitchen and 0.40 vs. 0.41 in Craftroom). Compared with No HITL, PA-HITL also improves Safe-SR from 0.20 to 0.35 in Kitchen and from 0.25 to 0.40 in Craftroom.
|
| |
| 09:30-10:30, Paper FrPo6P.41 | |
| Robotic Manipulator Trajectory Stabilization Via Imitation-Guided Reinforcement Learning |
|
| Lee, Hojeong | Korea University |
| Cho, Hyemin | Korea University |
| Lee, Seunghoon | Korea University |
| Park, Cheol Hoon | Korea University |
| Lee, Yong Jun | Korea University |
| Park, Jong-Chan | Korea University |
| Lim, Myo-Taeg | Korea University |
Keywords: Robotic Applications, Robot Mechanism and Control, Artificial Intelligence Systems
Abstract: Imitation learning enables agents to acquire skills by observing and learning from expert demonstrations, and has recently attracted significant attention in robotics. However, these demonstrations do not guarantee optimal behavior, and deploying trajectories generated by imitation learning policies in real-world environments can result in abrupt joint movements, instability, and compounding errors. To address these issues, we propose an imitation-guided reinforcement learning framework that leverages imitation learning trajectories as reference paths and learns a stable trajectory-tracking policy through reinforcement learning. Specifically, we generate reference joint trajectories using an action chunking transformer-based imitation learning policy, and further optimize a control policy via proximal policy optimization to track the trajectories while suppressing joint velocities and action variations. Experimental results demonstrate that the proposed framework reduces joint oscillations by approximately 26% compared to baseline imitation learning, confirming that reinforcement learning enhances control stability while maintaining the task-level behaviors acquired through imitation.
|
| |
| 09:30-10:30, Paper FrPo6P.42 | |
| Temporal Context Conditioning for Action Chunk Continuity in Vision-Language-Action Model |
|
| Shin, Dongin | Korea Electronics Technology Institute |
| Moon, JongSul | Korea Electronics Technology Institute |
| Kim, YoungOuk | Korea Electronics Technology Institute |
| Kim, Wonha | Kyung Hee University |
Keywords: Robotic Applications, Robot Vision, Artificial Intelligence Systems
Abstract: Vision-Language-Action (VLA) models have advanced robot action generation, yet existing approaches lack explicit modeling of temporal continuity at action chunk boundaries, leading to trajectory discontinuity and abrupt velocity changes in long-horizon tasks. This paper proposes Temporal Context Conditioning, a framework that injects the previous action trajectory and vision-language hidden representation into the Flow Matching-based Diffusion Transformer (DiT) via Feature-wise Linear Modulation (FiLM). A Boundary Gap Metric is introduced to quantify inter-chunk discontinuity. On the LIBERO benchmark, the proposed method improves the average success rate by 6.0% over the baseline and reduces the Boundary Gap by 15.8%.
|
| |
| 09:30-10:30, Paper FrPo6P.43 | |
| IRM-Guided Model Predictive Control for Mobile Manipulator |
|
| Choi, JungHyun | University of Seoul |
| Lee, Taegyeom | University of Seoul |
| Hwang, Myun Joong | University of Seoul |
Keywords: Robotic Applications, Robot Mechanism and Control
Abstract: This paper presents an inverse reachability map (IRM)-guided model predictive control framework for end-effector trajectory tracking by mobile manipulators. A precomputed three-dimensional IRM is used to evaluate base regions that are kinematically favorable for each desired end-effector pose. At each prediction step, the IRM is sliced into a two-dimensional base preference map and combined with environmental safety information using a multiplicative safety mask. The local weighted centroid around the predicted base position is then used as an IRM-guided reference in the mobile base MPC objective. The overall structure of the proposed framework is shown in Fig. 1. Unlike conventional IRM-based methods that mainly determine static base placements or select discrete candidate poses, the proposed method incorporates IRM information online to continuously guide the mobile base toward regions that support end-effector reachability and favorable arm manipulability. The manipulator controller tracks the global end-effector reference in the current base frame, while the base MPC accounts for the IRM-guided reference, heading consistency, and input regularization. Gazebo simulations with a Scout 2.0 mobile base and a Franka Panda manipulator show that the proposed IRM-MPC improves the base configuration compared with the baseline, as illustrated in Fig. 2, and reduces end-effector position tracking error while maintaining more consistent Yoshikawa manipulability, as shown in Fig. 3.
|
| |
| 09:30-10:30, Paper FrPo6P.44 | |
| Shared Autonomy for Teleoperation Using Probabilistic Estimation of Primitive Motion Intent |
|
| Lee, Haeseong | Seoul National University |
| Yoon, Junheon | Seoul National University |
| Lee, Yonghee | Seoul National University |
| Park, Jaeheung | Seoul National University |
Keywords: Robotic Applications, Human-Robot Interaction, Artificial Intelligence Systems
Abstract: This paper proposes a shared autonomy method that enables a teleoperated robot system to infer human intent. Teleoperation has significant potential for hazardous-environment operations and human demonstration data collection. However, teleoperation remains challenging due to limited feedback, communication latency, and embodiment mismatch between humans and robots. To improve teleoperation efficiency, various shared autonomy methods, such as policy blending, have been actively studied. However, existing approaches often require task-specific robot actions and may reduce the operator’s control authority. To address these limitations, the proposed method focuses on estimating the primitive motion intent of the operator, including translation-only, rotation-only, and combined motions. For this purpose, a lightweight estimator is developed to predict the probability of each motion primitive. Also, based on the probabilistic estimation, unintended human command is suppressed while preserving the control authority. Finally, real-robot experiments demonstrate that the proposed framework improves task efficiency during teleoperation.
|
| |
| 09:30-10:30, Paper FrPo6P.45 | |
| Actuator-Aware Inverse Kinematics with Joint-Limit Admissibility for Torque-Controlled Redundant Robots |
|
| Dastranj, Mohammad | Tampere University |
| Hejrati, Mahdi | Tampere University |
| Mattila, Jouni | Tampere University |
Keywords: Robotic Applications, Robot Mechanism and Control, Navigation, Guidance and Control
Abstract: This paper proposes actuator-aware inverse kinematics for torque-controlled redundant robots under joint-limit constraints. In the considered architecture, the inverse-kinematic output is not merely a purely kinematic joint-velocity command; it is the required joint velocity supplied to a downstream torque-level controller. Therefore, a small commanded task residual may not necessarily improve realized motion. The proposed method formulates a convex quadratic programming problem whose decision variable is the joint-level required velocity. Control barrier function style bounds impose reference-level joint-limit admissibility, while the task equation is handled through a penalized slack variable. Redundancy is resolved using a controller-compatibility objective that accounts for previous-command consistency and actuator torque-capacity weighting. The method is independent of the particular torque-level controller and can serve as an intermediate IK layer between an endpoint trajectory and a redundant robot controller. Experiments on a virtual-decomposition-controlled seven-degree-of-freedom upper-limb exoskeleton compare the method with standard inverse-kinematic baselines and a constrained task-preserving quadratic programming baseline. The results indicate substantially lower limit-pushing commands, bounded admissible required velocities, and reduced maximum position error while maintaining competitive realized task tracking in the tested trajectory, without modifying the downstream controller.
|
| |
| 09:30-10:30, Paper FrPo6P.46 | |
| Design and Dynamic Modeling of a Hybrid Compliant End-Effector with Feedforward Compensation for Robotic Surface Processing of Reclaimed Timber |
|
| Elsayed, Ahmed Sedky Mohamed | NTNU |
| Garammatikos, Sotirios | NTNU |
Keywords: Robotic Applications, Robot Mechanism and Control, Control Theory and Applications
Abstract: Automated surface preparation of reclaimed timber remains a key barrier to wider structural reuse. The main challenge is maintaining consistent contact force over geometrically uncertain surfaces that combine hard contaminants with a soft, damage-sensitive substrate. This paper presents the design and dynamic modeling of a hybrid compliant end-effector that pairs a passive spring-based compliance stage with an active motor-driven feedforward compensator. The mechanism comprises four linear guides, eight bearings, two compression springs (equivalent stiffness 500 N/m), a force sensor, and two anti-twisting guides. The system is modeled as a Simscape Multibody digital twin and evaluated under a ±5 mm sinusoidal surface-height disturbance with an 18 N force setpoint. The feedforward-augmented PID controller reduces the initial contact transient by 46.7%, steady-state RMS force error by 30.8% (from 1.03% to 0.71% of setpoint), and mean absolute error by 32.4% compared with PID alone.
|
| |
| 09:30-10:30, Paper FrPo6P.47 | |
| Demonstration Temporal Resolution Analysis for Visuomotor Diffusion Policy in Precision Electrical Component Assembly |
|
| Kim, Ikjune | Korea Atomic Energy Research Institute |
| Joo, Sungmoon | Korea Atomic Energy Research Institute |
Keywords: Robotic Applications, Artificial Intelligence Systems, Robot Mechanism and Control
Abstract: The automation of electrical component assembly remains challenging due to tight-tolerance insertion requirements. Purely visuomotor approaches offer practical deployment advantages, but introduce an underexplored data collection parameter: demonstration temporal resolution, which refers to the robot's travel distance between consecutive frames in the demonstration training data. It is defined as the per-frame end-effector displacement determined by the robot's end-effector speed and data sampling frequency. This work formalizes its constraints in diffusion-based imitation learning, showing that the feasible range is bounded below by the robot's kinematic repeatability (preventing state aliasing) and above by the task's geometric clearance (avoiding intra-step collisions). Experiments across nine temporal resolution settings with matched training-data volume identify the feasible range and derive experimentally calibrated data collection guidelines, validated on a PDU surrogate with 2.0 mm radial clearance across 150 trials, achieving 80%–92% per component and overall 85%, using the alternative hardware setup.
|
| |
| 09:30-10:30, Paper FrPo6P.48 | |
| Characterizing Hard-To-Automate Manufacturing Processes for Robotic Automation: Data-Driven Difficulty Factors and a Mapping to Robot Technologies |
|
| Lee, Jaeseon | KITECH |
| Im, Subin | SungKyunKwan University |
Keywords: Robotic Applications, Industrial Applications of Control, Artificial Intelligence Systems
Abstract: Korea has the world’s highest manufacturing robot density, yet many processes are still performed manually, and judging which ones are hard to automate has relied on inconsistent qualitative judgment. We build a quantitative basis for that judgment from 593 robotic-automation consulting cases (2016–2025). Of these, 353 non-trivial processes were assessed against eighteen difficulty factors spanning the object, task, and environment. Rather than scoring expert verdicts directly, we reverse-engineer the relationship between the experts’ factor checks and their grades through ordinal regression, yielding data-based weights and seven core factors (led by shape complexity and skill dependency) from which we derive an integer scoring scheme; suppressor factors are identified and excluded. The seven-factor scheme reproduces the expert grades about as well as the full model while remaining simpler (75.9% vs. 74.4% accuracy for high-or-above), and its scores follow the known difficulty ranking across process groups. We further map the core factors onto the enabling technologies (perception, grasping, contact/precision control, and task planning). The gap between the objective-factor model and an expert baseline quantifies how far the judgment can be made explicit, with a residual tied to experience, most clearly in skill dependency. External validation is planned through a national R&D program (2027–2030).
|
| |
| 09:30-10:30, Paper FrPo6P.49 | |
| STG 3.0: Time-Dependent Semantic Topological Graphs for Human-Aware Service Robots |
|
| Park, Jeong-Seop | Korea University |
| Lee, Yong Jun | Korea University |
| Park, Jong-Chan | Korea University |
| Ahn, Woo Jin | Inha University |
| Woo, Jong Jin | LG Electronics |
| Lim, Myo-Taeg | Korea University |
Keywords: Robotic Applications, Artificial Intelligence Systems, Human-Robot Interaction
Abstract: Household service robots must often locate a target user before providing services such as delivery or assistance. However, existing navigation and Vision–Language–Action (VLA) frameworks mainly rely on static semantic information and do not explicitly model human location over time. This paper presents STG 3.0, a time-dependent semantic topological graph for proactive human search. STG 3.0 combines a static Home-Tree representation with dynamic human nodes connected through time-dependent probabilistic edges learned from patrol observations. By integrating user-location probabilities into graph-based planning, the robot prioritizes locations where the target user is most likely to be found. Experiments in simulated and real indoor environments demonstrate reliable semantic mapping and reduced search time compared with conventional search strategies.
|
| |
| 09:30-10:30, Paper FrPo6P.50 | |
| Sensorless Admittance Control Based on Velocity Response Deviation under a Low Control Update Rate |
|
| Kim, Jiseong | Keimyung University |
| Jo, Jaemin | Keimyung University |
| Choi, Junghyun | Keimyung University |
Keywords: Control Theory and Applications, Human-Robot Interaction, Industrial Applications of Control
Abstract: This paper presents a sensorless admittance-control structure for a speed-controlled motor drive under a low control update rate. Velocity-response deviation between the reference and encoder-based angular velocity is converted into a torque-equivalent signal and applied to a virtual inertia-damping model without an external force/torque (F/T) sensor or internal drive signals. Experimental results confirmed force-responsive reference modification consistent with torque-sensor-based control and demonstrated feasibility under the low update-rate condition.
|
| |
| FrPo7P |
Symphony D, 4F |
| Poster Session 7 |
Poster Session |
| |
| 16:10-17:10, Paper FrPo7P.1 | |
| DigitForce: Channel-Decoupled Force Tokenization for Transformer-Based Humanoid Dexterous Manipulation |
|
| Kim, Beomjoon | Korea Electronics Technology Institute |
| Kim, Yunhan | Korea Electronics Technology Institute |
Keywords: Robotic Applications, Artificial Intelligence Systems, Sensors and Signal Processing
Abstract: Contact-rich humanoid manipulation depends on finger-localized contact cues that are difficult to infer from RGB and joint states. We present DIGITFORCE, a channel-decoupled state tokenizer that addresses a fused-state token bottleneck in Action Chunking Transformer (ACT)-based imitation learning. In the standard fused design, proprioception and force readings are projected into one state token, limiting channel-level access to transformer attention. DIGITFORCE keeps the RGB encoder, ACT backbone, decoder, action head, training objective, and deployment interface unchanged. Proprioception is encoded as one token, and each finger-associated force channel is encoded as a separate token. Force readings are used only as observations, with no change to the action space or controller. On a real Unitree G1 with dual five-finger Inspire Hand DFX hands performing a color-sequence block-stacking task, task success rates across 50 rollouts per policy are 62%, 76%, and 92% for Proprio-ACT, Concat-ACT, and DIGITFORCE, respectively, and DIGITFORCE shows shorter failed-grasp recovery intervals and smaller arm-retreat excursions than the fused-token baseline. Matched offline force-channel interventions show that DIGITFORCE produces sharper phase-dependent channel selectivity than the fused-token baseline and induces larger action shifts in hand-joint dimensions. The mean hand-to-arm sensitivity ratio increases from 1.05 in the fused-token baseline to 1.40 in DIGITFORCE. These results suggest that channel-decoupled force tokens provide channel-addressable contact cues for hand-joint action prediction, without tactile skin, a force-control head, or a new action interface.
|
| |
| 16:10-17:10, Paper FrPo7P.2 | |
| Energy-Cost-Based Prioritized Path Planning for Multiple Ground-Aerial Bimodal Robots |
|
| Lee, Dongeun | Kookmin University |
| Lee, Seung-Mok | Kookmin University |
Keywords: Robotic Applications, Navigation, Guidance and Control
Abstract: This paper presents an energy-cost-based prioritized path-planning method for multiple passive ground-aerial bimodal robots operating in obstacle environments. Passive bimodal robots can reduce energy consumption through ground locomotion and overcome obstacles through aerial locomotion, but aerial motion requires significantly higher energy to support the robot weight. To address this trade-off, the proposed method evaluates motion primitive costs using a mode-dependent energy model and constructs a Dijkstra-based heuristic from unit-distance energy costs. Robot paths are planned sequentially according to a fixed priority order using an energy-cost A∗ search over safe interval path planning(SIPP) states. The paths of higher-priority robots are stored in a reservation table and used as time-dependent constraints for lower-priority robots. Simulations were conducted in two obstacle environments: a narrow-gap environment that induces inter-robot congestion and a wall-obstacle environment that requires aerial locomotion. The results show that the proposed bimodal planner reduces travel time compared with ground-only planning and reduces energy consumption compared with aerial-only planning, demonstrating collision-free and energy-efficient multi-robot path planning.
|
| |
| 16:10-17:10, Paper FrPo7P.3 | |
| Automated Strawberry Packaging System Based on Vision Sensing and Re-Orientation Gripper Mechanism |
|
| Shin, WooSeong | Hanyang University |
| Lee, Minsu | Hanyang University |
| Choi, Jeongseok | Hanyang University |
| Seo, TaeWon | Hanyang University |
Keywords: Robotic Applications, Robot Vision, Robot Mechanism and Control
Abstract: This study proposes an integrated robotic system designed to pick randomly arranged harvested strawberries without damage and automatically place them into egg-tray-style containers. To address the growing trend of egg-tray-style strawberry packaging and the challenge of labor shortages, this system implements an automated process that combines mechanical adaptability with vision sensing. First, a two-fingered fin-ray gripper fabricated via silicone casting was utilized to achieve passive adaptive grasping of irregularly shaped strawberries while minimizing physical damage. At wrist, remote center of motion (RCM) mechanism wrist was introduced to ensure efficient placement angles within narrow containers. For the grasping strategy, YOLOv8-pose is employed to extract four keypoints(top, bottom, left, and right) of the strawberry. These points are used to determine the grasping position and angle and estimate the fruit's weight for grade classification. Experimental results confirm that the entire process—from recognition and grasping to weight estimation and packaging—is successfully performed within the integrated system, demonstrating the practical potential for automating strawberry packaging processes that currently rely heavily on manual labor.
|
| |
| 16:10-17:10, Paper FrPo7P.4 | |
| Pre-Execution Feasibility Checking for Language-Guided Robotic Manipulation Via 3D Scene Graph and Object Ontology |
|
| Kim, Haryeong | Sungkyunkwan University |
| Kuc, Tae-Yong | Sungkyunkwan University |
Keywords: Robotic Applications, Artificial Intelligence Systems, Robot Vision
Abstract: Pre-execution reasoning about whether a robot can successfully carry out a natural-language command remains a key challenge in language-guided service robotics. Existing approaches predominantly rely on geometry-based grasp planning while overlooking semantic object properties such as weight, movability, and gripper compatibility, often resulting in repeated attempts to execute physically infeasible actions. This paper presents a reasoning framework that integrates 3D Scene Graph construction with Object Ontology Extraction to assess the feasibility of commanded actions before execution. Using RGB-D scans, the framework automatically constructs per-object ontological representations across three layers—Symbolic, Explicit, and Implicit—providing the semantic grounding required for feasibility reasoning beyond geometric appearance. Lift feasibility is estimated by combining a category-level knowledge base with commonsense inference from a large language model; grasp feasibility is determined from point-cloud geometry relative to gripper constraints; movability is inferred from an implicit ontological attribute; and reachability is verified using the manipulator’s inverse kinematics solver. The extracted attributes and spatial relationships are organized into a 3D Scene Graph, enabling language-conditioned grounding of target objects. A feasibility function, F(o), jointly evaluates these four predicates to determine whether the commanded action should be executed. When any predicate is not satisfied, the framework returns interpretable feedback to the planning layer. Experiments conducted in Isaac Sim using a wheeled mobile manipulator equipped with a UR5 arm demonstrate that the proposed reasoning layer reduces unnecessary execution attempts from 37.5% to 10.0% and improves the task success rate from 62.5% to 85.0% compared with a geometry-only baseline.
|
| |
| 16:10-17:10, Paper FrPo7P.5 | |
| Real-Time Evaluation and Few-Shot Hardware Adaptation of a Vision-Based Tactile Sensor |
|
| Park, Juhee | KAIST(Korea Advanced Institute of Science and Technology) |
| Min, Seojung | KAIST |
| Kim, Jung | KAIST |
Keywords: Robotic Applications, Sensors and Signal Processing
Abstract: Vision-based tactile sensing is increasingly judged by offline test-set accuracy, yet the quantity that matters for robotics is accuracy under live contact on the deployed hardware. Reported accuracy can fall under real-time robotic contact and degrade further when the gel pad or sensor unit is replaced. We quantify both effects on a DIGIT sensor mounted on a 7-DOF Kinova Gen2 arm classifying five sandpaper grit levels (P100–P400). A model that reaches 96.76% on a static test set drops to 88.2% in real time on the same hardware, and falls to 78.0% on a previously unseen sensor-gel configuration. Swapping the gel pad alone accounts for most of the variation we observe across configurations seen during training. Fine-tuning on as few as 14 images per class recovers the hardware-induced gap, reaching a peak of 94.8% at 70 images per class. Offline accuracy is therefore an unreliable indicator of deployed performance, but the cost of adapting to a new hardware configuration can be small.
|
| |
| 16:10-17:10, Paper FrPo7P.6 | |
| Development of a Welding Torch Collision Avoidance Algorithm for Curved Weld Seam Tracking on 3D Pre-Deformed Aluminum Ship Structural Members |
|
| Park, In-Gyu | KIRO, Korea Institute of Robot and Convergence |
Keywords: Robotic Applications, Industrial Applications of Control, Control Devices and Instruments
Abstract: This paper proposes a welding torch collision avoidance algorithm for aluminum ship structural members with 3-dimensional curved lines. In the target CTV fabrication process, the aluminum member is pre-deformed to compensate for predicted welding distortion, and jig clamps are installed to maintain the deformed shape. These clamps can become obstacles during robotic welding. To address this problem, a collaborative robot-based welding system using the HCR14 robot was modeled, and a torch path generation method was developed. The curved weld seam was generalized using a cubic polynomial fitting algorithm based on the least squares method. The fitting accuracy was evaluated using RMSE, and the errors in the x, y, and z directions were 0.5656 mm, 0.2395 mm, and 7.9964×10−11 mm, respectively. Torch posture was controlled using Euler angles to follow the curvature-center direction while satisfying the work angle and travel angle. In addition, an artificial potential field method was applied to avoid collision by correcting the pitch angle corresponding to the work angle. The proposed algorithm is formulated in three-dimensional coordinates, however, its basic behavior was verified through a DAFUL-based simulation using a two-dimensional curved path in the x-y plane without height variation.
|
| |
| 16:10-17:10, Paper FrPo7P.7 | |
| Multi-Floor Navigation System for Indoor Autonomous Service Robot |
|
| Kwon, Do Young | Pukyong National University |
| Lee, Joon Ho | Pukyong National University |
| Choi, Woo Young | Pukyong National University |
Keywords: Robotic Applications, Rehabilitation Robot, Navigation, Guidance and Control
Abstract: This paper proposes a multi-floor navigation system considering the inter-floor movement of autonomous service robots in an indoor environment. Conventional navigation systems show effective performance in a single-floor environment, but for various services and stable driving, a navigation system considering inter-floor movement in multi-floor environments is required. For this purpose, we estimate the state of the robot based on the Kalman Filter using a dynamic model representing the motion that occurs during inter-floor movement from the Inertial Measurement Unit sensor. The floor decision of the robot is made using the Gaussian Mixture Model-based likelihood, in which the estimated state and covariance obtained from the KF are jointly considered. The floor decision result, including the moved distance and floor information, is integrated with the conventional navigation system, enabling the autonomous service robot to perform stable autonomous navigation across floors in an indoor environment. To validate the proposed multi-floor navigation system, a scenario-based experiment was conducted in an indoor multi-floor environment, where the autonomous service robot moved across four floors. The results confirmed that the robot accurately reached the target destination, demonstrating the feasibility of the proposed system.
|
| |
| 16:10-17:10, Paper FrPo7P.8 | |
| SONA: Scanning Once for N Simulation-Ready Assets from 2D Gaussian Splatting for Robotic Manipulation |
|
| Giri, Na | Korea Electronics Technology Institute |
| Kim, Beomjoon | Korea Electronics Technology Institute (KETI) |
| Oh, Saemyung | Korea Electronics Technology Institute |
Keywords: Robotic Applications, Robot Vision, Industrial Applications of Control
Abstract: Scalable simulation of real environments is increasingly central to robot learning, as policies are routinely trained and validated in digital twins (DT) before deployment on hardware. Learning manipulation policies in DT requires per-object simulator assets, which are conventionally modeled, textured, and assigned collision geometry and physical parameters by hand, and rebuilt whenever the scene changes. Asset preparation is thus a primary bottleneck in DT construction. Gaussian Splatting (GS) reconstructs photorealistic scenes from multi-view images, but its output is neither segmented into objects nor endowed with physical properties. This work addresses the scene-to-asset conversion as a distinct problem and presents SONA, a pipeline that converts a single GS reconstruction of a work cell into N per-object, simulation-ready assets. Each object is separated by a training-free vertex-labeling scheme that back-projects video segmentation masks onto a full-scene Truncated Signed Distance Field (TSDF) mesh, paired with a watertight visual mesh and a convex collision mesh in a common coordinate frame, and exported in the USD and URDF formats. On an industrial work cell, SONA reduces asset-generation time by up to 74% at N=6 relative to per-object scanning, and the exported assets reach a 94.2% pick success rate when used for policy learning in Isaac Lab without manual format conversion or physics-parameter calibration. On the one object captured in both scan modes, with identical policy initialization and training budget per asset, the scene-scan asset stays within 1.1% of the dedicated object-scan asset at a one-second hold criterion.
|
| |
| 16:10-17:10, Paper FrPo7P.9 | |
| Robot Tool Path Planning for Automated Peeling of Frozen Tuna |
|
| Jeong, Woong | Korea Institute of Industrial Technology, Korea University |
| Ahn, Woo Jin | Inha University |
| Shin, Myeongchan | KITECH |
| Lim, Myo-Taeg | Korea University |
| Nam, Yun Seok | Tech University of Korea |
| Lee, Kwang Hee | Korea Institute of Industrial Technology |
Keywords: Robotic Applications
Abstract: Bluefin tuna is commonly processed into high-value products, and therefore maintaining edible yield during processing is essential. Since the skin is directly attached to the edible flesh, skin removal critically determines the final yield. Nevertheless, the skin is still removed manually with a hand-held planar grinder, making the process labor-intensive and inconsistent in material removal. Achieving consistent material removal therefore requires robotic machining with quantitatively controlled tool paths. The proposed method combines specimen-specific surface path generation, curvature-adaptive iso-scallop spacing, surface-normal tool alignment, and collision-aware path trimming. The spacing between adjacent tool paths is determined directly from the local surface curvature and a prescribed scallop height, eliminating the need for iterative tuning of a specimen-specific global spacing. In NVIDIA Isaac Sim, the proposed method increased the peeling ratio from 92.97% with the equal-spacing baseline to 98.66%, while edible flesh loss increased from 1.96% to 3.04% and processing time increased from 55.38,s to 95.79,s.
|
| |
| 16:10-17:10, Paper FrPo7P.10 | |
| Integrated Mobile Snow-Removal Robot Platform Using FAST-LIO2-Based Localization and Web-Based Control |
|
| Park, BoJeong | KIRO |
| Park, Chanill | Korea Institute of Robotics & Technology Convergence (KIRO) |
| Jung, Eui-Jung | Korea Institute of Robot and Convergence |
Keywords: Robotic Applications, Navigation, Guidance and Control, Autonomous Vehicle Systems
Abstract: Current snow-removal operations are performed mainly on roads, whereas narrow sidewalks and alleys are difficult to access with large snow-removal equipment and remain highly dependent on human labor. In addition, snow-removal operations during the night and early-morning hours increase operator fatigue and safety risks. To address these issues, this paper presents a mobile snow-removal robot system using FAST-LIO2-based localization and web-based control. The proposed system integrates LiDAR-IMU-based mapping, PCD-map-based localization, traversability-costmap-based Nav2 navigation, web-based waypoint assignment, and brush and brine equipment control into a ROS2-based platform. In particular, a two-stage localization structure is applied by combining global PCD-map-based initial localization with FAST-LIO2-based local tracking, thereby considering both initial-pose uncertainty and computational load during operation. Experimental results show that the proposed switching structure reduced the computational load by 58.9% compared with the case in which global PCD-map-based localization was continuously executed without switching. In addition, the integrated operation of mapping, localization, waypoint navigation, web-based operation, and equipment command execution was verified. This study focuses on validating the operational feasibility and integrated system architecture of a mobile robot for snow-removal tasks in confined outdoor environments, rather than evaluating quantitative snow-clearing performance under actual snow-covered conditions.
|
| |
| 16:10-17:10, Paper FrPo7P.11 | |
| FreeRTOS-Based Particle Filter Implementation for Real-Time Dynamic Target Tracking on Resource-Constrained MCUs |
|
| Kim, Chaeyoung | Kumoh National Institute of Technology |
| Ryu, Seungha | Kumoh National Institute of Technology |
| Lee, Heoncheol | Kumoh National Institute of Technology |
Keywords: Robotic Applications, Robot Mechanism and Control, Artificial Intelligence Systems
Abstract: This paper investigates the real-time execution characteristics of particle filtering for dynamic target tracking on a resource-constrained embedded platform and presents a FreeRTOS-based stage-level execution architecture for coordinating computationally intensive particle-filter processing with time-critical control execution. Unlike single-task RTOS execution, the proposed architecture organizes the major particle-filter computations and the subsequent estimate computation into separate binary-semaphore-synchronized tasks while preserving their sequential data dependencies. The estimate computation derives the representative target state from the particle set for use by the control process and is implemented as a separate computational step following resampling. A fixed-priority configuration is further applied to differentiate control-related processing within the stage-level architecture. The proposed approach is evaluated on an ARM Cortex-M7-based STM32H743I MCU operating at 480 MHz through execution-time profiling, motor-control jitter analysis, and deadline-feasibility experiments. Four configurations are compared: Non-RTOS sequential execution, single-task RTOS execution, equal-priority stage-level RTOS execution, and the proposed priority-based stage-level RTOS execution. The proposed configuration exhibits the lowest motor-control jitter, demonstrating improved temporal isolation between PF computation and control processing. Under a 120-ms deadline, the empirical feasibility boundary occurs between 1186 and 1187 particles; 1186 particles satisfy all 1,000 measured cycles, whereas 1187 particles produce a 20% deadline-miss rate.
|
| |
| 16:10-17:10, Paper FrPo7P.12 | |
| HD-LIO: Bidirectional Degeneracy Feedback with Hierarchical Surfel Mapping for LiDAR-Inertial Odometry |
|
| Hong, Euntae | LGE |
| Choi, Sungjin | LG Electronics Inc |
| Cho, Beom-Jin | LG Electronics Inc |
| Noh, DongKi | LG Electronics Inc |
Keywords: Robotic Applications, Navigation, Guidance and Control, Robot Vision
Abstract: LiDAR-inertial odometry (LIO) provides accurate real-time localization and mapping for autonomous navigation. Its accuracy, however, can degrade in feature-sparse scenes where scan to map residuals provide weak directional constraints. The problem is more severe with narrow-FOV LiDAR because the scan geometry limits the observable directions and allows drift to enter both the state estimate and the map. Existing tightly coupled LIO methods mainly handle this effect in the state update, while the map can still integrate biased points along weakly constrained directions. Asymmetry between the estimator and the map can sustain a drift map feedback loop that is not addressed by estimator side degeneracy handling alone. We present HD-LIO, a hierarchical surfel map LIO system with bidirectional degeneracy feedback. HD-LIO uses degeneracy cues from the information matrix to attenuate fine voxel map updates along rank deficient directions, restore directional constraints through a per query voxel cache correspondence source, and preserve that auxiliary source with an update gate under persistent degeneracy. A surfel eigenspectrum reliability weight addresses local measurement uncertainty in each scan to map residual, complementing a directional linearized error model of the joint state map dynamics underlying these map side operations. We evaluate HD-LIO on public dataset splits covering Livox Avia, Livox Mid-360, and Ouster OS1-16 data. HD-LIO achieves the lowest mean trajectory error in every sensor specific sequence group against the baseline methods. On the narrow-FOV indoor Avia sequences, HD-LIO remains bounded over complete runs while the baselines diverge.
|
| |
| 16:10-17:10, Paper FrPo7P.13 | |
| DL-GPOM: Dual-Layer Gaussian Process Occupancy Mapping for Mixed-Evidence Preservation |
|
| Seo, Jaemin | KAIST |
| Oh, Hyondong | KAIST |
Keywords: Robotic Applications, Autonomous Vehicle Systems, Navigation, Guidance and Control
Abstract: An occupancy map should tell a motion planner where sensor evidence supports free space and where it supports obstacles. Standard occupancy representations combine free and occupied observations into a single posterior at each location. When free and occupied observations overlap and the free ones are more numerous, the posterior reports the space as free and the occupied observations are ignored---a failure that risks collision. Such overlaps become more frequent in open-world environments, where the robot encounters geometries and surfaces outside its sensor's calibration. We address this loss at the representation level. We model free and occupied observations with two independent Gaussian process (GP) layers---one trained only on free observations (mathbf{V}_f), the other only on occupied observations (mathbf{V}_o)---so that per-class evidence remains separate. From (mathbf{V}_f, mathbf{V}_o) we construct an occupancy representation in which minority occupied evidence is no longer absorbed by the free majority. We call this framework dual-layer GP occupancy mapping (DL-GPOM). To keep inference tractable at large scales, DL-GPOM combines Hilbert space GP approximation with patchwise GPU batching. On a public 2D LiDAR mapping benchmark, DL-GPOM produces the most accurate occupancy classification (AUC) and the most reliable occupancy probabilities (Brier score) among compared methods while running in real time on an embedded GPU.
|
| |
| 16:10-17:10, Paper FrPo7P.14 | |
| Multi-Robot Systems As Pedagogical Tools for Active Learning in Manufacturing Process Pedagogy |
|
| Farooq, Muhammad Umar | KAIST |
| Jang, Wonseok | KAIST |
| Suh, Mingyu | KAIST |
| Ban, Sahngjin | Korea Advanced Institute of Science and Technology |
| Jang, Young Jae | Korea Advanced Institute of Science and Technology |
Keywords: Robotic Applications, Autonomous Vehicle Systems
Abstract: Engineering pedagogy is increasingly adopting active learning methodologies to improve learning outcomes in modern classrooms. In some implementations, interdisciplinary modules are integrated to enhance cross-disciplinary knowledge while reinforcing students' understanding of the core course concepts. This work reports a case study of an active learning-based project that employs a multi-robot system in a manufacturing process course. The project was designed using a problem-based learning methodology, and a pseudo-shop-floor production line was developed on a testbed, with material handling coordinated by a multi-robot system. Enrolled students were divided into teams, where they operated the robots, controlled production, and calculated key performance indicators (KPIs) using the manufacturing process knowledge acquired during the course. Survey results indicated that students highly rated the multi-robot system-based project in terms of improving course learning and enabling the practical application of theoretical concepts. Instructor observations during interviews further suggested increased student motivation and improved conceptual understanding. Overall, the case study demonstrates that robots can be effectively integrated as pedagogical tools to bridge the gap between theory and practice in engineering pedagogy, potentially improving learning outcomes and student engagement.
|
| |
| 16:10-17:10, Paper FrPo7P.15 | |
| RUMBLE: Reinforcement Learning-Based Risk-Aware Unified Motor-Load Balancing and Locomotion Adaptation for Thermal Endurance in Quadruped Robots |
|
| Kim, Jangho | Daegu Gyeongbuk Institute of Science and Technology |
| Hong, Jinsong | DGIST |
| Oh, Sehoon | DGIST |
Keywords: Robotic Applications, Artificial Intelligence Systems, Industrial Applications of Control
Abstract: Quadruped robots are increasingly being considered for industrial tasks such as inspection, logistics, and payload transportation. In such tasks, heavy payloads, long operating times, and irregular terrain jointly impose repeated high-torque demands on electrically actuated motors, which can raise motor temperatures close to protection limits. Sudden activation of actuator protection caused by overheating can interrupt the task and may also create safety risks in industrial environments. This paper proposes a temperature-conditioned locomotion modulation framework that adjusts high-level gait and posture commands according to actuator temperature. Instead of uniformly scaling all motor commands or stopping the robot, the proposed Stage 2 module modulates variables such as body pitch, duty factor, and gait frequency to reduce current and torque usage of thermally critical actuators. The method is evaluated in simulation under global high-temperature, single-joint hot, and representative single-leg hot conditions. The results show that mean current and torque RMS decrease under high-temperature states, and that selected hot joints and a representative hot leg use less actuator effort than their all-normal references. These results suggest that the proposed framework has potential for extension to high-load legged-robot tasks such as long-duration payload transportation, industrial inspection, and disaster response.
|
| |
| 16:10-17:10, Paper FrPo7P.16 | |
| Autonomous Soil Monitoring with Vision-Language Guided Robotic Sensor Insertion |
|
| Shin, Chanhee | Chungbuk National University |
Keywords: Robotic Applications, Robot Mechanism and Control, Artificial Intelligence Systems
Abstract: This paper presents an autonomous in-situ soil moisture monitoring system that employs a 6-DoF manipulator mounted on a mobile robot platform. The system uses a vision-language model (VLM) for zero-shot wilted-plant detection and a three-stage vertical approach path that guides a soil moisture sensor into the target pot through sequential approach, insertion, and extraction stages. Conventional unconstrained inverse kinematics (IK)-based motion results in oblique entry trajectories, yielding a 24% collision rate and only 72% task success. By decoupling horizontal positioning from vertical insertion, the proposed three-stage path reduces the collision rate to 4% (an 83.3% reduction) and raises the success rate to 96%. Across 50 repeated trials, the system achieves a position repeatability of 2.4 mm (SD 0.9 mm) and a tilt error within 2.3°. The end-to-end pipeline completes each pot in 65.0 ± 9.2 s, reducing the total number of sensors required by 80%.
|
| |
| 16:10-17:10, Paper FrPo7P.17 | |
| Lightweight Semantic Map Merging Based on Overlapped Semantic Object Detection and Matching for Small Swarm Robot Systems |
|
| Kim, Geunhee | Kumoh National Institute of Technology |
| Lee, Heoncheol | Kumoh National Institute of Technology |
Keywords: Robotic Applications, Sensors and Signal Processing, Robot Vision
Abstract: As the utilization of multi-robot systems in indoor environments continues to expand, the importance of semantic map merging which integrates independent maps generated by individual robots into a comprehensive global map has increasingly gained prominence. However, map merging on small-scale robot platforms faces inherent challenges, including noise from object detection, uncertainties in communication data, and the limitations of constrained computational resources, making it difficult to simultaneously achieve both real-time performance and robustness. To address these limitations, this paper proposes a lightweight semantic map merging framework designed to achieve both computational efficiency and high registration accuracy. Experimental results using small-scale robots equipped with the Jetson Nano platform demonstrate that the proposed framework maintains high registration success rates and real-time processing efficiency, even within significantly constrained computational environments.
|
| |
| 16:10-17:10, Paper FrPo7P.18 | |
| A Unified Framework for Collision-Free Path Planning and Contact-Compliant Safety Control in Surgical Assistant Robots |
|
| Lee, Sanghoon | KAIST |
| Kim, Sungmin | Korea Advanced Institute of Science and Technology (KAIST) |
| Kim, JinYeol | Korea Advanced Institute of Science and Technology |
| Han, Seo Wook | Korean Advanced Institute of Science and Technology |
| Kim, Min Jun | KAIST |
Keywords: Robotic Applications, Biomedical Instruments and Systems, Robot Vision
Abstract: Surgical assistant robots can support surgeons in repetitive and physically constrained tasks, but their deployment in shared surgical workspaces requires safe operation around dynamic obstacles, delicate anatomical structures, and contact interactions. In such settings, robots must generate feasible motions toward surgical objectives while maintaining safe physical interaction under unexpected contact or human intervention. This paper presents a safety-aware planning and control framework that addresses both geometric and contact safety. The framework couples online obstacle perception and collision-aware motion planning with force-bounding low-level control. The planner adapts reference motion under obstacle, task, and robot constraints, while the controller regulates task-space interaction forces under disturbances and model uncertainty. Validation in a realistic surgical-assistance scenario demonstrates that the proposed framework can coordinate online motion adaptation and force-bounding execution in a contact-rich shared workspace.
|
| |
| 16:10-17:10, Paper FrPo7P.19 | |
| Low-Order Fluid Dynamics Estimation Encoder for Robotic Fluid Transport |
|
| Lee, Jun-Woo | Kookmin University |
| Kim, TaeHyun | Kookmin University |
| Kim, Sebeom | Kookmin University |
| Seo, Hyung-Tae | Kookmin University |
Keywords: Robotic Applications, Robot Vision, Artificial Intelligence Systems
Abstract: Robotic liquid transport requires the robot to account for fluid surface responses induced by container motion. Instead of reconstructing the full fluid state or directly estimating raw liquid properties, this paper focuses on estimating low-order response parameters relevant to downstream transport control. We propose a GRU-based Fluid Encoder that maps time-series fluid surface features to the damping ratio zeta and natural frequency omega_n. To introduce actuator-tracking dynamics into the excitation process, a MuJoCo-based virtual container system is used to convert commanded chirp motion into simulated container acceleration, which then drives the low-order sloshing response model. A chirp excitation spanning the target frequency range is adopted to improve parameter identifiability. The proposed encoder is further compared against two analytical frequency-domain baselines, referred to as Blind FRF and Privileged FRF. A control-relevance validation further indicates that the estimated parameters preserve information relevant to downstream motion adaptation.
|
| |
| 16:10-17:10, Paper FrPo7P.20 | |
| VR-Based Teleoperation and Data Collection Pipeline for VLA Training on a Dual-Arm Manipulator |
|
| Hong, SuHyun | University of Seoul |
| Hwang, Myun Joong | University of Seoul |
Keywords: Robotic Applications, Artificial Intelligence Systems
Abstract: Vision-Language-Action (VLA) models have recently emerged as a promising approach to learning robot manipulation policies from visual observations and language instructions. However, deploying a VLA model on a specific real-world robot requires a demonstration dataset compatible with the robot's kinematic structure, sensor configuration, and action representation. This paper presents a VR-controller-based teleoperation and data collection pipeline for VLA training on a dual-arm platform. The system supports independent teleoperation of the left and right PiPER manipulators and records multi-view RGB images, dual-arm robot states, actions, and language instructions in the LeRobotDataset v3.0 format. The proposed pipeline was evaluated through SmolVLA fine-tuning and real-robot execution on a controlled bottle pick-and-place task.
|
| |
| 16:10-17:10, Paper FrPo7P.21 | |
| When UAVs Meet IoT: Digital Twin-Based Training System |
|
| Jung, Sungwook | KETI (Korea Electronics Technology Institute) |
| Jung, Won-Seok | Korea Electronics Technology Institute |
| Choi, Sung-Chan | Korea Electronics Technology Institute |
| Sung, Nak-Myoung | Korea Electronics Technology Institute |
| Ahn, Il-Yeop | Korea Electronics Technology Institute |
Keywords: Robotic Applications, Autonomous Vehicle Systems, Control Devices and Instruments
Abstract: This paper presents a digital twin-based training system that integrates physical and software-in-the-loop UAVs with a ground control system and a three-dimensional simulation environment. A digital twin relay engine built on the oneM2M-compliant Mobius platform provides a common middleware layer for UAV telemetry exchange, command delivery, and physical--virtual state synchronization. Platform-specific flight-control interfaces are isolated within the relay layer, allowing the ground control system and simulator to process physical and virtual UAV data through a unified communication structure. The implemented system was evaluated using both SITL-based and physical UAV platforms. The experimental demonstration confirmed end-to-end telemetry delivery, ground control system integration, and the reflection of physical UAV states in the virtual environment.
|
| |
| 16:10-17:10, Paper FrPo7P.22 | |
| Design Progress of an Articulated Robotic Arm for Low-Payload Maintenance Tasks in KSTAR |
|
| Lee, Dohee | Korea Institute of Fusion Energy |
| Kim, Hong-Tack | Korea Institute of Fusion Energy |
| Park, Young Min | Korea Institute of Fusion Energy |
| Hong, Kwon Hee | Korea Institute of Fusion Energy |
| Her, Namil | KFE |
| Choi, Jungsup | SEOULTECH UNIVERSITY |
| Kim, Jinhyun | Seoul National University of Science and Technology |
| Kim, Beom Seok | Seoul National University of Science and Technology |
| Moon, Jeong Whan | KNR System |
| Ryew, Sung Moo | KnR Systems Inc |
Keywords: Robotic Applications
Abstract: Maintenance of fusion experimental devices like KSTAR is challenging due to harsh conditions such as high vacuum and temperatures. Typically, long downtime is required to cool and reduce radiation levels before possible human access. To address this, we designed an articulated robotic arm to perform low-payload maintenance tasks inside the KSTAR device. This paper presents the design progress of an articulated robotic arm, including system configuration, hardware design, and structural analysis, as part of a stage-wise approach to developing the arm. The arm is connected to the shuttle device and stored in a cask, totaling 11 m and 13 DoF. We conduct a finite element method (FEM) analysis to ensure design reliability and safety under target load conditions. We use the proposed robot system to perform essential maintenance operations for KSTAR, including visual inspection, debris removal, and other simple tasks. Our robot arm system has the potential to reduce maintenance preparation time, minimize radiation exposure to personnel, and contribute to improving the fusion experiment’s operational uptime and efficiency
|
| |
| 16:10-17:10, Paper FrPo7P.23 | |
| A Hybrid Robotic Transfer Platform for Manual-Inspection Baggage Logistics |
|
| Jung, Jooik | Incheon International Airport Corporation |
| Cho, Nam-Hyun | Incheon International Airport Corporation |
| Weon, Ihnsik | Korea Institute of Industrial Technology |
Keywords: Robotic Applications, Industrial Applications of Control, Process Control Systems
Abstract: This paper presents a hybrid robotic transfer platform for manual-inspection baggage logistics. In many airport baggage handling systems, baggage selected for manual inspection is still transported by workers between the main baggage handling flow and physically separated inspection areas. To address such recurring operational conditions, this paper reframes the problem as a transferable platform design challenge. It proposes a modular architecture that integrates legacy baggage and process interfaces, centralized orchestration, interoperable autonomous mobile robots (AMRs), conveyor and station handoff modules, vertical transfer interfacing, and safety and traceability functions. The hybrid concept separates repetitive handoff functions from mobile transport functions, thereby reducing technical risk while preserving flexibility across diverse site conditions. The manuscript summarizes the platform concept, layered system architecture, implementation rationale, and verification scenarios for phased deployment in airport environments.
|
| |
| 16:10-17:10, Paper FrPo7P.24 | |
| Human-Guided Reinforcement Learning-Based Manipulation for Multi-Purpose Agricultural Robots |
|
| Kim, Changjo | Chonnam National University |
| Kim, Gangmin | Chonnam National University |
| Choi, Jangju | Chonnam National University |
| Kim, Sumin | Chonnam National University |
| Park, Yonghyun | Gwangju Institute of Science and Technology (GIST) |
| Son, Hyoung Il | Chonnam National University |
Keywords: Robotic Applications, Artificial Intelligence Systems, Robot Mechanism and Control
Abstract: This study proposes a human-guided, reinforcement learning-based manipulation method for dual-arm agricultural robots performing various farming tasks in complex environments. By combining imitation learning and reinforcement learning, the proposed method enables the system to learn efficient and stable manipulation strategies while reducing reliance on manually designed control rules. First, the system learns human behaviors related to farming tasks through behavior cloning (BC), and then refines the learned policy using the proximal policy optimization (PPO) algorithm to enhance robustness and adaptability across diverse agricultural environments. This enables the system to mimic complex dual-arm cooperative behaviors, such as performing a primary task with one arm while removing obstacles with the other, allowing efficient operation even in complex and irregular environments. Through this study, agricultural robots are expected to perform tasks in a human-like manner across complex and diverse agricultural environments.
|
| |
| 16:10-17:10, Paper FrPo7P.25 | |
| Tactile-Based Reinforcement Learning for Visibility Enhancing Branch Pushing in Agricultural Robot |
|
| Kim, Gangmin | Chonnam National University |
| Jo, Yuseung | Chonnam National University |
| Kim, Changjo | Chonnam National University |
| Park, Jiwoon | Chonnam University |
| Ahn, Yuhyun | Chonnam National University |
| Park, Yonghyun | Gwangju Institute of Science and Technology (GIST) |
| Son, Hyoung Il | Chonnam National University |
Keywords: Robotic Applications, Artificial Intelligence Systems, Robot Mechanism and Control
Abstract: Agricultural robots in dense crop environments often suffer from visibility limitations caused by occluding leaves and branches. This study presents a tactile-based reinforcement learning for visibility enhancing branch pushing in agricultural robots. A soft tactile sensor module attached to a three-finger gripper estimates contact cues, including contact center, contact-direction tendency, and moment tendency during interaction with a deformable branch. These cues are used in a reinforcement learning policy trained in a physics-based branch simulation environment. In a 200 mm pushing task, the tactile-based policy achieved higher final visibility, required fewer steps to success, and reduced the end-effector tracking RMSE from 311.82 mm to 43.70 mm compared with the non-tactile condition. The results indicate that tactile-based reinforcement learning improves pushing accuracy during occlusion-clearing manipulation.
|
| |
| 16:10-17:10, Paper FrPo7P.26 | |
| Physics-Based Data Augmentation Using the Integral Imaging Method |
|
| You, Seungjin | Kyushu Institute of Technology |
| Cho, Myungjin | Hankyong National University |
| Lee, Min-Chul | Kyushu Institute of Technology |
Keywords: Sensors and Signal Processing
Abstract: As the demand for large-scale training data grows to enhance computer vision models, conventional methods such as manipulation-based augmentation or generative (artificial intelligence) AI synthesis face limitations, including corrupted physical consistency and the generation of optical hallucinations. To address these challenges, this paper proposes a physics-based data augmentation technique utilizing integral imaging technology. The proposed method interprets multi-view elemental images as spatial-angular sequential data and synthesizes high-resolution images from virtual viewpoints using the Virtual Focal Point with Beams (VFPB) model. Experimental results demonstrate that the augmented images achieve high quality, with an average peak signal-to-noise ratio (PSNR) of 38 and structural similarity index measure (SSIM) of 0.98 compared to the original reference. Furthermore, by precisely adjusting optical parameters such as the virtual focal distance and beam size, the generated datasets can flexibly inherit and expand the visual characteristics of the original multi-view images without requiring additional real-world data collection or complex computational rendering. This method presents a new physics-based learning paradigm that provides high-quality customized data optimized for domain-specific AI training, which is often difficult to address with general-purpose datasets.
|
| |
| 16:10-17:10, Paper FrPo7P.27 | |
| Photon Counting Stabilized Retinex Fusion for Low-Light Enhancement |
|
| Kwon, Sanghyeon | Kyushu Institute of Technology |
| Yeo, Gilsu | Kyushu Institute of Technology |
| Cho, Myungjin | Hankyong National University |
| Lee, Min-Chul | Kyushu Institute of Technology |
Keywords: Sensors and Signal Processing
Abstract: This paper proposes a photon counting stabilized Retinex method for low-light image enhancement in extremely dark environments. The proposed method first interprets the luminance component of the input image as a photon count representation and applies a Poisson-aware stabilization process to reduce shot noise caused by insufficient photon arrivals. The stabilized luminance is then processed using multi-scale Retinex (MSR) to generate a Retinex prior that compensates for illumination variations and enhances structural information. The photon stabilized luminance and Retinex prior are fused through a maximum a posteriori (MAP)-based weighted fusion scheme to restore brightness and fine details in dark regions. To further suppress noise amplified during enhancement, a mask-based post-processing stage for preserving the brightness is applied in the YCrCb color space, selectively reducing luminance and chrominance noise while preserving edges and structural details. Experiments using the See-in-the-Dark (SID) dataset show that the proposed method improves dark-region visibility and reduced color and background noise compared with conventional Retinex-based methods, and the peak signal-to-noise ratio (PSNR) and structural similarity index measure (SSIM) results prove the effectiveness of the proposed method.
|
| |
| 16:10-17:10, Paper FrPo7P.28 | |
| Fieldscale 2.0: Spatiotemporal Tail-Adaptive Rescaling for Thermal Infrared Images |
|
| Kim, Eunseo | Seoul National University |
| Gil, Hyeonjae | SNU |
| Rhee, Tai Hyoung | Seoul National University |
| Kim, Ayoung | Seoul National University |
| Lee, Dongjae | Seoul National University |
Keywords: Sensors and Signal Processing, Robot Vision
Abstract: Thermal infrared (TIR) imaging plays a vital role acrossdiverse illumination conditions, yet its native 14-or 16-bit format requires compression to standard 8-bit representations for visualization and downstream processing. While conventional linear and nonlinear tone-mapping approaches often cause severe darkening or detail loss, recent field-based rescaling methods offer better spatial adaptivity. However, these frameworks still suffer from residual artifacts, saturation, and temporal flickering. To resolve these limitations, we present Fieldscale 2.0, a spatiotemporal adaptive rescaling framework. Specifically, it integrates a three-knot field model to regulate tail distributions, a distribution-aware suppression module to prevent information loss from over-truncation, and an exponential moving average (EMA)-based temporal smoothing scheme for temporal stability. To validate its effectiveness, we conduct experiments across multiple thermal datasets and show that Fieldscale 2.0 achieves robust image and rescaling quality, improves temporal consistency in video sequences, and delivers better performance on a downstream task.
|
| |
| 16:10-17:10, Paper FrPo7P.29 | |
| Virtual Lens Array Synthesis from Multi-View Images Via IRIM |
|
| Lee, Hyeongju | Kyushu Institute of Technology |
| Cho, Myungjin | Hankyong National University |
| Lee, Min-Chul | Kyushu Institute of Technology |
Keywords: Sensors and Signal Processing
Abstract: Integral imaging acquires the 4D light field through a physical lens array, which is difficult to fabricate and subject to optical errors such as vignetting and inter-lens crosstalk. In contrast, lens-array displays for integral imaging are commercially available with diverse specifications. This work introduces a deterministic framework that bridges this asymmetry: iRIM, the inverse of Ray-based Inverse Mapping (RIM), synthesizes the elemental image (EI) that any virtual lens array---across diverse (N_L, N_P, l, f) specifications---would have captured from a single multi-view input. The synthesis follows from the bijection established by the forward RIM and is therefore free of quantization error and learned components. We parameterize the virtual lens-array configuration using the physical display geometry, aligning the synthesized EI with the display side for direct 3D visualization. On a Blender-simulated multi-view scene, iRIM produces uniformly clean sub-aperture images (SAIs) across the entire 6X6 angular grid, while a simulated Lytro-style lens-array camera baseline on the same input retains as little as 6.7% effective coverage at the corner SAIs, quantifying the practical gap closed by deterministic synthesis.
|
| |
| 16:10-17:10, Paper FrPo7P.30 | |
| Algebraic Elliptical Extent Estimation for Robust Extended Object Tracking from Partial Contour Observations |
|
| Lee, In Ho | Korea Institute of Industrial Techology |
Keywords: Sensors and Signal Processing, Robotic Applications, Process Control Systems
Abstract: This paper presents a robust extended object tracking (EOT) algorithm that estimates the full elliptical extent of a target from partial contour observations. In practical sensing geometries only the sensor-facing side of the object is observable, so the conventional assumption of measurements distributed over the entire surface no longer holds. The proposed framework models the sensor--object geometry to identify the visible contour, applies an affine-invariant whitening transformation to the measurements, and solves the ellipse-fitting problem in closed form via singular value decomposition, with an explicit ellipse-validity check and an ellipse-specific fallback fit; the recovered parameters are assimilated as linear pseudo-measurements by a standard Kalman filter. On a synthetic maneuvering trajectory the method attains a Gaussian Wasserstein (GW) error of 9.15 m at 0.94 ms per scan, against 64.84--70.62 m for RMM, MEM-EKF*, and PAKF. On a real LiDAR car-tracking sequence of 49 scans with 58--261 points per scan, in which only the right and rear facets are visible (visibility ratio 47--48%), the method attains mean semi-axis errors of 0.16 m and 0.08 m and an orientation error of 2.7 deg against the vehicle dimensions, outperforming the baselines. A sensitivity study delineates the applicable scope: single-scan reconstruction remains reliable for visibility ratios down to approximately 0.3.
|
| |
| 16:10-17:10, Paper FrPo7P.31 | |
| Roll-Rotated Imaging-Sonar Scanning for Elevation-Ambiguity-Free Mapping of Steep Underwater Terrain for Underwater Inspection |
|
| Ku, Bonchul | Pohang University of Science and Technology |
| Kim, Jason | HEROLab (in Univ. POSTECH) |
| Kim, Seungmin | Pohang University of Science and Technology |
| Song, Young-woon | Pohang University of Science and Technology (POSTECH) |
| Yu, Son-Cheol | Pohang University of Science and Technology (POSTECH) |
Keywords: Sensors and Signal Processing, Robotic Applications
Abstract: Underwater 3D mapping with a forward-looking imaging sonar suffers from elevation ambiguity. Each acoustic return gives range and bearing but not elevation angle, so on steeply sloped terrain returns are placed at the wrong height and form a false slope. Prior methods remove this ambiguity by adding a second sonar, which requires an extra sensor and data association between the two devices. This paper presents an active mapping method that uses a single imaging sonar and rotates it about its roll axis on demand. While an occupancy map is built online, steep regions likely to be distorted by the ambiguity are detected and re-scanned with the sonar rolled ninety degrees about its look axis, so that height is resolved by the well-measured bearing angle. The rolled observations then carve the false slope from the map through a probabilistic negative-only update. In a Stonefish simulation, the method maps steep terrain more accurately than fixed tilts of sixty and ninety degrees, lowering the RMSE against the ground truth by about 14% and the number of cells with height error above 0.5 m by about 57% over the better baseline.
|
| |
| 16:10-17:10, Paper FrPo7P.32 | |
| What Can Lines Tell Us? Parking-Line-Derived Orientation Constraints for Parking-Lot Mapping |
|
| Sangmi, Hyeon | Jeonbuk National University |
| Sunghwan, Jeong | Korea Electronics Technology Institute |
| Choi, Kyoungho | Ekonexon |
| Jo, HyungGi | Jeonbuk National University |
Keywords: Sensors and Signal Processing, Autonomous Vehicle Systems, Navigation, Guidance and Control
Abstract: LiDAR-inertial odometry (LIO) in ground-dominant parking lots can suffer from biased road-surface correspondences caused by parked vehicles, curbs, and partially occluded ground returns. We exploit intensity-selected parking-line points to estimate a local road-surface plane and incorporate its normal into an iterated error-state Kalman filter (IESKF) as an orientation-only pseudo-measurement. The same plane also defines a one-sided map-insertion gate that suppresses below-plane outliers. On a real parking-lot sequence collected by an EV charging robot, the proposed method achieved the lowest Mean Map Entropy (MME) and the smallest final-revisit vertical displacement magnitude among the evaluated methods. The complete pipeline required 51.11 ms per scan on an NVIDIA Jetson AGX Orin.
|
| |
| 16:10-17:10, Paper FrPo7P.33 | |
| Communication-Blackout-Aware MPC for Mobile Relay Robot Deployment in Indoor Disaster Environments |
|
| Kim, Dong Ju | Pukyong National University |
| Kim, Sung Jae | Pukyoung National University |
| Suh, Jinho | Pukyong National University |
Keywords: Sensors and Signal Processing, Robotic Applications
Abstract: Reliable wireless communication is essential for teleoperated mobile robots operating in indoor disaster environments. However, walls, shelves, corridors, and non-line-of-sight propagation can cause rapid degradation of the received signal strength indicator (RSSI), resulting in communication blackout and unstable teleoperation. This paper proposes a learned communication-blackout-aware model predictive control (MPC) framework for mobile relay robot deployment. A lightweight blackout-risk estimator is trained using simulated RSSI transition samples and predicts the probability of future blackout from safe RSSI, RSSI rate, inter-node distance, obstacle blocking count, and RSSI deficit. The estimated blackout risk is incorporated into the MPC cost function together with RSSI violation and safety-guard terms, and an RSSI-safeguarded action selection limits excessive instantaneous RSSI degradation caused by the learned risk term. Simulation results in a warehouse-type indoor environment show that the proposed method achieves the highest average worst-link RSSI, the lowest RSSI violation count, and the lowest mean blackout risk compared with distance-based deployment, connectivity-aware reactive control, and RSSI-constrained MPC.
|
| |
| 16:10-17:10, Paper FrPo7P.34 | |
| State of Charge Estimation under Temperature Mismatch Via Voltage-Domain Adaptation and SOC-Domain Residual Learning |
|
| Cheon, Kibum | Chungbuk National University |
| Haejun, Kim | Chungbuk National University |
| Shin, Jongho | Chungbuk National University |
Keywords: Sensors and Signal Processing, Control Theory and Applications, Artificial Intelligence Systems
Abstract: State of charge (SOC) estimation based on equivalent circuit models (ECMs) and Kalman filters is widely used in battery management systems. However, its accuracy degrades when the operating temperature differs from that of the identified model. Because constructing a complete temperature-specific library of open-circuit voltage (OCV)–SOC tables and ECM parameters is often impractical, a nominal model identified at a reference temperature must be reused at other operating temperatures. This paper presents an online SOC estimation framework that addresses such temperature mismatch through sequential compensation. Starting from a nominal extended Kalman filter (EKF), the ohmic resistance is first updated online using recursive least squares (RLS) to reduce voltage-domain mismatch. The remaining SOC-domain residual is then corrected by a lightweight neural network (NN), and a low-pass filter stabilizes the correction. On the CALCE INR 18650-20R dataset, applying a 25°C nominal model to 45°C dynamic discharge profiles, the proposed estimator reduces the SOC RMSE from 0.2911% for the 25°C fixed EKF to 0.0740%. The estimator also remains suitable for real-time operation.
|
| |
| 16:10-17:10, Paper FrPo7P.35 | |
| Weakly Supervised rPPG Learning Using ECG-To-PPG Signal Translation Diffusion Model |
|
| Oh, Seungmin | Jeonbuk National University |
| Choi, Jiho | Jeonbuk National University |
| Lee, Sang Jun | Jeonbuk National University |
Keywords: Sensors and Signal Processing, Biomedical Instruments and Systems, Artificial Intelligence Systems
Abstract: Remote photoplethysmography (rPPG) estimates physiological signals from facial videos without contact sensors. Supervised rPPG methods require reliable photoplethysmography (PPG) labels, but PPG signals measured in driving environments can be distorted by motion artifacts, illumination changes, vehicle vibration, and contact pressure variation. These noisy labels can degrade rPPG training by providing inaccurate pulse patterns. To address this problem, this paper proposes a weakly supervised rPPG learning framework using ECG-to-PPG signal translation. The proposed framework generates pseudo PPG signals from ECG signals and uses them as supervision signals for rPPG model training. For pseudo PPG generation, this paper uses Physiological Frequency-Consistency Region-Disentangled Diffusion Model (PFC-RDDM), which incorporates PPG peak- and trough-based ROIs, ECG R-peak condition information, and physiological loss functions. Experimental results show that PFC-RDDM improves ECG-to-PPG generation performance and that generated pseudo PPG labels can be used for rPPG training on ECG-only driving data.
|
| |
| 16:10-17:10, Paper FrPo7P.36 | |
| Information-Aware Source Selection for Open-Loop Drone Response Prediction |
|
| Park, Hyeongjun | Hanyang University |
| Kang, Chang Mook | Hanyang University |
Keywords: Sensors and Signal Processing, Artificial Intelligence Systems, Robotic Applications
Abstract: This extended abstract examines source-selected response prediction for an unmanned aerial vehicle (UAV) simulated in PX4 software-in-the-loop (SITL). Flight logs are reorganized into transition samples containing velocity and angular-rate states, offboard command setpoints, tracking-error terms, and finite-difference derivative labels. A Ridge predictor is trained from source excitation logs after candidate source windows are ranked by a composite score reflecting trajectory diversity, tracking mismatch, response magnitude, and command variation. Prediction quality is assessed by recursively propagating states in open loop, rather than by relying only on one-step derivative fitting. The experiments show that source-selected Ridge is most effective on aggressive velocity-reversal maneuvers, while persistence remains competitive for smoother or mixed trajectories. Controlled checks with multilayer perceptron (MLP) base and residual models further indicate that improved local fitting does not automatically produce stable multi-step prediction. These results emphasize rollout-oriented validation and maneuver-dependent source selection for compact UAV response models.
|
| |
| 16:10-17:10, Paper FrPo7P.37 | |
| LiDAR-Guided Geometric Filtering for Noise Removal in W-Band Radar Odometry |
|
| Lee, Dongje | SEADRONIX |
| Kim, Hanguen | Seadronix Corp |
| Jang, Hyesu | Seadronix |
Keywords: Sensors and Signal Processing, Navigation, Guidance and Control, Robotic Applications
Abstract: W-band radar is increasingly adopted for perception in robotics, with its robustness under adverse weather and its extended range relative to camera and LiDAR. Prior radar approaches rely on intensity-based detection, assuming noise returns are weak. In practice, radar cannot separate spurious signals since noise frequently matches the intensity of real structures. Thus, we propose a heterogeneous sensor fusion method that uses geometric line features extracted from LiDAR. Line features serve as a spatial constraint, retaining only radar returns originating from physical structures. Across maritime and urban ground datasets, we filtered out the noise and verified that the discarded noise is nearly indistinguishable from the retained returns in intensity. Supplying the refined returns to state-of-the-art radar odometry algorithms reduces the absolute pose error (APE) of every baseline.
|
| |
| 16:10-17:10, Paper FrPo7P.38 | |
| Detecting Fabric Deformation Using Piezoelectric Soft Sensors |
|
| Lee, TaeKyoung | Korea University |
| Cha, Youngsu | Korea University |
Keywords: Sensors and Signal Processing
Abstract: In this paper, we propose a sensing system for detection of fabric deformation using multiple tendon-inspired sensors. Specifically, the piezoelectric sensors in the sensing system are positioned on the corner of a fabric in a fan-shaped configuration. Additionally, sewing threads are connected to the sensors and are tied to the edges of the sensing area, transferring tensile force to the sensors. Furthermore, a depth camera is installed to calculate the ground-truth curvature of the fabric. A series of experiments is conducted with various experimental conditions, including the shape of supporting plates and pulling distances, to evaluate the relationship between the sensor output and the corresponding curvatures. From these results, the curvature of the fabric is found to be estimable using the relationship between the voltage and the curvature. Furthermore, a virtual simulation of fabric deformation is conducted to visualize the real-time change in fabric surface, demonstrating the feasibility of the sensing system. Moreover, we discuss the theoretical expectation for the tendon-inspired sensors.
|
| |
| 16:10-17:10, Paper FrPo7P.40 | |
| Accuracy Enhancement of Finite-Time Spectral Observers Via Exact Partial Discretization for Periodic Signal Estimation |
|
| Hong, Woosuk | Tokyo University of Science |
| Murakami, Madoka | Tokyo University of Science |
| Nakamura, Hisakazu | Tokyo University of Science |
|
|
| |
| 16:10-17:10, Paper FrPo7P.41 | |
| GEONJI-3D: A Multimodal Dataset for 3D Perception in Unstructured Outdoor Environments |
|
| Shin, Jaeho | Jeonbuk National University |
| Ha, Jiwon | Jeonbuk National University |
| Hwang, JunHyeon | Jeonbuk National University |
| Park, Jaebyung | Jeonbuk National University |
Keywords: Sensors and Signal Processing, Robotic Applications, Robot Vision
Abstract: Autonomous robots operating in unstructured outdoor environments require robust perception capabilities to safely navigate complex terrain with slopes, vegetation, rocks, tree branches, and dynamic obstacles. Although several datasets have been proposed for autonomous driving and off-road perception, multimodal datasets collected using real ground robot platforms in forest-like unstructured environments remain limited. In this paper, we introduce GEONJI-3D, a multimodal dataset for 3D perception and traversability estimation in unstructured outdoor environments. The dataset was collected using a tracked ground robot platform equipped with a stereo RGB-D camera, a 3D LiDAR sensor, and an IMU at the Geonji Mountain Research Forest of Jeonbuk National University. GEONJI-3D consists of two driving scenarios that include well-maintained walking trails, narrow and rough paths, revisited section, vegetation, tree branches, and dynamic obstacles. The dataset provides synchronized RGB images, depth images, LiDAR point clouds, and IMU measurements in the ROS 2 bag format. By providing real-world multimodal sensor data from challenging outdoor environments, GEONJI-3D can support research on sensor fusion, terrain perception, traversability estimation, and robust autonomous navigation for ground robots.
|
| |
| 16:10-17:10, Paper FrPo7P.42 | |
| Seat Occupancy Localization Using Differential Range-Angle Energy Distribution in mmWave Radar |
|
| Kim, Sumin | Changwon National Unversity |
| Park, Minjoo | Changwon National Univesrity |
| Ryu, Haeun | Changwon National University |
| Gim, Juhui | Changwon National University |
Keywords: Sensors and Signal Processing, Autonomous Vehicle Systems
Abstract: This paper proposes a training-free seat occupancy localization method using mmWave radar based on seat-specific energy partitioning in range-angle (RA) maps. An empty-cabin RA map is first utilized as a static reference to suppress structural reflections and extract occupant-induced energy variations. The resulting differential RA-map energy is then partitioned into seat-specific regions using range- and angle-based boundary estimation with sigmoid weighting functions to accommodate boundary uncertainty and occupant motion. Occupancy localization is subsequently performed using normalized energy scores computed within each seat region. Experimental validation in a four-seat cabin mock-up environment demonstrates that the proposed framework can reliably localize seat occupancy under various occupant configurations while maintaining lower computational complexity than conventional radar-processing approaches. The results indicate that the proposed method provides an effective and computationally efficient solution for radar-based seat occupancy localization without requiring point-cloud generation, model training, or retraining procedures.
|
| |
| 16:10-17:10, Paper FrPo7P.43 | |
| Emergency Vehicle Siren Recognition, Azimuth Estimation, and Distance Estimation Using Multichannel Acoustic Features |
|
| An, Hyogeon | Tech University of Korea |
| Park, Sangjun | Tech University of Korea |
| Park, Seongkeun | Tech University of Korea |
Keywords: Sensors and Signal Processing, Artificial Intelligence Systems, Autonomous Vehicle Systems
Abstract: This paper proposes a model that simultaneously performs event recognition, azimuth estimation, and distance estimation for emergency vehicle siren sounds using multichannel acoustic features. Conventional SELD-based models have been effectively used for sound event detection and direction-of-arrival estimation; however, distance information of the sound source should also be considered to determine the approach status of an emergency vehicle. Therefore, in this study, a Distance Head is added to a CST-Former-based SELD architecture, enabling a single model to jointly predict siren events, azimuth, and distance information. Mel-spectrograms and GCC-based multichannel acoustic features are used as input features, and the channel-, frequency-, and time-axis attention mechanisms of the CST-Former are employed to learn spatial information and time-frequency patterns from acoustic signals. For distance estimation, log-distance regression is applied to reduce the effect of scale differences in distance values, and a distance mask is used so that the loss is computed only for frames with valid distance labels. The applicability of the proposed model is verified by evaluating sound recognition performance, azimuth estimation error, and distance estimation error using multichannel siren data acquired in a real-world environment.
|
| |
| 16:10-17:10, Paper FrPo7P.44 | |
| A Multimodal Integrated Data Synchronization and Acquisition System Based on Heterogeneous Interfaces for Driver Monitoring System Evaluation |
|
| Son, Huigyeong | Korea Automotive Technology Institute |
| Lee, Hun | Korea Automotive Technology Institute |
| Oh, Young-dal | KATECH |
| Ryu, DongWoon | Korea Automotive Technology Institute |
| Park, Sunhong | Korea Automotive Technology Institute |
Keywords: Sensors and Signal Processing, Human-Robot Interaction, Biomedical Instruments and Systems
Abstract: This paper proposes a multimodal data synchronization and acquisition system for the training and evaluation of driver monitoring systems (DMS). The proposed system receives multi-channel camera video from an integrated ECU over RTSP (Real-Time Streaming Protocol) and applies a channel-wise, thread-based concurrent storage structure designed to minimize frame drops. In addition, diverse heterogeneous data such as EEG, biosignals, seat pressure, vehicle CAN, and gaze data are synchronously acquired in a temporally aligned form through trigger and communication interfaces. Through a pilot test conducted in a real-vehicle environment, the stable synchronized acquisition performance of multi-channel video and sensor data was confirmed, and it was verified that the proposed system can be used as a general-purpose data acquisition platform for various DMS studies.
|
| |
| 16:10-17:10, Paper FrPo7P.45 | |
| RGB Image Reconstruction Via Radar-IMU Fusion for Human Localization and Pose Recognition |
|
| Kim, Seungyeon | Sungshin Women's University |
| Yoo, Jaehyun | Sungshin Women's University |
Keywords: Sensors and Signal Processing, Artificial Intelligence Systems
Abstract: The importance of Human Activity Recognition (HAR) is increasing in applications such as indoor monitoring, smart homes, and home care. Frequency-Modulated ContinuousWave (FMCW) radar is widely used as an indoor sensing method because it provides range and velocity information while minimizing privacy intrusion. However, a radar-only method is limited in accurately reconstructing human localization and pose when reflected signals are weak. To address this issue, this paper proposes a FiLM-based radar-IMU fusion model that uses range-Doppler map sequences and wristworn IMU feature sequences. The proposed model generates FiLM parameters from IMU features and uses them to modulate radar feature maps. This allows motion and pose information to be incorporated into the radar representation, enabling RGB image reconstruction of human localization and pose. The experimental setup considers two dynamic activities, walking and crouching walking, over a distance range of 2–7 m, and two static poses, standing and sitting, at distances of 2 m and 7 m. Experimental results show that the proposed FiLM-based fusion method provides more stable RGB-space reconstruction for human localization and pose recognition than a radar-only method.
|
| |
| 16:10-17:10, Paper FrPo7P.46 | |
| DCL-DAE: Dilated CNN and Bi-LSTM Based IMU Denoising Autoencoder for Robust Robot Manipulator Control |
|
| Heo, Ji-yoon | Sejong University |
| Woo, Hyunsoo | Sejong University |
Keywords: Sensors and Signal Processing, Robotic Applications, Artificial Intelligence Systems
Abstract: While IMU sensors are critical for precise robot manipulator control, diverse disturbance environments induce high-frequency, nonlinear noise that compromises trajectory stability. Conventional filtering and standard deep learning models often degrade temporal resolution or incur high computational costs. To address these challenges, we propose DCL-DAE (Dilated CNN-LSTM Denoising Autoencoder), which integrates dilated convolutions with Bi-LSTM network. By adopting dilated convolutions, the encoder expands the receptive field while preserving fine-grained time-series features. Furthermore, the Bi-LSTM latent space captures bidirectional temporal contexts to isolate and filter out diverse disturbance anomalies. Evaluated on the CASPER dataset under five distinct disturbance conditions, DCL-DAE achieved an average trajectory reconstruction improvement of 68.01%, with an error suppression rate of up to 89.38% in the z-axis acceleration. Qualitative validations via PyBullet physics simulation further confirmed that the proposed architecture robustly mitigates end-effector jitter, directly translating numerical denoising gains into smooth and stable physical control.
|
| |
| 16:10-17:10, Paper FrPo7P.47 | |
| Inference-Time Temporal Max Pooling for Robust PMSM Drive-Module Fault Diagnosis under Sensor Relocation |
|
| Youn, Donggyu | UST(University of Science and Technology |
| Jeung, Deokgi | Korea Institute of Machinery and Materials |
| Sin, MinKi | Korea Institute of Machinery & Materials |
| Cho, Jang Ho | Korea Institute of Machinery & Materials |
Keywords: Sensors and Signal Processing, Robotic Applications, Artificial Intelligence Systems
Abstract: Sensor relocation alters vibration transmission paths, degrading vibration-based fault diagnosis of permanent magnet synchronous motor (PMSM) drive modules and increasing false alarms. This study evaluates inference-time temporal post-processing for robust multi-class diagnosis under sensor relocation and physically induced disturbances. A triaxial vibration dataset covering five diagnostic conditions and four sensor locations was used to train a lightweight 1D-CNN–BiLSTM model (20.20 M MACs per segment) at a single source location and evaluate it at three other locations, including two strictly unseen test locations. Raw inference, temporal majority voting (TMV), temporal moving average (TMA), and temporal max pooling (TMP) were compared. At 500 rpm with a representative window of (w=40), TMP achieved 95.28% and 99.85% accuracy with 0.00% false alarm rate at the two unseen locations, demonstrating improved cross-location reliability without target-location training data.
|
| |
| 16:10-17:10, Paper FrPo7P.48 | |
| Physics-Guided Acoustic Anomaly Detection with Environment Noise Modeling for Submarine Rotating Machinery |
|
| Kim, KwangSik | Inha University |
| Kim, Edam | Department of Naval Architecture and Ocean Engineering, Inha University |
| Kim, Yoo-Lim | Naval Ship System R&D Team, Hanwha Ocean Co. Ltd |
| Roh, Young-Ki | Naval Ship System R&D Team, Hanwha Ocean Co. Ltd |
| Lee, Won-Joon | Naval Ship System R&D Team, Hanwha Ocean Co. Ltd |
| Lee, JangHyun | Department of Naval Architecture and Ocean Engineering, Inha University |
Keywords: Sensors and Signal Processing, Artificial Intelligence Systems
Abstract: This study proposes an unsupervised anomaly detection framework for the early detection of abnormal operating conditions in submarine rotating machinery under limited-data environments. Because acoustic signals are contaminated by underwater ambient noise and onboard mechanical noise, robust performance against environmental noise is essential. A physics-based submarine noise model was developed by incorporating colored background noise, structure-borne resonance noise, band-limited auxiliary noise, tonal components, and sensor noise. Realistic noise-augmented datasets were generated under signal-to-noise ratio (SNR) conditions ranging from −6 to 0 dB. To interpret the physical characteristics of the acoustic signals, a structural dynamics model of the pump–mount–support structure was established. Modal analysis identified the natural frequencies and mode shapes, while harmonic response analysis revealed resonance-prone frequency regions and vibration amplification at different rotational speeds. For edge-computing applications, three unsupervised anomaly detection algorithms were evaluated: (i) a Gaussian Mixture Model (GMM) using statistical MFCC features, (ii) an Ensemble Autoencoder using statistical acoustic features, and (iii) a Conv1D-based Ensemble Autoencoder using log Mel-spectrogram sequences. Performance was evaluated using AUC, F1-score, and computational efficiency. The GMM achieved competitive performance with low computational cost, whereas the Conv1D-based model provided higher detection accuracy by exploiting temporal acoustic patterns at the expense of greater computational complexity. These results highlight the importance of selecting an appropriate anomaly detection algorithm by balancing detection performance and computational efficiency for resource-constrained edge-computing environments.
|
| |