| |
| FrPo6P |
Symphony D, 4F |
| Poster Session 6 |
Poster Session |
| |
| 09:30-10:30, Paper FrPo6P.1 | |
| A Feasibility Study on In-Fingertip Null-Space Control for Vial Cap Engagement and Screwing |
|
| Heo, Jeong | Hanyang University |
| Seo, TaeWon | Hanyang University |
| Hwang, Donghyun | Korea Institute of Science and Technology |
Keywords: Robot Mechanism and Control
Abstract: Robotic automation in laboratory environments can involve object handling in confined workspaces, where surrounding structures restrict large wrist rotations and arm motions. This paper presents the feasibility of the alignment and insertion preceding vial cap engagement and screwing by formulating the problem as an object-handling task and solving it with a task-priority based control framework that performs manipulation through finger-level motion of a dexterous hand, without relying on arm-centric movement. Grasp geometry that is represented by the three pairwise fingertip distances is maintained as the primary task while object translation and rotation are generated within its null-space as the secondary task. The object proxy pose is estimated from fingertip positions via forward kinematics, without any external pose sensor. The framework was validated on hardware with the arm wrist held fixed, so that all object motion was produced by finger-level manipulation. In single-axis experiments, we measured a sub-millimeter primary grasp deviation of 0.59 mm on average, and in task-level demonstration the robotic hand aligned and inserted the object while preserving the grasp geometry. These results support the feasibility of finger-level in-hand manipulation for object handling in workspace-constrained laboratory settings.
|
| |
| 09:30-10:30, Paper FrPo6P.2 | |
| Optimal Path Search for Robot-Arm-Based Magnetic Control of a Wireless Soil Sampler |
|
| Kang, SeonKyu | Chungbuk National University |
| Kim, Jayoung | Chungbuk National University |
Keywords: Robot Mechanism and Control, Robotic Applications, Control Devices and Instruments
Abstract: This study proposes a method for determining optimal positions and orientations of a robotic end-effector used in curved-path drilling of a wireless soil sampler. The system consists of an actuation magnet attached to the robotic end-effector and a permanent magnet embedded in the soil sampler. By controlling the end-effector pose, magnetic force and torque are generated to wirelessly propel and steer the sampler during drilling. Unlike conventional Edelman-auger-type soil samplers that are limited to straight-line penetration, the proposed method enables curved drilling trajectories for obstacle avoidance in underground environments. The drilling path is discretized into a sequence of steps along a refer-ence trajectory connecting the start and goal points. At each step, candidate end-effector poses are evaluated based on the magnetic force and torque required to propel and steer the sampler, which are estimated using a magnetic dipole model and an agarose (0.6 wt%) gel resistance model. Robot reachability is verified through inverse kinematics. A two-stage optimization identifies feasible poses satisfying both magnetic actuation and kinematic constraints, then selects the pose that best follows the reference path while minimizing joint motion. The proposed method enables curved-path drilling planning under coupled magnetic and robotic constraints for obstacle avoidance soil sampling.
|
| |
| 09:30-10:30, Paper FrPo6P.3 | |
| Preparation of a Digital Twin Model of Harmonic Drive Gear System Using Nonlinear Finite Element Analysis for Physical AI Applications |
|
| Park, Moon-Woo | Korea Institute for Robot Industry Advancement |
Keywords: Robot Mechanism and Control, Robotic Applications
Abstract: This paper presents a digital twin modeling approach for a harmonic drive reduction gear system based on nonlinear finite element analysis (FEA) for future Physical AI applications. Unlike conventional rigid-body transmission models, the proposed model explicitly represents the elastic deformation of the flexspline induced by the wave generator and the nonlinear tooth engagement between the flexspline and circular spline. A commercial harmonic drive reducer model was selected as the target system. A detailed finite element model was developed using HyperWorks/OptiStruct, including wave generator, flexspline, circular spline, and bearing assemblies. Nonlinear surface-to-surface contact formulations were employed to simulate deformation propagation and torque transmission mechanisms. Static structural analysis was performed under the average torque limit. The simulation results demonstrated realistic deformation patterns, tooth contact distributions, and stress concentrations. Maximum von-Mises stress of the flexspline was predicted as 549 MPa, remaining below the material yield strength of SCM440 steel. The developed model provides a physics-based foundation for digital twin construction and future reduced-order modeling techniques applicable to Physical AI systems. Future research will focus on integrating simulation data with machine learning frameworks to enable real-time state estimation, predictive maintenance, and intelligent motion planning of robotic systems utilizing harmonic drive reducers
|
| |
| 09:30-10:30, Paper FrPo6P.4 | |
| Simulation-Based Diffusion Policy Training Framework for Precision Task |
|
| Park, Sangyong | Kyungpook National University |
| Ko, Yeongmin | Kyungpook National University |
Keywords: Robot Mechanism and Control
Abstract: Precision manipulation is essential in modern robotic applications such as assembly, cutting, and inspection. Especially in cutting, simple position error can lead to project failure or damage of valuable component. Consequently, developing reliable manipulation policies for precision is required and remains a challenge. and this kind of task requires huge amount of data generated by physical robot manipulation. But this kind of data consumes time and resources. So this paper proposes a way to train precision manipulation task via simulation. The simulation can generate 100 trajectories in 511seconds and generated datasets contains three categories Video, joints, and trajectory. Each category contains as much related information as possible so it can be used for multiple model. In this paper trained model is diffuser. Result of the test shows diffuser can be trained with data generated by simulation and able to create trajectory. but It's only been tested in simulation environment, further research is required to test how it works with actual hardware. Contribution of paper is Propose a simulation for high-precision manipulation task training, Using open sourced manipulator(SO101) shows high-precision manipulation task can be performed with simple hardware and can be used in various systems, proves model can be trained with gathered data and able to generate trajectory.
|
| |
| 09:30-10:30, Paper FrPo6P.5 | |
| Real-Time Control of Tendon-Driven Continuum Robot Using MPPI and Residual Jacobian-Based Online Learning |
|
| Lee, Dongjun | Daegu Gyeongbuk Institute of Science and Technology(DGIST) |
| Kim, DongWook | Daegu Gyeongbuk Institute of Science and Technology (DGIST) |
Keywords: Robot Mechanism and Control, Control Theory and Applications, Artificial Intelligence Systems
Abstract: Tendon-driven continuum robots (TDCRs) are widely used in confined operating systems due to their thin shape, flexibility. Several modeling and control methods have been used for TDCRs, such as Cosserat rod model and model predictive control (MPC). However, Cosserat rod models and MPC have limitations in real-time control and handling the nonlinear behavior of TDCRs due to computational complexity and dynamics linearization, respectively. In this paper, to address the two problems, we employ a PCC model based residual radial basis function network (RBFN). Although the PCC model is computationally efficient due to its simple approximation, it cannot fully explain modeling errors; therefore, we augment with a residual RBFN to compensate for the errors that the PCC model cannot account for. Moreover, we employ a model predictive path integral (MPPI) controller that uses a Jacobian-based PCC–residual RBFN model as the rollout model, and the residual RBFN model is updated online using buffer data. We validate the proposed method on a handmade one-segment TDCR hardware consisting of four tendons, with vision-based tracking using two webcams. The reference trajectory tracking test shows that the proposed residual RBFN method reduces the mean tracking error, RMS tracking error, and model prediction error by 47.0%, 41.6%, and approximately 50%, respectively.
|
| |
| 09:30-10:30, Paper FrPo6P.6 | |
| RCM-Consistent Admittance Control with Inverse-Dynamics QP for Hands-On Robotic Manipulation |
|
| Jeong, Jaehun | Korea Advanced Institute of Science and Technology (KAIST) |
| Park, Seongsu | KAIST |
| Lee, Sanghoon | KAIST |
| Kim, Min Jun | KAIST |
Keywords: Robot Mechanism and Control, Robotic Applications, Human-Robot Interaction
Abstract: Robot-assisted minimally invasive surgery (RAMIS) often requires a surgical tool to pivot about a trocar while allowing direct hands-on manipulation by the surgeon. During kinesthetic guidance, however, operator-applied forces and moments act as external inputs to the robot. Their interaction with nonlinear robot dynamics and physical limits can result in remote center of motion (RCM) constraint violation or infeasible torque commands. This paper proposes an RCM-constrained control framework for hands-on manipulation that combines an RCM-consistent admittance filter with inverse-dynamics quadratic programming (QP). The admittance filter generates a joint-space reference trajectory from the measured external wrench while enforcing the RCM constraint at the position, velocity, and acceleration levels. The inverse-dynamics QP then computes torque commands that track the reference trajectory while considering the joint position, velocity, and torque limits. Baumgarte stabilization is incorporated into the acceleration-level RCM formulation to reduce errors caused by numerical drift and saturation. Simulation studies using a 6-DoF manipulator compare the proposed method with two baseline controllers under insertion-roll interaction scenarios. Overall, the results show that the proposed controller reduces both the RCM error and its time derivative while accounting for the prescribed joint torque limits in the inverse-dynamics QP.
|
| |
| 09:30-10:30, Paper FrPo6P.7 | |
| Dual-Attachment Cable Suspension System for Zero-Gravity Emulation of Space Manipulators |
|
| Shin, Seungmin | Korea Advanced Institute of Science and Technology |
| Kim, Jiwon | KAIST |
| Kim, Min Jun | KAIST |
Keywords: Robot Mechanism and Control, Robotic Applications, Control Devices and Instruments
Abstract: Space manipulators are increasingly used in various on-orbit operations. Accordingly, reliable on-ground validation of such systems has become increasingly important. However, because space manipulators are designed for a microgravity environment, their actuators are not capable of supporting gravity loads, which makes on-ground testing challenging. This paper proposes a dual-attachment cable suspension system for zero-gravity emulation of a 7-DOF space manipulator. For each attachment point, the required external force for gravity compensation is computed and mapped to feasible cable tensions under pull-only and tension-limit constraints. By applying cable-generated external forces at two different links, the proposed system achieves gravity compensation for the six dominant gravitational joint torques. The approach is evaluated in simulation with the CAESAR manipulator through joint-space trajectory tracking and a pick-and-place task. The results show that the proposed system successfully compensates for gravitational torques and produces motion close to zero-gravity behavior, while maintaining feasible cable tensions throughout the task.
|
| |
| 09:30-10:30, Paper FrPo6P.8 | |
| Learning to Utilize Passive Toe Dynamics in Bipedal Walking Via Adversarial Motion Priors |
|
| Kim, Minseok | Gwangju Institute of Science and Technology |
| Hur, Pilwon | Gwangju Institute of Science and Technology |
Keywords: Robot Mechanism and Control, Artificial Intelligence Systems, Robotic Applications
Abstract: Passive toe joints can provide compliance during stance, but their unactuated motion is difficult to induce because it depends on robot dynamics and ground contact. This paper presents a motion imitation approach based on Adversarial Motion Priors (AMP) for inducing passive toe-joint motion in bipedal walking without toe torque commands or prescribed toe trajectories. Human walking motion is retargeted to a toe-jointed bipedal robot and used for policy training. The passive toe joints are excluded from the policy action space and are not assigned reference trajectories. Instead, their motion emerges through passive spring-damper dynamics, robot motion, and ground contact. To prevent passive toe motion from being penalized as an imitation mismatch, passive toe states are masked in the discriminator observation. We compare the proposed condition with a fixed-toe baseline and a no-masking ablation. Simulation results across three independently trained random seeds show clear stance-phase passive toe dorsiflexion and toe loading. The no-masking ablation produces smaller toe dorsiflexion, suggesting that discriminator masking promotes greater passive-toe utilization. These results suggest that passive toe utilization can emerge from AMP-based motion imitation without explicit toe commands or prescribed toe trajectories.
|
| |
| 09:30-10:30, Paper FrPo6P.9 | |
| Hall-Only Wheel Friction-Change Detection: A Feasibility Study |
|
| Lee, Chang-Hwan | Gyeongkuk National Univ |
Keywords: Robot Mechanism and Control, Autonomous Vehicle Systems, Robotic Applications
Abstract: Detecting changes in wheel–terrain interaction, such as traction loss, slip, or contact loss, is essential for the safe autonomous operation of mobile robots. Conventional approaches typically rely on additional sensors such as inertial measurement units, force/torque sensors, or high-resolution encoders, which increase cost and integration complexity. This paper presents a simulation-based feasibility study of detecting wheel friction changes using only the three-phase Hall sensors already embedded in commercial in-wheel brushless DC (BLDC) motors, without any additional sensing hardware. The proposed pipeline combines a phase-locked loop (PLL) speed estimator, which reconstructs wheel angular velocity from quantized Hall edge timings, with a generalized momentum observer (GMO) that estimates the disturbance torque acting on the wheel axis. Because the observer residual converges to the negative of the disturbance torque, a change in rolling friction appears directly as a measurable shift in the residual, enabling friction-change detection without explicit acceleration measurement. The pipeline is evaluated under open-loop constant-torque operation across five operating points, from 20% to 100% of rated torque. The residual converges to the true disturbance torque with a steady-state error below 0.4%, independently of the operating point. Three representative disturbance events—an external impulse, a partial traction loss (slip), and a complete contact loss (drop-off)—are shown to be distinguishable in the time-domain residual through their amplitude and recovery behavior. A comparison against a 1000-pulse-per-revolution encoder shows that the Hall-plus-PLL configuration matches or exceeds the encoder over 80% of the operating envelope, while requiring no additional sensors. These results indicate that Hall-only sensing is sufficient for reliable wheel friction-change detection over a wide operating range, providing a practical, deployable basis for proprioceptive terrain monitoring. Ongoing work extends the approach to four-wheel platforms with active step-torque probing for quantitative friction estimation.
|
| |
| 09:30-10:30, Paper FrPo6P.10 | |
| Force-Manipulability-Aware Whole-Body Optimization for Quadruped Manipulators |
|
| Choi, Kanghyeon | University of Seoul |
| Hwang, Myun Joong | University of Seoul |
Keywords: Robot Mechanism and Control, Robotic Applications
Abstract: Quadruped manipulators require whole-body postures that ensure both end-effector reachability and effective force transmission during contact tasks. This paper presents a posture optimization method based on directional force manipulability for pulling tasks. Given a target end-effector pose and pulling direction, the method searches over the object-relative body distance, body height, and body pitch. For each candidate, arm and leg inverse kinematics are solved, and candidates are filtered using end-effector pose accuracy and support feasibility constraints. The feasible posture with the maximum directional force manipulability is selected as the whole-body configuration. The method is evaluated in Isaac Sim using a Go1 quadruped equipped with a PiPER manipulator and randomly sampled target poses. The results show that feasible postures can be generated across diverse targets and that body placement substantially affects directional force manipulability even when the same target end-effector pose is satisfied.
|
| |
| 09:30-10:30, Paper FrPo6P.11 | |
| Predictive versus Reconstructive Terrain Representations for Quadrupedal Locomotion under Depth Degradation |
|
| Kim, Taehyeong | Kyungpook National University |
| Lee, Sangmoon | Kyungpook National University |
Keywords: Robot Mechanism and Control, Artificial Intelligence Systems, Robot Vision
Abstract: Terrain-aware legged locomotion increasingly depends on a single onboard depth camera, yet real depth degrades through holes, occlusion, and dropout. Reconstruction-based encoders learn a latent by reproducing a height map, forcing it to account for every observed value and tying its quality to the sensor. We ask whether predictive representations—which predict the future latent rather than reconstructing the terrain—degrade more gracefully. We present a locomotion backbone with a swappable representation head for a controlled comparison: a fixed PointNet encoder, a proprioceptive encoder that also estimates base velocity, a recurrent memory, and an implicit terrain-aware RL policy are held identical, and only the head is switched between reconstruction and a JEPA-style latent prediction. On the Unitree Go2 in Isaac Lab, both heads walk under clean depth, but as it is corrupted the reconstructive representation degrades sharply while the predictive one degrades gracefully and keeps walking—consistently across dropout, holes, and occlusion.
|
| |
| 09:30-10:30, Paper FrPo6P.12 | |
| Analysis of the Propulsion Performance of a Coaxial Counter-Rotating UAV with Propeller Blade Numbers |
|
| Kang, ChanHwi | Poongsan |
| Yoo, SungJoo | Poongsan |
| Song, Yi-Hwa | Poongsan |
| Jang, JaeHun | Pukyong National University |
Keywords: Robot Mechanism and Control, Robotic Applications, Navigation, Guidance and Control
Abstract: Coaxial counter-rotating propulsion systems provide high thrust density within a compact airframe, making them suitable for uncrewed aerial vehicles (UAVs). However, the aerodynamic interaction between the upper and lower rotors causes propulsion characteristics that differ from those observed in single-motor tests. This study investigates the effect of propeller blade number on the propulsion performance of a coaxial counter-rotating UAV. A preliminary experiment was conducted using a single motor to compare the performance of two-blade and four-blade propellers under identical operating conditions. The results showed that the two-blade propeller generated higher thrust than the four-blade propeller. To evaluate the actual flight performance, both propeller configurations were applied to a coaxial counterrotating UAV, and flight tests were performed under identical flight conditions. Flight efficiency was evaluated based on flight tests conducted from takeoff until the battery voltage decreased to 17V. Although the two-blade propeller exhibited superior thrust performance in the single-motor experiment, the four-blade propeller achieved a longer flight time during the flight test. These results indicate that single-motor performance alone is insufficient for evaluating the propulsion efficiency of coaxial counter-rotating UAVs, and that aerodynamic interactions between coaxial rotors should be considered when selecting propeller configurations.
|
| |
| 09:30-10:30, Paper FrPo6P.13 | |
| A Fuzzy Null-Space Controller for Human-Like Elbow Configuration in Dual-Arm 7-DOF Humanoid Manipulators Based on Palm Orientation |
|
| Lee, Hunjo | Korea University of Science and Technology, Korea Institute of Industrial Technology |
| Yang, Gi-Hun | KITECH |
Keywords: Robot Mechanism and Control, Robotic Applications, Control Theory and Applications
Abstract: For a 7-DOF anthropomorphic manipulator, the redundant degree of freedom left by the end-effector task must be resolved to obtain a human-like posture. Prior approaches rely on either a structure-specific analytic swivelangle model or a posture prior learned from human motion-capture data. This paper proposes a lightweight null-space controller that infers a preferred elbow direction directly from the real-time palm orientation using a rule-based fuzzy model, combined with manipulability maximization. Since all variables are defined purely from forward kinematics and the Jacobian, the method needs no structure-specific derivation and no training data, and is readily portable across 7-DOF arm designs. On a dual-arm UFactory xArm7 platform in MuJoCo, the elbow direction converges to the fuzzy-inferred target as the palm orientation is swept, while the end-effector position error remains negligible.
|
| |
| 09:30-10:30, Paper FrPo6P.15 | |
| DUET-SLAM: Dual-Flow Consistency for Dense Monocular SLAM in Dynamic Environments with Geometry-Aware Foundation Model |
|
| Jeon, Jinwoo | KAIST |
| Seo, Dong-Uk | Korea Advanced Institute of Science and Technology |
| Myung, Hyun | KAIST (Korea Advanced Institute of Science and Technology) |
Keywords: Robot Vision, Robotic Applications, Artificial Intelligence Systems
Abstract: Geometry-aware foundation models have recently enabled monocular visual simultaneous localization and mapping (SLAM) to perform camera tracking and dense reconstruction without known camera intrinsics. In SLAM systems that leverage geometry-aware foundation models, estimating dense pointmap correspondences is essential for pose optimization. However, such correspondences are typically estimated under a static-scene assumption, and moving objects can corrupt dense correspondences and degrade pose estimation in dynamic environments. To tackle this problem, we propose DUET-SLAM, a training-free dense monocular SLAM framework that improves the robustness of SLAM systems based on geometry-aware foundation models in dynamic environments. Our key insight is that dense correspondences can be interpreted as an implicit optical flow. We introduce a training-free flow-consistency mechanism that compares the implicit flow with optical flow estimated using an off-the-shelf network. The resulting flow discrepancy is utilized to identify motion-inconsistent correspondences associated with dynamic objects. Experiments on the dynamic sequences of the TUM RGB-D dataset demonstrate that DUET-SLAM improves camera pose estimation accuracy compared with the baseline.
|
| |
| 09:30-10:30, Paper FrPo6P.16 | |
| Depth-Supervised Event Gaussian Splatting for Accurate Surface Reconstruction |
|
| Kim, Yunsoo | KAIST |
| Lee, Taeji | Korea Advanced Institute of Science and Technology |
| Myung, Hyun | KAIST (Korea Advanced Institute of Science and Technology) |
Keywords: Robot Vision, Sensors and Signal Processing, Robotic Applications
Abstract: Event cameras have emerged as a compelling sensing modality for novel view synthesis, offering high temporal resolution, low latency, and high dynamic range. However, existing event-based 3D Gaussian splatting methods suffer from inaccurate surface geometry, as flat surface regions with uniform appearance generate sparse or no events, leaving interior surface areas severely under-supervised. In this paper, we propose depth-supervised event Gaussian splatting (DSEGS), a framework that integrates two complementary geometric supervision signals into event-based 3D Gaussian splatting. First, for each event frame, we estimate a dense depth map using a pretrained monocular depth estimation model and align it with the rendered depth via scale-shift fitting. Second, we enforce multi-frame normal consistency by computing world-space surface normals from rendered depth maps and penalizing angular differences at depth-warped corresponding points across frames, leveraging the constraint that the same physical surface must exhibit consistent normals from any viewpoint. Experiments on the Replica dataset demonstrate that DSEGS significantly improves surface accuracy over existing event-based methods.
|
| |
| 09:30-10:30, Paper FrPo6P.17 | |
| Reliability-Weighted Active Gaussian Reconstruction under Camera Pose Uncertainty |
|
| Oh, Sangcheol | Kwangwoon University |
| Oh, Junghyun | Kwangwoon University |
Keywords: Robot Vision, Navigation, Guidance and Control, Artificial Intelligence Systems
Abstract: Active Gaussian reconstruction aims to efficiently reconstruct a scene by selecting informative viewpoints under a limited observation budget. Recent Gaussian Splatting-based active reconstruction methods guide viewpoint se- lection using the current map state, but they commonly assume reliable camera poses during map updating. In practical robotic scenarios, camera poses can contain translational and rotational errors, and observations acquired with inaccurate poses may be incorrectly accumulated into the Gaussian map. Since the updated map state is reused for subsequent view- point selection, such unreliable accumulation can distort the confidence map and degrade reconstruction performance. In this paper, we propose a reliability-weighted map update scheme for active Gaussian reconstruction under camera pose uncertainty. The proposed method keeps the viewpoint selection process of ActiveGS unchanged, while modifying how newly acquired RGB-D observations are accumulated into the Gaussian map. The input observation is first locally aligned with the current Gaussian map, and a frame-level pose reliability score is estimated from the residual depth discrepancy and geometric sensitivity. This reliability score, together with distance-dependent attenuation, adjusts the observation con- tribution of each Gaussian primitive during confidence update. Experiments on Replica indoor scenes under controlled pose-noise settings show that the proposed method consistently improves both peak signal-to-noise ratio (PSNR) and completion ratio compared with ActiveGS.
|
| |
| 09:30-10:30, Paper FrPo6P.18 | |
| 360DVIO: Deep Visual-Inertial Odometry Using a 360-Degree Camera |
|
| Lee, Seunghun | KAIST (Korea Advanced Institute of Science and Technology) |
| Nam, Jihun | KAIST (Korea Advanced Institute of Science and Technology) |
| Myung, Hyun | KAIST (Korea Advanced Institute of Science and Technology) |
Keywords: Robot Vision, Navigation, Guidance and Control, Sensors and Signal Processing
Abstract: 360-degree cameras keep the whole surroundings in view even under aggressive 6-DoF motion and are therefore well suited for state estimation on unmanned aerial vehicles (UAVs) and mobile robots. In such dynamic scenarios, metric scale recovery is critical, yet the scale ambiguity inherent to monocular vision remains, and the equirectangular projection (ERP) distorts scene geometry unevenly across latitude, further complicating depth estimation from visual cues. At the same time, the full-sphere field of view is geometrically complementary to inertial measurements and supports stable gravity estimation and scale initialization in visual-inertial odometry (VIO). Existing 360-degree VIO methods rely on handcrafted features that degrade under illumination changes and low-texture conditions, while learning-based 360-degree odometry remains visual-only and unable to recover metric scale. We propose a learning-based VIO framework for a single 360-degree camera. It integrates a pre-trained recurrent network for equirectangular patch correspondence into a tightly coupled factor graph, with the reprojection residual adapted to ERP spherical geometry. Experiments on public benchmark datasets show that the proposed framework recovers metric-scale trajectories and achieves the lowest trajectory error on four of the five evaluated sequences, including low-light and aggressive-motion settings.
|
| |
| 09:30-10:30, Paper FrPo6P.19 | |
| Geometry-Preserving Sim-To-Real Data Generation for Centroid-Stable Visual Servoing in EV Battery Disassembly |
|
| Park, Seong Eun | Hanyang University, THOTHInc |
| Lee, Sang Hyoung | Korea Institute of Industrial Technology |
| Cho, Nam Jun | Hanyang University |
| Kwon, Taesoo | Carnegie Mellon University |
Keywords: Robot Vision, Process Control Systems, Industrial Applications of Control
Abstract: End-of-life EV battery disassembly requires precise vision-guided bolt alignment under hazardous and non-standardized conditions. In Image-Based Visual Servoing (IBVS), the predicted bolt centroid directly drives robot motion; therefore, even a small centroid bias can exceed socket tolerances and cause alignment failure. However, accurately annotated centroid data are difficult to obtain because labeling requires expert knowledge and direct access to each battery pack. To address this bottleneck, we propose a geometry-preserving sim-to-real data generation method that separates bolt geometry from appearance variation. Pixel-accurate rendered geometry and ground-truth centroids are fixed first, and diffusion-based appearance variation is introduced only through constrained composition. A geometric consistency filter, ShapeCons, then removes samples with residual boundary or center-mark distortion that could shift centroid predictions. In real-robot experiments, the proposed method achieves 89.5% closed-loop IBVS alignment success at the ≤2 px criterion, yielding a 42.1 percentage-point improvement over an appearance-augmentation baseline without geometry filtering. These results show that geometric consistency in synthetic data is critical for reliable visual servoing, even when conventional detection accuracy remains high.
|
| |
| 09:30-10:30, Paper FrPo6P.20 | |
| UToM: Uncertainty-Aware Token Merging for Efficient Semantic Segmentatio |
|
| Kim, Wanhee | Korea Advanced Institute of Science and Technology |
| Sung, Chang Ki | KAIST |
| Myung, Hyun | KAIST (Korea Advanced Institute of Science and Technology) |
Keywords: Robot Vision, Artificial Intelligence Systems, Autonomous Vehicle Systems
Abstract: Vision Transformer-based semantic segmentation achieves strong performance, but its quadratic self-attention cost makes high-resolution dense prediction computationally expensive for resource-constrained systems. Token merging reduces this cost by combining redundant tokens, but existing merging criteria often rely primarily on feature similarity, which may not fully capture semantic reliability in dense prediction. To address this limitation, this paper proposes UToM, an uncertainty-aware token merging method for efficient semantic segmentation. UToM predicts token-level semantic probabilities from intermediate features using a lightweight auxiliary head and derives pair-level semantic confidence for candidate token pairs. This confidence is combined with feature similarity to compute an uncertainty-aware merge score. By suppressing semantically unreliable merges and encouraging merges between tokens that confidently support the same semantic class, UToM improves the accuracy-efficiency trade-off in semantic segmentation. Experimental results show that semantic confidence provides an effective guidance signal for reliable token merging and enables improved accuracy-efficiency trade-offs in dense prediction.
|
| |
| 09:30-10:30, Paper FrPo6P.21 | |
| VLM-Driven Characterization of Action Points in Unknown Environments for Autonomous Robots |
|
| Häuselmann, Ramona | Luleĺ University of Technology |
| Valdes Saucedo, Mario Alberto | Lulea University of Technology |
| Kanellakis, Christoforos | LTU |
| Nikolakopoulos, George | Luleĺ University of Technology |
Keywords: Robot Vision, Robotic Applications, Artificial Intelligence Systems
Abstract: Field robot deployments increasingly rely on scene representations that characterize actionable environment elements and their relationship to robot capabilities and tasks. This paper presents a hierarchical perception-reasoning architecture based on a VLM-driven scene representation that outputs the state and location of observed action points in the visited environment. A vision-language-based reasoning framework fuses RGB, depth, and voxelized occupancy representation into a spatial world model, extracting high-level action points with semantic context annotations. These VLM-proposed scene highlights are stored in a candidate database and actively refined using a best-view selection module that selects viewpoints to improve observation quality. A task-level reasoning module then operates over the action point representation to annotate them with additional information such as semantic properties relevant to robot missions. Real-world experiments in indoor environments demonstrate robust operation, improved viewpoint selection, and scene understanding for autonomous field robots.
|
| |
| 09:30-10:30, Paper FrPo6P.22 | |
| Chunk-Based Diffusion VLA with State-Free Inference for Dual-Arm Humanoid Robot Simulation |
|
| Sanaullah, Sanaullah | Korea Institute of Machinery and Materials |
| Kumar, Abhishek | Korea Institute of Machinery and Material |
| Abbasi, Saad Jamshed | Pusan National University |
| Kim, Jeong Yong | Korea Institute of Machinery and Materials |
| Han, Byung-Kil | Korea Institute of Machinery and Materials |
| Park, Dongil | Korea Institute of Machinery and Materials (KIMM) |
Keywords: Robot Vision, Artificial Intelligence Systems, Robotic Applications
Abstract: We evaluate StarVLA — a Qwen2.5-VL-3B vision-language backbone with a flow-matching diffusion action head — on the Robotis SG2 AI Worker in NVIDIA IsaacLab, comparing two separately fine-tuned variants: a state conditioned variant, trained and evaluated with proprioceptive joint state as an additional input, and a state-free variant, trained and evaluated on image-language pairs only, across 50 fully randomized trials per variant on a single manipulation task. The state-free variant achieves 66% task success versus 28% for the state-conditioned variant (z = 3.81, p < 0.001, two-proportion test), which we hypothesize is related to the added learning burden of jointly modeling proprioceptive state alongside vision and language during fine-tuning, though we do not establish this mechanism causally. Both models overfit beyond their optimal checkpoints (200k and 30k steps respectively) despite training to 400k, producing overconfident and unstable outputs — a practical finding for VLA checkpoint selection. Self-correction, where the robot visually re-assessed and redirected to the correct target, was observed in 14% of state-free trials exclusively. This study establishes a simulation baseline on this task and platform for VLA design choices prior to real-robot transfer.
|
| |
| 09:30-10:30, Paper FrPo6P.23 | |
| Mapless Indoor Room Finding Using Directional Signs and Room Number Recognition |
|
| Choi, Eun Hyuk | Department of Electronic Engineering, Kyung Hee University |
| Kim, Donghan | Kyung Hee University |
Keywords: Robot Vision, Navigation, Guidance and Control, Artificial Intelligence Systems
Abstract: Finding a target room in an indoor corridor environment is difficult to solve using only obstacle avoidance or local goal following. The robot must read directional signs at intersections, select the corridor branch that contains the target room, and then verify the room numbers beside the doors to reach the destination. In this paper, we propose a mapless semantic navigation system for indoor room finding without using a pre-built metric map. The proposed system uses directional signs and room numbers as semantic landmarks. Based on RGB camera perception, it determines both the corridor branch to follow and whether the target room has been reached. A YOLO-World-based detector first detects directional signs and door-plate regions. Then, a VLM and OCR read the room-number ranges on directional signs and the observed room numbers on door plates. The recognized semantic information is stored in a Scene Graph-based memory. The Rule Engine compares the target room number, sign ranges, and observed room-number sequence to determine the driving intent. Branch lock and target tracking conditions are also introduced to reduce incorrect branch selection and false approaches caused by intermediate room numbers. Experiments in an Isaac Sim indoor corridor environment show that the proposed system can find a target room using semantic landmarks without a prior map or predefined goal coordinates.
|
| |
| 09:30-10:30, Paper FrPo6P.24 | |
| Projective Monocular Depth Correction for 2D Occupancy Grid Mapping |
|
| Yoon, Chaehyun | Sookmyung Women's University |
| Lee, Alex | Sookmyung Women’s University |
Keywords: Robot Vision, Navigation, Guidance and Control, Sensors and Signal Processing
Abstract: 2D occupancy grid maps describe where a robot can travel and where obstacles are located, and are widely used for efficient robot navigation services. For cost-effective service robots, it is useful to update such maps from camera-based sensing rather than relying only on metric range sensors. Recent monocular depth models can convert RGB images into dense depth estimates, from which obstacle-height range measurements can be extracted as a camera-derived scan for 2D mapping. However, this scan is not a metric range scan. Frame-wise act either on the robot pose through SE(2) scan matching or on the whole scan through a single uniform scale, but these models cannot fully explain camera-induced range distortion. We propose a range-measurement correction method for camera-derived 2D occupancy mapping. A visual trajectory prior provides the initial map-frame pose, which is refined by a local SE(2) residual, while the camera-derived range measurements are corrected by a regularized determinant-normalized SL(3) measurement transform in scaled bearing and log inverse-range coordinates. The corrected ranges are fused by conservative ray casting along the original bearings without deforming the accumulated map. We evaluate the method against SE(2)-only and Sim(2)-scale baselines on indoor robot sequences, using LiDAR maps only as ground-truth references.
|
| |
| 09:30-10:30, Paper FrPo6P.25 | |
| Uncertainty-Aware Filtering for Robotic Grasping in Limited Observation |
|
| Choi, Youngtae | Pukyong National University |
| Lee, Munhaeng | Pukyong National University |
| Suh, Jinho | Pukyong National University |
Keywords: Robot Vision, Artificial Intelligence Systems, Robotic Applications
Abstract: Robust robotic grasping in limited observation environments remains challenging due to sensor occlusions and incomplete point clouds. Although deterministic shape completion networks can reconstruct missing geometry, they fail to quantify epistemic uncertainty, frequently leading to risky grasp proposals on hallucinated surface artifacts. To address this issue, we present a novel uncertainty-aware grasp filtering framework that directly bridges geometric perception and manipulation task spaces. We integrate Monte Carlo Dropout into a PoinTr backbone to estimate point-wise epistemic uncertainty based on the spatial variance of multiple stochastic forward passes. These point-wise uncertainties are then propagated into the grasp domain by aggregating them within the 3D bounding box of gripper candidates generated by GraspGen, formulating a density-independent grasp-space uncertainty metric. Leveraging a Top-K_s ranking strategy, our framework isolates unreliable surface artifacts and establishes a safety-verified grasp pool to select the optimal pose. Physics-based simulations in MuJoCo demonstrate that our framework substantially enhances both geometric reconstruction quality and downstream grasp success rates over conventional baselines.
|
| |
| 09:30-10:30, Paper FrPo6P.26 | |
| Occupancy-Grounded Room Segmentation for Hierarchical 3D Scene Graphs |
|
| Cueto Zumaya, Carlos Roberto | University of Turku |
| Catalano, Iacopo | University of Turku |
| Peńa Queralta, Jorge | Zurich University of Applied Sciences |
| Bessa, Wallace M. | Federal University of Rio Grande Do Norte |
Keywords: Robot Vision, Artificial Intelligence Systems, Robotic Applications
Abstract: Hierarchical 3D scene graphs (3DSGs) for indoor robots organize geometric and semantic information across spatial scales, with a room layer that connects object-level perception to room-scale reasoning. Existing systems construct this layer from different spatial substrates (e.g., place clusters, wall planes, or segmentation outputs), and as a result, room nodes are not evaluated on a common geometric criterion. We present an occupancy-grounded 3DSG pipeline in which room nodes are anchored to tracked free-space regions derived from occupancy decomposition, giving each room an explicit polygonal footprint. We evaluate the pipeline on 12 Matterport3D scenes by matching predicted room polygons to annotated room instances and compare against Hydra, a representative state-of-the-art place-connectivity baseline. The results show that occupancy-grounded anchoring recovers substantially more room instances than place-connectivity construction, at the cost of lower precision, and that wall-accurate room boundaries remain an open problem for both methods. Code is available at https://github.com/crcz25/OccuSG.
|
| |
| 09:30-10:30, Paper FrPo6P.27 | |
| HEDGe-Net: Geometric Priors and a Design Pattern for Off-Road 3D LiDAR Semantic Segmentation |
|
| Ji, Youngseo | Gachon University |
| An, Jhonghyun | Gachon University |
Keywords: Robot Vision, Autonomous Vehicle Systems, Artificial Intelligence Systems
Abstract: Off-road 3D LiDAR semantic segmentation departs sharply from the urban regime that dominates the literature: vegetation is vertically continuous, the ground is class-bearing, and sparse distant safety-critical structures (poles, fences) dissolve into noise. We present a systematic study and a strong baseline on RELLIS-3D. (i) We characterize off-road failure as a four-mode taxonomy concentrating 96.4% of off-diagonal error mass, mapping each mode to a missing geometric prior and a specific inductive-bias mismatch of urban-tuned backbones. (ii) We report a new off-road state of the art - 49.24% float32 mIoU (the official, saturation-prone metric; 54.66% under a saturation-free int64 recount we introduce, for which no literature value yet exists) - from a single PT-v3 + 3D-SSL (Concerto) model on par with our prior five-member ensemble (49.04 / 54.45) - one model replacing five; a property of that recipe, not of the priors in (iii). A matched-recipe decomposition shows off-road performance is representation-bound: self-supervised initialization supplies nearly all of the gain, while post-hoc architectural priors give diminishing returns. (iii) We distill the discipline of safely augmenting a warm-init backbone into a reusable Design Pattern, instantiated by two transparently-reported geometric priors (HSE, DIG) whose headline effect is indistinguishable from zero at our seed budget - itself evidence for the representation-bound finding, not a separate performance claim.
|
| |
| 09:30-10:30, Paper FrPo6P.28 | |
| 3D Reconstruction of Tangled Deformable Linear Objects Using RGB-Guided Depth Fusion |
|
| Yoon, Youngbin | Chonnam National University |
| Hong, Ayoung | Chonnam National University |
Keywords: Robot Vision, Robotic Applications
Abstract: Accurate 3D shape reconstruction of Deformable Linear Objects (DLOs) such as ropes and cables is a fundamental prerequisite for robotic manipulation tasks, including knot-tying and cable routing. Existing methods rely either on color-based segmentation, which fails for monochrome objects, or provide only 2D centerline estimates without depth information, making it difficult to resolve crossing ambiguities under self-occlusion. We propose a training-free pipeline that reconstructs the full 3D shape of a monochrome DLO from RGB-D input. Boundary detection is performed by fusing RGB-estimated depth from Depth Anything V2 (DAV2) with sensor-measured depth from an Intel RealSense D435, where the two signals complement each other to capture boundaries that are otherwise difficult to detect. The detected boundaries partition the DLO mask into interior regions, from which centerline segments are connected via a direction-aware greedy matching algorithm. When the topology cannot be uniquely determined from a single observation, alternative candidates are maintained and refined with additional observations. Quantitative evaluation shows that our method achieves an F1 score of 97.8%, outperforming Canny, PiDiNet, and DiffusionEdge, and qualitative comparisons with mBEST and RT-DLO confirm its advantage for reconstructing monochrome tangled ropes.
|
| |
| 09:30-10:30, Paper FrPo6P.29 | |
| HD-LIO: Bidirectional Degeneracy Feedback with Hierarchical Surfel Mapping for LiDAR-Inertial Odometry |
|
| Hong, Euntae | LGE |
| Choi, Sungjin | LG Electronics Inc |
| Cho, Beom-Jin | LG Electronics Inc |
| Noh, DongKi | LG Electronics Inc |
Keywords: Robotic Applications, Navigation, Guidance and Control, Robot Vision
Abstract: LiDAR-inertial odometry (LIO) provides accurate real-time localization and mapping for autonomous navigation. Its accuracy, however, can degrade in feature-sparse scenes where scan to map residuals provide weak directional constraints. The problem is more severe with narrow-FOV LiDAR because the scan geometry limits the observable directions and allows drift to enter both the state estimate and the map. Existing tightly coupled LIO methods mainly handle this effect in the state update, while the map can still integrate biased points along weakly constrained directions. Asymmetry between the estimator and the map can sustain a drift map feedback loop that is not addressed by estimator side degeneracy handling alone. We present HD-LIO, a hierarchical surfel map LIO system with bidirectional degeneracy feedback. HD-LIO uses degeneracy cues from the information matrix to attenuate fine voxel map updates along rank deficient directions, restore directional constraints through a per query voxel cache correspondence source, and preserve that auxiliary source with an update gate under persistent degeneracy. A surfel eigenspectrum reliability weight addresses local measurement uncertainty in each scan to map residual, complementing a directional linearized error model of the joint state map dynamics underlying these map side operations. We evaluate HD-LIO on public dataset splits covering Livox Avia, Livox Mid-360, and Ouster OS1-16 data. HD-LIO achieves the lowest mean trajectory error in every sensor specific sequence group against the baseline methods. On the narrow-FOV indoor Avia sequences, HD-LIO remains bounded over complete runs while the baselines diverge.
|
| |
| 09:30-10:30, Paper FrPo6P.30 | |
| Reflection-Robust Stereo Visual-Inertial Odometry Via Kalman-Filtered Ground-Height Estimation |
|
| Hwang, Uihyun | KAIST |
| Shin, Sungjae | Korea Advanced Institute of Science and Technology (KAIST) |
| Kim, Dongjae | KAIST |
| Myung, Hyun | KAIST (Korea Advanced Institute of Science and Technology) |
Keywords: Robot Vision, Autonomous Vehicle Systems, Artificial Intelligence Systems
Abstract: Visual-inertial odometry relies on visual features that correspond to 3D points with consistent projections across sequential frames. In reflective indoor-floor environments, however, specular floor patterns can be detected and triangulated as virtual features below the physical ground plane, producing unreliable feature associations and inconsistent visual residuals. This paper proposes a lightweight reflection-aware front-end for feature-based stereo VIO. Rather than suppressing all floor-region features, the proposed method estimates the ground height in real time and rejects only feature candidates whose inferred 3D positions are below the estimated ground height. A floor-region mask restricts the validation process to reflection-prone regions, while stereo height uncertainty and pose uncertainty are used to construct frame-level ground-height measurements. A scalar Kalman filter with Normalized Innovation Squared (NIS) gating then provides a stable ground-height estimate, and an uncertainty-aware safety margin is used for selective below-ground virtual feature rejection. Experiments on real reflective-floor indoor sequences show that the proposed front-end generally improves trajectory accuracy in reflection-dominant environments while maintaining accurate ground-height estimates and recovering from actual ground-height changes.
|
| |
| 09:30-10:30, Paper FrPo6P.31 | |
| Boundary-Aware Picking Center Generation for Foreign Object Removal in Continuous Dried-Red-Pepper Conveyor Frames |
|
| Shin, Sumin | Korea Institute of Industrial Technology (KITECH) |
| Lee, Yong Jun | Korea University |
| Inpyo, Lee | Korea Institute of Industrial Technology |
| Ahn, Woo Jin | Inha University |
| Myotaeg, Lim | Korea University |
| Lee, Kwang Hee | Korea Institute of Industrial Technology |
Keywords: Robot Vision, Robotic Applications, Industrial Applications of Control
Abstract: To bridge anomaly detection outputs with robotic manipulation, this work presents a decision-making layer tailored for spatial boundary resolution on a dried-red-pepper transport line. Standard frame-by-frame analysis often yields redundant picking actions because the same anomaly reappears across overlapping image boundaries. To mitigate this, our approach transforms candidate coordinates into localized targets via a smallest enclosing circle (SEC) formulation constrained by the effective suction radius Rend . Furthermore, the boundary region is partitioned into an overlap-body and a tail-zone, enabling dynamic assignment between immediate frame execution and cross-frame carry-over. In real conveyor test runs, the framework successfully streamlined target candidates from 29.68 to 12.58 target centers per frame while securing a finalized coverage rate of 1.000. Integrated physical testing demonstrated an overall extraction accuracy of 90.0% alongside a low latency of 97.32 ms for the vision-to-command pipeline.
|
| |
| 09:30-10:30, Paper FrPo6P.32 | |
| Bridging the Sim-To-Real Gap in Quadrupedal Visual Navigation Via 3D Gaussian Splatting |
|
| Lee, Minseok | Handong Global University |
| Han, Seongyu | Handong Global University |
| Mun, Huijae | University |
| Hwang, Sung Soo | Handong Global University |
Keywords: Robot Vision, Navigation, Guidance and Control, Robotic Applications
Abstract: Deploying visual navigation on physical quadrupedal robots is challenging due to the visual sim-to-real gap and locomotion-induced camera jitter. This jitter degrades feature tracking and depth consistency. This paper presents a pipeline that bridges these perceptual and dynamic gaps to enable virtual parameter optimization. In the simulation stage, we reconstruct a 3D Gaussian Splatting (3DGS) digital twin of a campus corridor and use a reinforcement learning (RL) locomotion policy to generate camera bobbing noise. Under these simulated visual and dynamic disturbances, the parameters of the RTAB-Map visual SLAM and Nav2 stacks are optimized. For deployment, the RL policy is bypassed, and velocity commands are routed directly to the physical robot’s built-in Sport API via a ROS 2 bridge. Real-world experiments demonstrate that parameters optimized under simulated jitter transfer directly to the physical platform, supporting visual mapping and qualitative obstacle-aware replanning without manual recalibration.
|
| |
| 09:30-10:30, Paper FrPo6P.33 | |
| Brick3D: Brick Detection and Localization Based on VLMs for Masonry Construction Drones |
|
| Valdes Saucedo, Mario Alberto | Lulea University of Technology |
| Stamatopoulos, Marios-Nektarios | Luleĺ University of Technology |
| Kanellakis, Christoforos | LTU |
| Nikolakopoulos, George | Luleĺ University of Technology |
Keywords: Robot Vision, Robotic Applications, Artificial Intelligence Systems
Abstract: This paper presents Brick3D, a monocular camera-based perception framework for UAV-assisted masonry construction that enables the detection, condition assessment, and metric localization of construction bricks using only onboard visual sensing. The proposed system leverages vision-language models (VLMs) to perform zero-shot semantic segmentation, allowing robust brick detection without task-specific retraining and enabling adaptation to varying brick appearances and construction environments. A coverage-based integrity assessment method evaluates the correspondence between segmented regions and their expected geometry, classifying bricks as undamaged, damaged, or uncertain under partial visibility conditions. For localization, the framework combines refined corner observations with a Perspective-n-Point (PnP) formulation to estimate the full 6-DoF pose of each brick relative to the onboard camera. To improve monocular localization accuracy, an offline nonlinear calibration model and Kalman filtering stage are introduced to compensate for systematic bias and reduce temporal noise. Experimental validation in both controlled laboratory and real-world indoor and outdoor environments demonstrates reliable brick detection, robust damage classification, and accurate localization.
|
| |
| 09:30-10:30, Paper FrPo6P.34 | |
| Real-Time Classification and Pose Estimation of Humans and Legged Humanoids |
|
| Kang, GyeongRim | JeonBuk National University |
| Kim, KangGeon | Jeonbuk National University |
| Jo, HyungGi | Jeonbuk National University |
Keywords: Robot Vision, Human-Robot Interaction, Artificial Intelligence Systems
Abstract: Humans and legged humanoids may be observed in the same workspace. The difficult cases considered in this study occur when their image trajectories cross and a temporary occlusion interrupts a track. After re-entry, a detection can then be associated with the wrong entity. We address this case with a single-RGB pipeline that performs 2D pose estimation, classification, tracking, and identity recovery together in near real-time. A single-stage multi-person pose estimator provides the keypoints, MobileNetV3-Small supplies track-level class evidence through temporal voting, and a dual-track Re-ID stage handles lost identities. On the recorded occlusion and re-entry sequences, the classifier reached 98.99% accuracy and a macro F1 score of 0.986. The complete pipeline used 181 MB of GPU VRAM on the target platform.
|
| |
| 09:30-10:30, Paper FrPo6P.35 | |
| Adaptive Object Localization with Depth Reliability-Aware Innovation-Gated Kalman Filter for Small Swarm Robot Systems |
|
| Kim, Seonggeon | Kumoh National Institute of Technology |
| Lee, Heoncheol | Kumoh National Institute of Technology |
Keywords: Robot Vision, Robotic Applications, Sensors and Signal Processing
Abstract: This paper addresses unstable RGB-D-based object localization on resource-constrained small robot platforms. Noisy depth measurements from individual robots can produce inconsistent object position estimates, which may subsequently degrade object-level map merging. RGB-D depth measurements are susceptible to noise, occlusion, and transient depth spikes, causing inconsistent object positions and duplicated landmarks in terms of map merging. To address this, we propose Depth Reliability-aware Innovation-Gated Adaptive Kalman filter (DRIG-AKF), which stabilizes object position estimates by quantifying depth reliability from valid depth ratio, depth variance, and object distance, and incorporating it into an adaptive measurement covariance. An innovation-gated mechanism suppresses transient depth spikes via Mahalanobis innovation distance gating, and the stabilized positions and their covariance estimates can be transmitted as inputs to a subsequent map-merging stage. The proposed depth reliability model and innovation gating are directly validated using a separate recording with raw per-pixel depth measurements. The overall framework is evaluated on three real-world datasets using bounding-box-derived reliability proxies due to the absence of per-pixel depth logs, demonstrating an average position standard deviation reduction of 66.9% over EMA smoothing and 27.5% over a KF+Hungarian baseline, with reductions of up to 79.5% and 54.1%, respectively, in datasets where the reliability indicators more accurately reflect depth measurement quality.
|
| |
| 09:30-10:30, Paper FrPo6P.36 | |
| Preserving Fine Details in Semantic-Map-Conditioned LiDAR Generation Via Pixel-Space Diffusion |
|
| Lee, Jaewon | Korea Institute of Industrial Technology |
| Kim, Jiwoong | Korea Institute of Industrial Technology(KITECH) |
Keywords: Robot Vision, Artificial Intelligence Systems, Autonomous Vehicle Systems
Abstract: LiDAR sensors provide precise 3D measurements of the surrounding environment and are widely used for perception in autonomous driving and robotics. However, collecting real driving data that covers diverse and rare situations is costly, which has motivated research on synthetic LiDAR scene generation using diffusion models. To reduce computation, most existing methods generate scenes in a compressed latent space and decode them back to the original resolution. This decoding is lossy, and the loss becomes critical when a semantic map is used as a condition, since the condition is also compressed and its detailed structure is not fully preserved. In this paper, we adapt PixelDiT, a pixel- space diffusion transformer, to conditional LiDAR range-image generation. Instead of encoding the range image and the semantic map into a compressed latent space, we inject the semantic map as tokens at native resolution. This helps the generated scan retain thin and small-scale structures that a latent-space model tends to blur. On the SemanticKITTI Semantic-Map-to-LiDAR task, the proposed method follows the semantic condition more faithfully than a latent-space baseline.
|
| |
| 09:30-10:30, Paper FrPo6P.37 | |
| Keyframe-Based LiDAR Scan-To-Map Registration in a Greenhouse Corridor |
|
| Park, Taewook | Ulsan National Institute of Science & Technology |
| Lee, Seunghoon | Korea Advanced Institute of Science and Technology |
| Yun, Won-Jae | METAFARMERS Inc |
| Lee, Kyu-Wha | METAFARMERS Inc |
| Oh, Hyondong | KAIST |
Keywords: Robot Vision, Autonomous Vehicle Systems
Abstract: Greenhouse robots require reliable localization despite unavailable or unreliable Global Navigation Satellite System signals and limited geometric features along crop corridors. We present a LiDAR-based localization method that registers multi-session scans to a unified prebuilt point cloud map using point-to-plane iterative closest point registration and LiDAR--inertial odometry priors. Evaluation in a 50 m-long commercial tomato greenhouse corridor demonstrates consistent multi-session localization in a common map frame, with position and orientation RMSEs below 3.5 cm and 4.8 degree, respectively.
|
| |
| 09:30-10:30, Paper FrPo6P.38 | |
| OR-MR: Structure-Aware Multi-Candidate Planning for Multi-Depot Grid Coverage |
|
| Seo, JangHo | Kyungpook National University |
| Lee, Joonwoo | Kyungpook National University |
Keywords: Robotic Applications, Navigation, Guidance and Control, Artificial Intelligence Systems
Abstract: Multi-depot multi-robot coverage path planning aims to cover all free cells while each robot starts from and returns to its own depot, making the longest closed route the mission-time bottleneck. Existing grid-based planners usually commit to a single decomposition or growth rule, and their makespan can change significantly with obstacle bottlenecks and depot placement. This paper presents OR-MR (Ownership-Routed Multi-Robot), a structure-aware bounded heuristic portfolio for 4-connected grid coverage. OR-MR constructs solutions on a common 2×2 supercell substrate, applies a rule-based Cut Rule using map-cut and depot-geometry features to activate a small set of synchronized balanced supercell growth (SBSG) and Peak-Balance candidates (observed K=1–13 per scenario), and selects the final plan by route-level closed makespan. On 18 corner-anchored scenarios over six 50×50 maps, OR-MR improves the closed makespan in most cases against three representative baselines, with the largest gains on the split-rich Map 4.
|
| |
| 09:30-10:30, Paper FrPo6P.39 | |
| ELite++: Efficient Lifelong LiDAR Mapping Via Window-Based Registrationand Voxel Ratio Maps |
|
| Lee, GeonHui | Seoul National University |
| Gil, Hyeonjae | SNU |
| Kim, Ayoung | Seoul National University |
| Jung, Minwoo | Seoul National University |
Keywords: Robotic Applications, Autonomous Vehicle Systems
Abstract: Long-term LiDAR mapping requires clean and consistent static maps across repeated sessions in dynamic environments. Existing lifelong mapping pipelines combine session alignment, dynamic filtering, and map update, but repeated dense point-wise updates and candidate-dependent alignment can limit scalability. To tackle these issues, we present ELite++, an efficient extension of ELite that combines automatic window-based registration, aligned-session dynamic object removal, and compact Voxel Ratio Map integration. ELite++ removes transient objects while preserving static structures through hit-exposure voxel stability, local ground protection, and structure-aware refinement. Experiments on diverse LiDAR datasets show improved dynamic object removal over existing DOR methods, preserved multi-session alignment quality, and runtime reduced to 11% of the original pipeline.
|
| |
| 09:30-10:30, Paper FrPo6P.40 | |
| Human-In-The-Loop Replanning for Instruction-Guided Object Retrieval in Fire Environments |
|
| Song, Yuri | Kwangwoon University |
| Kim, Yeonjin | Kwangwoon University |
| Oh, Junghyun | Kwangwoon University |
Keywords: Robotic Applications, Human-Robot Interaction, Navigation, Guidance and Control
Abstract: Disaster-response robots must operate under evolving hazards while respecting user instructions. In fire environments, predicted object risk can conflict with the retrieval priority implied by a natural-language instruction: a lower-priority target may become urgent, whereas automatic risk-based replanning may violate the user’s intent. This paper formulates this problem as instruction-conditioned object retrieval under dynamic fire risk. We propose a Priority- Aware Human-in-the-Loop (PA-HITL) replanning framework that extracts target objects and a preferred order, estimates predicted object risk, and presents reorder suggestions when risk-based urgency conflicts with instruction-implied priority. The final execution target is determined through accept/reject decisions instead of automatic reordering. Experiments in the HAZARD fire environment show that, compared with fully automatic risk-based replanning, PA-HITL improves SR from 0.55 to 0.60 in Kitchen and from 0.65 to 0.71 in Craftroom, while maintaining comparable Safe-SR (0.35 vs. 0.36 in Kitchen and 0.40 vs. 0.41 in Craftroom). Compared with No HITL, PA-HITL also improves Safe-SR from 0.20 to 0.35 in Kitchen and from 0.25 to 0.40 in Craftroom.
|
| |
| 09:30-10:30, Paper FrPo6P.41 | |
| Robotic Manipulator Trajectory Stabilization Via Imitation-Guided Reinforcement Learning |
|
| Lee, Hojeong | Korea University |
| Cho, Hyemin | Korea University |
| Lee, Seunghoon | Korea University |
| Park, Cheol Hoon | Korea University |
| Lee, Yong Jun | Korea University |
| Park, Jong-Chan | Korea University |
| Lim, Myo-Taeg | Korea University |
Keywords: Robotic Applications, Robot Mechanism and Control, Artificial Intelligence Systems
Abstract: Imitation learning enables agents to acquire skills by observing and learning from expert demonstrations, and has recently attracted significant attention in robotics. However, these demonstrations do not guarantee optimal behavior, and deploying trajectories generated by imitation learning policies in real-world environments can result in abrupt joint movements, instability, and compounding errors. To address these issues, we propose an imitation-guided reinforcement learning framework that leverages imitation learning trajectories as reference paths and learns a stable trajectory-tracking policy through reinforcement learning. Specifically, we generate reference joint trajectories using an action chunking transformer-based imitation learning policy, and further optimize a control policy via proximal policy optimization to track the trajectories while suppressing joint velocities and action variations. Experimental results demonstrate that the proposed framework reduces joint oscillations by approximately 26% compared to baseline imitation learning, confirming that reinforcement learning enhances control stability while maintaining the task-level behaviors acquired through imitation.
|
| |
| 09:30-10:30, Paper FrPo6P.42 | |
| Temporal Context Conditioning for Action Chunk Continuity in Vision-Language-Action Model |
|
| Shin, Dongin | Korea Electronics Technology Institute |
| Moon, JongSul | Korea Electronics Technology Institute |
| Kim, YoungOuk | Korea Electronics Technology Institute |
| Kim, Wonha | Kyung Hee University |
Keywords: Robotic Applications, Robot Vision, Artificial Intelligence Systems
Abstract: Vision-Language-Action (VLA) models have advanced robot action generation, yet existing approaches lack explicit modeling of temporal continuity at action chunk boundaries, leading to trajectory discontinuity and abrupt velocity changes in long-horizon tasks. This paper proposes Temporal Context Conditioning, a framework that injects the previous action trajectory and vision-language hidden representation into the Flow Matching-based Diffusion Transformer (DiT) via Feature-wise Linear Modulation (FiLM). A Boundary Gap Metric is introduced to quantify inter-chunk discontinuity. On the LIBERO benchmark, the proposed method improves the average success rate by 6.0% over the baseline and reduces the Boundary Gap by 15.8%.
|
| |
| 09:30-10:30, Paper FrPo6P.43 | |
| IRM-Guided Model Predictive Control for Mobile Manipulator |
|
| Choi, JungHyun | University of Seoul |
| Lee, Taegyeom | University of Seoul |
| Hwang, Myun Joong | University of Seoul |
Keywords: Robotic Applications, Robot Mechanism and Control
Abstract: This paper presents an inverse reachability map (IRM)-guided model predictive control framework for end-effector trajectory tracking by mobile manipulators. A precomputed three-dimensional IRM is used to evaluate base regions that are kinematically favorable for each desired end-effector pose. At each prediction step, the IRM is sliced into a two-dimensional base preference map and combined with environmental safety information using a multiplicative safety mask. The local weighted centroid around the predicted base position is then used as an IRM-guided reference in the mobile base MPC objective. The overall structure of the proposed framework is shown in Fig. 1. Unlike conventional IRM-based methods that mainly determine static base placements or select discrete candidate poses, the proposed method incorporates IRM information online to continuously guide the mobile base toward regions that support end-effector reachability and favorable arm manipulability. The manipulator controller tracks the global end-effector reference in the current base frame, while the base MPC accounts for the IRM-guided reference, heading consistency, and input regularization. Gazebo simulations with a Scout 2.0 mobile base and a Franka Panda manipulator show that the proposed IRM-MPC improves the base configuration compared with the baseline, as illustrated in Fig. 2, and reduces end-effector position tracking error while maintaining more consistent Yoshikawa manipulability, as shown in Fig. 3.
|
| |
| 09:30-10:30, Paper FrPo6P.44 | |
| Shared Autonomy for Teleoperation Using Probabilistic Estimation of Primitive Motion Intent |
|
| Lee, Haeseong | Seoul National University |
| Yoon, Junheon | Seoul National University |
| Lee, Yonghee | Seoul National University |
| Park, Jaeheung | Seoul National University |
Keywords: Robotic Applications, Human-Robot Interaction, Artificial Intelligence Systems
Abstract: This paper proposes a shared autonomy method that enables a teleoperated robot system to infer human intent. Teleoperation has significant potential for hazardous-environment operations and human demonstration data collection. However, teleoperation remains challenging due to limited feedback, communication latency, and embodiment mismatch between humans and robots. To improve teleoperation efficiency, various shared autonomy methods, such as policy blending, have been actively studied. However, existing approaches often require task-specific robot actions and may reduce the operator’s control authority. To address these limitations, the proposed method focuses on estimating the primitive motion intent of the operator, including translation-only, rotation-only, and combined motions. For this purpose, a lightweight estimator is developed to predict the probability of each motion primitive. Also, based on the probabilistic estimation, unintended human command is suppressed while preserving the control authority. Finally, real-robot experiments demonstrate that the proposed framework improves task efficiency during teleoperation.
|
| |
| 09:30-10:30, Paper FrPo6P.45 | |
| Actuator-Aware Inverse Kinematics with Joint-Limit Admissibility for Torque-Controlled Redundant Robots |
|
| Dastranj, Mohammad | Tampere University |
| Hejrati, Mahdi | Tampere University |
| Mattila, Jouni | Tampere University |
Keywords: Robotic Applications, Robot Mechanism and Control, Navigation, Guidance and Control
Abstract: This paper proposes actuator-aware inverse kinematics for torque-controlled redundant robots under joint-limit constraints. In the considered architecture, the inverse-kinematic output is not merely a purely kinematic joint-velocity command; it is the required joint velocity supplied to a downstream torque-level controller. Therefore, a small commanded task residual may not necessarily improve realized motion. The proposed method formulates a convex quadratic programming problem whose decision variable is the joint-level required velocity. Control barrier function style bounds impose reference-level joint-limit admissibility, while the task equation is handled through a penalized slack variable. Redundancy is resolved using a controller-compatibility objective that accounts for previous-command consistency and actuator torque-capacity weighting. The method is independent of the particular torque-level controller and can serve as an intermediate IK layer between an endpoint trajectory and a redundant robot controller. Experiments on a virtual-decomposition-controlled seven-degree-of-freedom upper-limb exoskeleton compare the method with standard inverse-kinematic baselines and a constrained task-preserving quadratic programming baseline. The results indicate substantially lower limit-pushing commands, bounded admissible required velocities, and reduced maximum position error while maintaining competitive realized task tracking in the tested trajectory, without modifying the downstream controller.
|
| |
| 09:30-10:30, Paper FrPo6P.46 | |
| Design and Dynamic Modeling of a Hybrid Compliant End-Effector with Feedforward Compensation for Robotic Surface Processing of Reclaimed Timber |
|
| Elsayed, Ahmed Sedky Mohamed | NTNU |
| Garammatikos, Sotirios | NTNU |
Keywords: Robotic Applications, Robot Mechanism and Control, Control Theory and Applications
Abstract: Automated surface preparation of reclaimed timber remains a key barrier to wider structural reuse. The main challenge is maintaining consistent contact force over geometrically uncertain surfaces that combine hard contaminants with a soft, damage-sensitive substrate. This paper presents the design and dynamic modeling of a hybrid compliant end-effector that pairs a passive spring-based compliance stage with an active motor-driven feedforward compensator. The mechanism comprises four linear guides, eight bearings, two compression springs (equivalent stiffness 500 N/m), a force sensor, and two anti-twisting guides. The system is modeled as a Simscape Multibody digital twin and evaluated under a ±5 mm sinusoidal surface-height disturbance with an 18 N force setpoint. The feedforward-augmented PID controller reduces the initial contact transient by 46.7%, steady-state RMS force error by 30.8% (from 1.03% to 0.71% of setpoint), and mean absolute error by 32.4% compared with PID alone.
|
| |
| 09:30-10:30, Paper FrPo6P.47 | |
| Demonstration Temporal Resolution Analysis for Visuomotor Diffusion Policy in Precision Electrical Component Assembly |
|
| Kim, Ikjune | Korea Atomic Energy Research Institute |
| Joo, Sungmoon | Korea Atomic Energy Research Institute |
Keywords: Robotic Applications, Artificial Intelligence Systems, Robot Mechanism and Control
Abstract: The automation of electrical component assembly remains challenging due to tight-tolerance insertion requirements. Purely visuomotor approaches offer practical deployment advantages, but introduce an underexplored data collection parameter: demonstration temporal resolution, which refers to the robot's travel distance between consecutive frames in the demonstration training data. It is defined as the per-frame end-effector displacement determined by the robot's end-effector speed and data sampling frequency. This work formalizes its constraints in diffusion-based imitation learning, showing that the feasible range is bounded below by the robot's kinematic repeatability (preventing state aliasing) and above by the task's geometric clearance (avoiding intra-step collisions). Experiments across nine temporal resolution settings with matched training-data volume identify the feasible range and derive experimentally calibrated data collection guidelines, validated on a PDU surrogate with 2.0 mm radial clearance across 150 trials, achieving 80%–92% per component and overall 85%, using the alternative hardware setup.
|
| |
| 09:30-10:30, Paper FrPo6P.48 | |
| Characterizing Hard-To-Automate Manufacturing Processes for Robotic Automation: Data-Driven Difficulty Factors and a Mapping to Robot Technologies |
|
| Lee, Jaeseon | KITECH |
| Im, Subin | SungKyunKwan University |
Keywords: Robotic Applications, Industrial Applications of Control, Artificial Intelligence Systems
Abstract: Korea has the world’s highest manufacturing robot density, yet many processes are still performed manually, and judging which ones are hard to automate has relied on inconsistent qualitative judgment. We build a quantitative basis for that judgment from 593 robotic-automation consulting cases (2016–2025). Of these, 353 non-trivial processes were assessed against eighteen difficulty factors spanning the object, task, and environment. Rather than scoring expert verdicts directly, we reverse-engineer the relationship between the experts’ factor checks and their grades through ordinal regression, yielding data-based weights and seven core factors (led by shape complexity and skill dependency) from which we derive an integer scoring scheme; suppressor factors are identified and excluded. The seven-factor scheme reproduces the expert grades about as well as the full model while remaining simpler (75.9% vs. 74.4% accuracy for high-or-above), and its scores follow the known difficulty ranking across process groups. We further map the core factors onto the enabling technologies (perception, grasping, contact/precision control, and task planning). The gap between the objective-factor model and an expert baseline quantifies how far the judgment can be made explicit, with a residual tied to experience, most clearly in skill dependency. External validation is planned through a national R&D program (2027–2030).
|
| |
| 09:30-10:30, Paper FrPo6P.49 | |
| STG 3.0: Time-Dependent Semantic Topological Graphs for Human-Aware Service Robots |
|
| Park, Jeong-Seop | Korea University |
| Lee, Yong Jun | Korea University |
| Park, Jong-Chan | Korea University |
| Ahn, Woo Jin | Inha University |
| Woo, Jong Jin | LG Electronics |
| Lim, Myo-Taeg | Korea University |
Keywords: Robotic Applications, Artificial Intelligence Systems, Human-Robot Interaction
Abstract: Household service robots must often locate a target user before providing services such as delivery or assistance. However, existing navigation and Vision–Language–Action (VLA) frameworks mainly rely on static semantic information and do not explicitly model human location over time. This paper presents STG 3.0, a time-dependent semantic topological graph for proactive human search. STG 3.0 combines a static Home-Tree representation with dynamic human nodes connected through time-dependent probabilistic edges learned from patrol observations. By integrating user-location probabilities into graph-based planning, the robot prioritizes locations where the target user is most likely to be found. Experiments in simulated and real indoor environments demonstrate reliable semantic mapping and reduced search time compared with conventional search strategies.
|
| |
| 09:30-10:30, Paper FrPo6P.50 | |
| Sensorless Admittance Control Based on Velocity Response Deviation under a Low Control Update Rate |
|
| Kim, Jiseong | Keimyung University |
| Jo, Jaemin | Keimyung University |
| Choi, Junghyun | Keimyung University |
Keywords: Control Theory and Applications, Human-Robot Interaction, Industrial Applications of Control
Abstract: This paper presents a sensorless admittance-control structure for a speed-controlled motor drive under a low control update rate. Velocity-response deviation between the reference and encoder-based angular velocity is converted into a torque-equivalent signal and applied to a virtual inertia-damping model without an external force/torque (F/T) sensor or internal drive signals. Experimental results confirmed force-responsive reference modification consistent with torque-sensor-based control and demonstrated feasibility under the low update-rate condition.
|
| |
| FrPo7P |
Symphony D, 4F |
| Poster Session 7 |
Poster Session |
| |
| 16:10-17:10, Paper FrPo7P.1 | |
| DigitForce: Channel-Decoupled Force Tokenization for Transformer-Based Humanoid Dexterous Manipulation |
|
| Kim, Beomjoon | Korea Electronics Technology Institute |
| Kim, Yunhan | Korea Electronics Technology Institute |
Keywords: Robotic Applications, Artificial Intelligence Systems, Sensors and Signal Processing
Abstract: Contact-rich humanoid manipulation depends on finger-localized contact cues that are difficult to infer from RGB and joint states. We present DIGITFORCE, a channel-decoupled state tokenizer that addresses a fused-state token bottleneck in Action Chunking Transformer (ACT)-based imitation learning. In the standard fused design, proprioception and force readings are projected into one state token, limiting channel-level access to transformer attention. DIGITFORCE keeps the RGB encoder, ACT backbone, decoder, action head, training objective, and deployment interface unchanged. Proprioception is encoded as one token, and each finger-associated force channel is encoded as a separate token. Force readings are used only as observations, with no change to the action space or controller. On a real Unitree G1 with dual five-finger Inspire Hand DFX hands performing a color-sequence block-stacking task, task success rates across 50 rollouts per policy are 62%, 76%, and 92% for Proprio-ACT, Concat-ACT, and DIGITFORCE, respectively, and DIGITFORCE shows shorter failed-grasp recovery intervals and smaller arm-retreat excursions than the fused-token baseline. Matched offline force-channel interventions show that DIGITFORCE produces sharper phase-dependent channel selectivity than the fused-token baseline and induces larger action shifts in hand-joint dimensions. The mean hand-to-arm sensitivity ratio increases from 1.05 in the fused-token baseline to 1.40 in DIGITFORCE. These results suggest that channel-decoupled force tokens provide channel-addressable contact cues for hand-joint action prediction, without tactile skin, a force-control head, or a new action interface.
|
| |
| 16:10-17:10, Paper FrPo7P.2 | |
| Energy-Cost-Based Prioritized Path Planning for Multiple Ground-Aerial Bimodal Robots |
|
| Lee, Dongeun | Kookmin University |
| Lee, Seung-Mok | Kookmin University |
Keywords: Robotic Applications, Navigation, Guidance and Control
Abstract: This paper presents an energy-cost-based prioritized path-planning method for multiple passive ground-aerial bimodal robots operating in obstacle environments. Passive bimodal robots can reduce energy consumption through ground locomotion and overcome obstacles through aerial locomotion, but aerial motion requires significantly higher energy to support the robot weight. To address this trade-off, the proposed method evaluates motion primitive costs using a mode-dependent energy model and constructs a Dijkstra-based heuristic from unit-distance energy costs. Robot paths are planned sequentially according to a fixed priority order using an energy-cost A∗ search over safe interval path planning(SIPP) states. The paths of higher-priority robots are stored in a reservation table and used as time-dependent constraints for lower-priority robots. Simulations were conducted in two obstacle environments: a narrow-gap environment that induces inter-robot congestion and a wall-obstacle environment that requires aerial locomotion. The results show that the proposed bimodal planner reduces travel time compared with ground-only planning and reduces energy consumption compared with aerial-only planning, demonstrating collision-free and energy-efficient multi-robot path planning.
|
| |
| 16:10-17:10, Paper FrPo7P.3 | |
| Automated Strawberry Packaging System Based on Vision Sensing and Re-Orientation Gripper Mechanism |
|
| Shin, WooSeong | Hanyang University |
| Lee, Minsu | Hanyang University |
| Choi, Jeongseok | Hanyang University |
| Seo, TaeWon | Hanyang University |
Keywords: Robotic Applications, Robot Vision, Robot Mechanism and Control
Abstract: This study proposes an integrated robotic system designed to pick randomly arranged harvested strawberries without damage and automatically place them into egg-tray-style containers. To address the growing trend of egg-tray-style strawberry packaging and the challenge of labor shortages, this system implements an automated process that combines mechanical adaptability with vision sensing. First, a two-fingered fin-ray gripper fabricated via silicone casting was utilized to achieve passive adaptive grasping of irregularly shaped strawberries while minimizing physical damage. At wrist, remote center of motion (RCM) mechanism wrist was introduced to ensure efficient placement angles within narrow containers. For the grasping strategy, YOLOv8-pose is employed to extract four keypoints(top, bottom, left, and right) of the strawberry. These points are used to determine the grasping position and angle and estimate the fruit's weight for grade classification. Experimental results confirm that the entire process—from recognition and grasping to weight estimation and packaging—is successfully performed within the integrated system, demonstrating the practical potential for automating strawberry packaging processes that currently rely heavily on manual labor.
|
| |
| 16:10-17:10, Paper FrPo7P.4 | |
| Pre-Execution Feasibility Checking for Language-Guided Robotic Manipulation Via 3D Scene Graph and Object Ontology |
|
| Kim, Haryeong | Sungkyunkwan University |
| Kuc, Tae-Yong | Sungkyunkwan University |
Keywords: Robotic Applications, Artificial Intelligence Systems, Robot Vision
Abstract: Pre-execution reasoning about whether a robot can successfully carry out a natural-language command remains a key challenge in language-guided service robotics. Existing approaches predominantly rely on geometry-based grasp planning while overlooking semantic object properties such as weight, movability, and gripper compatibility, often resulting in repeated attempts to execute physically infeasible actions. This paper presents a reasoning framework that integrates 3D Scene Graph construction with Object Ontology Extraction to assess the feasibility of commanded actions before execution. Using RGB-D scans, the framework automatically constructs per-object ontological representations across three layers—Symbolic, Explicit, and Implicit—providing the semantic grounding required for feasibility reasoning beyond geometric appearance. Lift feasibility is estimated by combining a category-level knowledge base with commonsense inference from a large language model; grasp feasibility is determined from point-cloud geometry relative to gripper constraints; movability is inferred from an implicit ontological attribute; and reachability is verified using the manipulator’s inverse kinematics solver. The extracted attributes and spatial relationships are organized into a 3D Scene Graph, enabling language-conditioned grounding of target objects. A feasibility function, F(o), jointly evaluates these four predicates to determine whether the commanded action should be executed. When any predicate is not satisfied, the framework returns interpretable feedback to the planning layer. Experiments conducted in Isaac Sim using a wheeled mobile manipulator equipped with a UR5 arm demonstrate that the proposed reasoning layer reduces unnecessary execution attempts from 37.5% to 10.0% and improves the task success rate from 62.5% to 85.0% compared with a geometry-only baseline.
|
| |
| 16:10-17:10, Paper FrPo7P.5 | |
| Real-Time Evaluation and Few-Shot Hardware Adaptation of a Vision-Based Tactile Sensor |
|
| Park, Juhee | KAIST(Korea Advanced Institute of Science and Technology) |
| Min, Seojung | KAIST |
| Kim, Jung | KAIST |
Keywords: Robotic Applications, Sensors and Signal Processing
Abstract: Vision-based tactile sensing is increasingly judged by offline test-set accuracy, yet the quantity that matters for robotics is accuracy under live contact on the deployed hardware. Reported accuracy can fall under real-time robotic contact and degrade further when the gel pad or sensor unit is replaced. We quantify both effects on a DIGIT sensor mounted on a 7-DOF Kinova Gen2 arm classifying five sandpaper grit levels (P100–P400). A model that reaches 96.76% on a static test set drops to 88.2% in real time on the same hardware, and falls to 78.0% on a previously unseen sensor-gel configuration. Swapping the gel pad alone accounts for most of the variation we observe across configurations seen during training. Fine-tuning on as few as 14 images per class recovers the hardware-induced gap, reaching a peak of 94.8% at 70 images per class. Offline accuracy is therefore an unreliable indicator of deployed performance, but the cost of adapting to a new hardware configuration can be small.
|
| |
| 16:10-17:10, Paper FrPo7P.6 | |
| Development of a Welding Torch Collision Avoidance Algorithm for Curved Weld Seam Tracking on 3D Pre-Deformed Aluminum Ship Structural Members |
|
| Park, In-Gyu | KIRO, Korea Institute of Robot and Convergence |
Keywords: Robotic Applications, Industrial Applications of Control, Control Devices and Instruments
Abstract: This paper proposes a welding torch collision avoidance algorithm for aluminum ship structural members with 3-dimensional curved lines. In the target CTV fabrication process, the aluminum member is pre-deformed to compensate for predicted welding distortion, and jig clamps are installed to maintain the deformed shape. These clamps can become obstacles during robotic welding. To address this problem, a collaborative robot-based welding system using the HCR14 robot was modeled, and a torch path generation method was developed. The curved weld seam was generalized using a cubic polynomial fitting algorithm based on the least squares method. The fitting accuracy was evaluated using RMSE, and the errors in the x, y, and z directions were 0.5656 mm, 0.2395 mm, and 7.9964×10−11 mm, respectively. Torch posture was controlled using Euler angles to follow the curvature-center direction while satisfying the work angle and travel angle. In addition, an artificial potential field method was applied to avoid collision by correcting the pitch angle corresponding to the work angle. The proposed algorithm is formulated in three-dimensional coordinates, however, its basic behavior was verified through a DAFUL-based simulation using a two-dimensional curved path in the x-y plane without height variation.
|
| |
| 16:10-17:10, Paper FrPo7P.7 | |
| Fieldscale 2.0: Spatiotemporal Tail-Adaptive Rescaling for Thermal Infrared Images |
|
| Kim, Eunseo | Seoul National University |
| Gil, Hyeonjae | SNU |
| Rhee, Tai Hyoung | Seoul National University |
| Kim, Ayoung | Seoul National University |
| Lee, Dongjae | Seoul National University |
Keywords: Sensors and Signal Processing, Robot Vision
Abstract: Thermal infrared (TIR) imaging plays a vital role acrossdiverse illumination conditions, yet its native 14-or 16-bit format requires compression to standard 8-bit representations for visualization and downstream processing. While conventional linear and nonlinear tone-mapping approaches often cause severe darkening or detail loss, recent field-based rescaling methods offer better spatial adaptivity. However, these frameworks still suffer from residual artifacts, saturation, and temporal flickering. To resolve these limitations, we present Fieldscale 2.0, a spatiotemporal adaptive rescaling framework. Specifically, it integrates a three-knot field model to regulate tail distributions, a distribution-aware suppression module to prevent information loss from over-truncation, and an exponential moving average (EMA)-based temporal smoothing scheme for temporal stability. To validate its effectiveness, we conduct experiments across multiple thermal datasets and show that Fieldscale 2.0 achieves robust image and rescaling quality, improves temporal consistency in video sequences, and delivers better performance on a downstream task.
|
| |
| 16:10-17:10, Paper FrPo7P.8 | |
| Virtual Lens Array Synthesis from Multi-View Images Via IRIM |
|
| Lee, Hyeongju | Kyushu Institute of Technology |
| Cho, Myungjin | Hankyong National University |
| Lee, Min-Chul | Kyushu Institute of Technology |
Keywords: Sensors and Signal Processing
Abstract: Integral imaging acquires the 4D light field through a physical lens array, which is difficult to fabricate and subject to optical errors such as vignetting and inter-lens crosstalk. In contrast, lens-array displays for integral imaging are commercially available with diverse specifications. This work introduces a deterministic framework that bridges this asymmetry: iRIM, the inverse of Ray-based Inverse Mapping (RIM), synthesizes the elemental image (EI) that any virtual lens array---across diverse (N_L, N_P, l, f) specifications---would have captured from a single multi-view input. The synthesis follows from the bijection established by the forward RIM and is therefore free of quantization error and learned components. We parameterize the virtual lens-array configuration using the physical display geometry, aligning the synthesized EI with the display side for direct 3D visualization. On a Blender-simulated multi-view scene, iRIM produces uniformly clean sub-aperture images (SAIs) across the entire 6X6 angular grid, while a simulated Lytro-style lens-array camera baseline on the same input retains as little as 6.7% effective coverage at the corner SAIs, quantifying the practical gap closed by deterministic synthesis.
|
| |
| 16:10-17:10, Paper FrPo7P.9 | |
| Algebraic Elliptical Extent Estimation for Robust Extended Object Tracking from Partial Contour Observations |
|
| Lee, In Ho | Korea Institute of Industrial Techology |
Keywords: Sensors and Signal Processing, Robotic Applications, Process Control Systems
Abstract: This paper presents a robust extended object tracking (EOT) algorithm that estimates the full elliptical extent of a target from partial contour observations. In practical sensing geometries only the sensor-facing side of the object is observable, so the conventional assumption of measurements distributed over the entire surface no longer holds. The proposed framework models the sensor--object geometry to identify the visible contour, applies an affine-invariant whitening transformation to the measurements, and solves the ellipse-fitting problem in closed form via singular value decomposition, with an explicit ellipse-validity check and an ellipse-specific fallback fit; the recovered parameters are assimilated as linear pseudo-measurements by a standard Kalman filter. On a synthetic maneuvering trajectory the method attains a Gaussian Wasserstein (GW) error of 9.15 m at 0.94 ms per scan, against 64.84--70.62 m for RMM, MEM-EKF*, and PAKF. On a real LiDAR car-tracking sequence of 49 scans with 58--261 points per scan, in which only the right and rear facets are visible (visibility ratio 47--48%), the method attains mean semi-axis errors of 0.16 m and 0.08 m and an orientation error of 2.7 deg against the vehicle dimensions, outperforming the baselines. A sensitivity study delineates the applicable scope: single-scan reconstruction remains reliable for visibility ratios down to approximately 0.3.
|
| |
| 16:10-17:10, Paper FrPo7P.10 | |
| Roll-Rotated Imaging-Sonar Scanning for Elevation-Ambiguity-Free Mapping of Steep Underwater Terrain for Underwater Inspection |
|
| Ku, Bonchul | Pohang University of Science and Technology |
| Kim, Jason | HEROLab (in Univ. POSTECH) |
| Kim, Seungmin | Pohang University of Science and Technology |
| Song, Young-woon | Pohang University of Science and Technology (POSTECH) |
| Yu, Son-Cheol | Pohang University of Science and Technology (POSTECH) |
Keywords: Sensors and Signal Processing, Robotic Applications
Abstract: Underwater 3D mapping with a forward-looking imaging sonar suffers from elevation ambiguity. Each acoustic return gives range and bearing but not elevation angle, so on steeply sloped terrain returns are placed at the wrong height and form a false slope. Prior methods remove this ambiguity by adding a second sonar, which requires an extra sensor and data association between the two devices. This paper presents an active mapping method that uses a single imaging sonar and rotates it about its roll axis on demand. While an occupancy map is built online, steep regions likely to be distorted by the ambiguity are detected and re-scanned with the sonar rolled ninety degrees about its look axis, so that height is resolved by the well-measured bearing angle. The rolled observations then carve the false slope from the map through a probabilistic negative-only update. In a Stonefish simulation, the method maps steep terrain more accurately than fixed tilts of sixty and ninety degrees, lowering the RMSE against the ground truth by about 14% and the number of cells with height error above 0.5 m by about 57% over the better baseline.
|
| |
| 16:10-17:10, Paper FrPo7P.11 | |
| What Can Lines Tell Us? Parking-Line-Derived Orientation Constraints for Parking-Lot Mapping |
|
| Hyeon, Sangmi | Jeonbuk National University |
| Jeong, Sunghwan | Korea Electronics Technology Institute |
| Choi, Kyoungho | Ekonexon |
| Jo, HyungGi | Jeonbuk National University |
Keywords: Sensors and Signal Processing, Autonomous Vehicle Systems, Navigation, Guidance and Control
Abstract: LiDAR-inertial odometry (LIO) in ground-dominant parking lots can suffer from biased road-surface correspondences caused by parked vehicles, curbs, and partially occluded ground returns. We exploit intensity-selected parking-line points to estimate a local road-surface plane and incorporate its normal into an iterated error-state Kalman filter (IESKF) as an orientation-only pseudo-measurement. The same plane also defines a one-sided map-insertion gate that suppresses below-plane outliers. On a real parking-lot sequence collected by an EV charging robot, the proposed method achieved the lowest Mean Map Entropy (MME) and the smallest final-revisit vertical displacement magnitude among the evaluated methods. The complete pipeline required 51.11 ms per scan on an NVIDIA Jetson AGX Orin.
|
| |
| 16:10-17:10, Paper FrPo7P.12 | |
| State of Charge Estimation under Temperature Mismatch Via Voltage-Domain Adaptation and SOC-Domain Residual Learning |
|
| Cheon, Kibum | Chungbuk National University |
| Haejun, Kim | Chungbuk National University |
| Shin, Jongho | Chungbuk National University |
Keywords: Sensors and Signal Processing, Control Theory and Applications, Artificial Intelligence Systems
Abstract: State of charge (SOC) estimation based on equivalent circuit models (ECMs) and Kalman filters is widely used in battery management systems. However, its accuracy degrades when the operating temperature differs from that of the identified model. Because constructing a complete temperature-specific library of open-circuit voltage (OCV)–SOC tables and ECM parameters is often impractical, a nominal model identified at a reference temperature must be reused at other operating temperatures. This paper presents an online SOC estimation framework that addresses such temperature mismatch through sequential compensation. Starting from a nominal extended Kalman filter (EKF), the ohmic resistance is first updated online using recursive least squares (RLS) to reduce voltage-domain mismatch. The remaining SOC-domain residual is then corrected by a lightweight neural network (NN), and a low-pass filter stabilizes the correction. On the CALCE INR 18650-20R dataset, applying a 25°C nominal model to 45°C dynamic discharge profiles, the proposed estimator reduces the SOC RMSE from 0.2911% for the 25°C fixed EKF to 0.0740%. The estimator also remains suitable for real-time operation.
|
| |
| 16:10-17:10, Paper FrPo7P.13 | |
| Weakly Supervised rPPG Learning Using ECG-To-PPG Signal Translation Diffusion Model |
|
| Oh, Seungmin | Jeonbuk National University |
| Choi, Jiho | Jeonbuk National University |
| Lee, Sang Jun | Jeonbuk National University |
Keywords: Sensors and Signal Processing, Biomedical Instruments and Systems, Artificial Intelligence Systems
Abstract: Remote photoplethysmography (rPPG) estimates physiological signals from facial videos without contact sensors. Supervised rPPG methods require reliable photoplethysmography (PPG) labels, but PPG signals measured in driving environments can be distorted by motion artifacts, illumination changes, vehicle vibration, and contact pressure variation. These noisy labels can degrade rPPG training by providing inaccurate pulse patterns. To address this problem, this paper proposes a weakly supervised rPPG learning framework using ECG-to-PPG signal translation. The proposed framework generates pseudo PPG signals from ECG signals and uses them as supervision signals for rPPG model training. For pseudo PPG generation, this paper uses Physiological Frequency-Consistency Region-Disentangled Diffusion Model (PFC-RDDM), which incorporates PPG peak- and trough-based ROIs, ECG R-peak condition information, and physiological loss functions. Experimental results show that PFC-RDDM improves ECG-to-PPG generation performance and that generated pseudo PPG labels can be used for rPPG training on ECG-only driving data.
|
| |
| 16:10-17:10, Paper FrPo7P.14 | |
| Information-Aware Source Selection for Open-Loop Drone Response Prediction |
|
| Park, Hyeongjun | Hanyang University |
| Kang, Chang Mook | Hanyang University |
Keywords: Sensors and Signal Processing, Artificial Intelligence Systems, Robotic Applications
Abstract: This extended abstract examines source-selected response prediction for an unmanned aerial vehicle (UAV) simulated in PX4 software-in-the-loop (SITL). Flight logs are reorganized into transition samples containing velocity and angular-rate states, offboard command setpoints, tracking-error terms, and finite-difference derivative labels. A Ridge predictor is trained from source excitation logs after candidate source windows are ranked by a composite score reflecting trajectory diversity, tracking mismatch, response magnitude, and command variation. Prediction quality is assessed by recursively propagating states in open loop, rather than by relying only on one-step derivative fitting. The experiments show that source-selected Ridge is most effective on aggressive velocity-reversal maneuvers, while persistence remains competitive for smoother or mixed trajectories. Controlled checks with multilayer perceptron (MLP) base and residual models further indicate that improved local fitting does not automatically produce stable multi-step prediction. These results emphasize rollout-oriented validation and maneuver-dependent source selection for compact UAV response models.
|
| |
| 16:10-17:10, Paper FrPo7P.15 | |
| LiDAR-Guided Geometric Filtering for Noise Removal in W-Band Radar Odometry |
|
| Lee, Dongje | SEADRONIX |
| Kim, Hanguen | Seadronix Corp |
| Jang, Hyesu | Seadronix |
Keywords: Sensors and Signal Processing, Navigation, Guidance and Control, Robotic Applications
Abstract: W-band radar is increasingly adopted for perception in robotics, with its robustness under adverse weather and its extended range relative to camera and LiDAR. Prior radar approaches rely on intensity-based detection, assuming noise returns are weak. In practice, radar cannot separate spurious signals since noise frequently matches the intensity of real structures. Thus, we propose a heterogeneous sensor fusion method that uses geometric line features extracted from LiDAR. Line features serve as a spatial constraint, retaining only radar returns originating from physical structures. Across maritime and urban ground datasets, we filtered out the noise and verified that the discarded noise is nearly indistinguishable from the retained returns in intensity. Supplying the refined returns to state-of-the-art radar odometry algorithms reduces the absolute pose error (APE) of every baseline.
|
| |
| 16:10-17:10, Paper FrPo7P.16 | |
| Detecting Fabric Deformation Using Piezoelectric Soft Sensors |
|
| Lee, TaeKyoung | Korea University |
| Cha, Youngsu | Korea University |
Keywords: Sensors and Signal Processing
Abstract: In this paper, we propose a sensing system for detection of fabric deformation using multiple tendon-inspired sensors. Specifically, the piezoelectric sensors in the sensing system are positioned on the corner of a fabric in a fan-shaped configuration. Additionally, sewing threads are connected to the sensors and are tied to the edges of the sensing area, transferring tensile force to the sensors. Furthermore, a depth camera is installed to calculate the ground-truth curvature of the fabric. A series of experiments is conducted with various experimental conditions, including the shape of supporting plates and pulling distances, to evaluate the relationship between the sensor output and the corresponding curvatures. From these results, the curvature of the fabric is found to be estimable using the relationship between the voltage and the curvature. Furthermore, a virtual simulation of fabric deformation is conducted to visualize the real-time change in fabric surface, demonstrating the feasibility of the sensing system. Moreover, we discuss the theoretical expectation for the tendon-inspired sensors.
|
| |
| 16:10-17:10, Paper FrPo7P.17 | |
| Accuracy Enhancement of Finite-Time Spectral Observers Via Exact Partial Discretization for Periodic Signal Estimation |
|
| Hong, Woosuk | Tokyo University of Science |
| Murakami, Madoka | Tokyo University of Science |
| Nakamura, Hisakazu | Tokyo University of Science |
|
|
| |
| 16:10-17:10, Paper FrPo7P.18 | |
| GEONJI-3D: A Multimodal Dataset for 3D Perception in Unstructured Outdoor Environments |
|
| Shin, Jaeho | Jeonbuk National University |
| Ha, Jiwon | Jeonbuk National University |
| Hwang, JunHyeon | Jeonbuk National University |
| Park, Jaebyung | Jeonbuk National University |
Keywords: Sensors and Signal Processing, Robotic Applications, Robot Vision
Abstract: Autonomous robots operating in unstructured outdoor environments require robust perception capabilities to safely navigate complex terrain with slopes, vegetation, rocks, tree branches, and dynamic obstacles. Although several datasets have been proposed for autonomous driving and off-road perception, multimodal datasets collected using real ground robot platforms in forest-like unstructured environments remain limited. In this paper, we introduce GEONJI-3D, a multimodal dataset for 3D perception and traversability estimation in unstructured outdoor environments. The dataset was collected using a tracked ground robot platform equipped with a stereo RGB-D camera, a 3D LiDAR sensor, and an IMU at the Geonji Mountain Research Forest of Jeonbuk National University. GEONJI-3D consists of two driving scenarios that include well-maintained walking trails, narrow and rough paths, revisited section, vegetation, tree branches, and dynamic obstacles. The dataset provides synchronized RGB images, depth images, LiDAR point clouds, and IMU measurements in the ROS 2 bag format. By providing real-world multimodal sensor data from challenging outdoor environments, GEONJI-3D can support research on sensor fusion, terrain perception, traversability estimation, and robust autonomous navigation for ground robots.
|
| |
| 16:10-17:10, Paper FrPo7P.19 | |
| Seat Occupancy Localization Using Differential Range-Angle Energy Distribution in mmWave Radar |
|
| Kim, Sumin | Changwon National Unversity |
| Park, Minjoo | Changwon National Univesrity |
| Ryu, Haeun | Changwon National University |
| Gim, Juhui | Changwon National University |
Keywords: Sensors and Signal Processing, Autonomous Vehicle Systems
Abstract: This paper proposes a training-free seat occupancy localization method using mmWave radar based on seat-specific energy partitioning in range-angle (RA) maps. An empty-cabin RA map is first utilized as a static reference to suppress structural reflections and extract occupant-induced energy variations. The resulting differential RA-map energy is then partitioned into seat-specific regions using range- and angle-based boundary estimation with sigmoid weighting functions to accommodate boundary uncertainty and occupant motion. Occupancy localization is subsequently performed using normalized energy scores computed within each seat region. Experimental validation in a four-seat cabin mock-up environment demonstrates that the proposed framework can reliably localize seat occupancy under various occupant configurations while maintaining lower computational complexity than conventional radar-processing approaches. The results indicate that the proposed method provides an effective and computationally efficient solution for radar-based seat occupancy localization without requiring point-cloud generation, model training, or retraining procedures.
|
| |
| 16:10-17:10, Paper FrPo7P.20 | |
| SONA: Scanning Once for N Simulation-Ready Assets from 2D Gaussian Splatting for Robotic Manipulation |
|
| Giri, Na | Korea Electronics Technology Institute |
| Kim, Beomjoon | Korea Electronics Technology Institute (KETI) |
| Oh, Saemyung | Korea Electronics Technology Institute |
Keywords: Robotic Applications, Robot Vision, Industrial Applications of Control
Abstract: Scalable simulation of real environments is increasingly central to robot learning, as policies are routinely trained and validated in digital twins (DT) before deployment on hardware. Learning manipulation policies in DT requires per-object simulator assets, which are conventionally modeled, textured, and assigned collision geometry and physical parameters by hand, and rebuilt whenever the scene changes. Asset preparation is thus a primary bottleneck in DT construction. Gaussian Splatting (GS) reconstructs photorealistic scenes from multi-view images, but its output is neither segmented into objects nor endowed with physical properties. This work addresses the scene-to-asset conversion as a distinct problem and presents SONA, a pipeline that converts a single GS reconstruction of a work cell into N per-object, simulation-ready assets. Each object is separated by a training-free vertex-labeling scheme that back-projects video segmentation masks onto a full-scene Truncated Signed Distance Field (TSDF) mesh, paired with a watertight visual mesh and a convex collision mesh in a common coordinate frame, and exported in the USD and URDF formats. On an industrial work cell, SONA reduces asset-generation time by up to 74% at N=6 relative to per-object scanning, and the exported assets reach a 94.2% pick success rate when used for policy learning in Isaac Lab without manual format conversion or physics-parameter calibration. On the one object captured in both scan modes, with identical policy initialization and training budget per asset, the scene-scan asset stays within 1.1% of the dedicated object-scan asset at a one-second hold criterion.
|
| |
| 16:10-17:10, Paper FrPo7P.21 | |
| Robot Tool Path Planning for Automated Peeling of Frozen Tuna |
|
| Jeong, Woong | Korea Institute of Industrial Technology, Korea University |
| Ahn, Woo Jin | Inha University |
| Shin, Myeongchan | KITECH |
| Lim, Myo-Taeg | Korea University |
| Nam, Yun Seok | Tech University of Korea |
| Pyo, Dongbum | Korea Institute of Industrial Technology |
| Lee, Kwang Hee | Korea Institute of Industrial Technology |
Keywords: Robotic Applications
Abstract: Bluefin tuna is commonly processed into high-value products, and therefore maintaining edible yield during processing is essential. Since the skin is directly attached to the edible flesh, skin removal critically determines the final yield. Nevertheless, the skin is still removed manually with a hand-held planar grinder, making the process labor-intensive and inconsistent in material removal. Achieving consistent material removal therefore requires robotic machining with quantitatively controlled tool paths. The proposed method combines specimen-specific surface path generation, curvature-adaptive iso-scallop spacing, surface-normal tool alignment, and collision-aware path trimming. The spacing between adjacent tool paths is determined directly from the local surface curvature and a prescribed scallop height, eliminating the need for iterative tuning of a specimen-specific global spacing. In NVIDIA Isaac Sim, the proposed method increased the peeling ratio from 92.97% with the equal-spacing baseline to 98.66%, while edible flesh loss increased from 1.96% to 3.04% and processing time increased from 55.38,s to 95.79,s.
|
| |
| 16:10-17:10, Paper FrPo7P.22 | |
| Integrated Mobile Snow-Removal Robot Platform Using FAST-LIO2-Based Localization and Web-Based Control |
|
| Park, BoJeong | KIRO |
| Park, Chanill | Korea Institute of Robotics & Technology Convergence (KIRO) |
| Jung, Eui-Jung | Korea Institute of Robot and Convergence |
Keywords: Robotic Applications, Navigation, Guidance and Control, Autonomous Vehicle Systems
Abstract: Current snow-removal operations are performed mainly on roads, whereas narrow sidewalks and alleys are difficult to access with large snow-removal equipment and remain highly dependent on human labor. In addition, snow-removal operations during the night and early-morning hours increase operator fatigue and safety risks. To address these issues, this paper presents a mobile snow-removal robot system using FAST-LIO2-based localization and web-based control. The proposed system integrates LiDAR-IMU-based mapping, PCD-map-based localization, traversability-costmap-based Nav2 navigation, web-based waypoint assignment, and brush and brine equipment control into a ROS2-based platform. In particular, a two-stage localization structure is applied by combining global PCD-map-based initial localization with FAST-LIO2-based local tracking, thereby considering both initial-pose uncertainty and computational load during operation. Experimental results show that the proposed switching structure reduced the computational load by 58.9% compared with the case in which global PCD-map-based localization was continuously executed without switching. In addition, the integrated operation of mapping, localization, waypoint navigation, web-based operation, and equipment command execution was verified. This study focuses on validating the operational feasibility and integrated system architecture of a mobile robot for snow-removal tasks in confined outdoor environments, rather than evaluating quantitative snow-clearing performance under actual snow-covered conditions.
|
| |
| 16:10-17:10, Paper FrPo7P.23 | |
| A Multimodal Integrated Data Synchronization and Acquisition System Based on Heterogeneous Interfaces for Driver Monitoring System Evaluation |
|
| Son, Huigyeong | Korea Automotive Technology Institute |
| Lee, Hun | Korea Automotive Technology Institute |
| Oh, Young-dal | KATECH |
| Ryu, DongWoon | Korea Automotive Technology Institute |
| Park, Sunhong | Korea Automotive Technology Institute |
Keywords: Sensors and Signal Processing, Human-Robot Interaction, Biomedical Instruments and Systems
Abstract: This paper proposes a multimodal data synchronization and acquisition system for the training and evaluation of driver monitoring systems (DMS). The proposed system receives multi-channel camera video from an integrated ECU over RTSP (Real-Time Streaming Protocol) and applies a channel-wise, thread-based concurrent storage structure designed to minimize frame drops. In addition, diverse heterogeneous data such as EEG, biosignals, seat pressure, vehicle CAN, and gaze data are synchronously acquired in a temporally aligned form through trigger and communication interfaces. Through a pilot test conducted in a real-vehicle environment, the stable synchronized acquisition performance of multi-channel video and sensor data was confirmed, and it was verified that the proposed system can be used as a general-purpose data acquisition platform for various DMS studies.
|
| |
| 16:10-17:10, Paper FrPo7P.24 | |
| RGB Image Reconstruction Via Radar-IMU Fusion for Human Localization and Pose Recognition |
|
| Kim, Seungyeon | Sungshin Women's University |
| Yoo, Jaehyun | Sungshin Women's University |
Keywords: Sensors and Signal Processing, Artificial Intelligence Systems
Abstract: The importance of Human Activity Recognition (HAR) is increasing in applications such as indoor monitoring, smart homes, and home care. Frequency-Modulated ContinuousWave (FMCW) radar is widely used as an indoor sensing method because it provides range and velocity information while minimizing privacy intrusion. However, a radar-only method is limited in accurately reconstructing human localization and pose when reflected signals are weak. To address this issue, this paper proposes a FiLM-based radar-IMU fusion model that uses range-Doppler map sequences and wristworn IMU feature sequences. The proposed model generates FiLM parameters from IMU features and uses them to modulate radar feature maps. This allows motion and pose information to be incorporated into the radar representation, enabling RGB image reconstruction of human localization and pose. The experimental setup considers two dynamic activities, walking and crouching walking, over a distance range of 2–7 m, and two static poses, standing and sitting, at distances of 2 m and 7 m. Experimental results show that the proposed FiLM-based fusion method provides more stable RGB-space reconstruction for human localization and pose recognition than a radar-only method.
|
| |
| 16:10-17:10, Paper FrPo7P.25 | |
| FreeRTOS-Based Particle Filter Implementation for Real-Time Dynamic Target Tracking on Resource-Constrained MCUs |
|
| Kim, Chaeyoung | Kumoh National Institute of Technology |
| Ryu, Seungha | Kumoh National Institute of Technology |
| Lee, Heoncheol | Kumoh National Institute of Technology |
Keywords: Robotic Applications, Robot Mechanism and Control, Artificial Intelligence Systems
Abstract: This paper investigates the real-time execution characteristics of particle filtering for dynamic target tracking on a resource-constrained embedded platform and presents a FreeRTOS-based stage-level execution architecture for coordinating computationally intensive particle-filter processing with time-critical control execution. Unlike single-task RTOS execution, the proposed architecture organizes the major particle-filter computations and the subsequent estimate computation into separate binary-semaphore-synchronized tasks while preserving their sequential data dependencies. The estimate computation derives the representative target state from the particle set for use by the control process and is implemented as a separate computational step following resampling. A fixed-priority configuration is further applied to differentiate control-related processing within the stage-level architecture. The proposed approach is evaluated on an ARM Cortex-M7-based STM32H743I MCU operating at 480 MHz through execution-time profiling, motor-control jitter analysis, and deadline-feasibility experiments. Four configurations are compared: Non-RTOS sequential execution, single-task RTOS execution, equal-priority stage-level RTOS execution, and the proposed priority-based stage-level RTOS execution. The proposed configuration exhibits the lowest motor-control jitter, demonstrating improved temporal isolation between PF computation and control processing. Under a 120-ms deadline, the empirical feasibility boundary occurs between 1186 and 1187 particles; 1186 particles satisfy all 1,000 measured cycles, whereas 1187 particles produce a 20% deadline-miss rate.
|
| |
| 16:10-17:10, Paper FrPo7P.26 | |
| Comparison of SAM-Based Models for Tomato Stem Segmentation in Robotic Harvesting |
|
| Park, Jeongmo | Jeonbuk National University |
| Seo, YongSeong | Jeonbuk National University |
| Park, Jaebyung | Jeonbuk National University |
Keywords: Robot Vision
Abstract: Accurate tomato stem segmentation is essential for determining cutting positions in robotic harvesting systems. Object detection can be used to localize target regions, whereas image segmentation provides pixel-level shape information for representing elongated stem structures. Several SAM-based models have recently been introduced for promptable segmentation, but their relative performance on tomato stems has not been sufficiently examined. This study compares four SAM-based models—MobileSAM, SAM, SAM2, and SAM3—using the same box prompts obtained from YOLOv8m-seg predictions. Quantitative and qualitative evaluations assess segmentation accuracy, boundary preservation, and inference time. SAM achieves the highest IoU and Dice scores, while SAM2 provides similar segmentation performance with a shorter inference time. MobileSAM shows greater boundary deviations in some cases, whereas SAM3 produces false-positive regions and misses substantial portions of the stem, resulting in lower segmentation performance. These results show that the performance of ROI-guided stem segmentation varies across SAM-based models and that model selection is important for robotic harvesting applications.
|
| |
| 16:10-17:10, Paper FrPo7P.27 | |
| DL-GPOM: Dual-Layer Gaussian Process Occupancy Mapping for Mixed-Evidence Preservation |
|
| Seo, Jaemin | KAIST |
| Oh, Hyondong | KAIST |
Keywords: Robotic Applications, Autonomous Vehicle Systems, Navigation, Guidance and Control
Abstract: An occupancy map should tell a motion planner where sensor evidence supports free space and where it supports obstacles. Standard occupancy representations combine free and occupied observations into a single posterior at each location. When free and occupied observations overlap and the free ones are more numerous, the posterior reports the space as free and the occupied observations are ignored---a failure that risks collision. Such overlaps become more frequent in open-world environments, where the robot encounters geometries and surfaces outside its sensor's calibration. We address this loss at the representation level. We model free and occupied observations with two independent Gaussian process (GP) layers---one trained only on free observations (mathbf{V}_f), the other only on occupied observations (mathbf{V}_o)---so that per-class evidence remains separate. From (mathbf{V}_f, mathbf{V}_o) we construct an occupancy representation in which minority occupied evidence is no longer absorbed by the free majority. We call this framework dual-layer GP occupancy mapping (DL-GPOM). To keep inference tractable at large scales, DL-GPOM combines Hilbert space GP approximation with patchwise GPU batching. On a public 2D LiDAR mapping benchmark, DL-GPOM produces the most accurate occupancy classification (AUC) and the most reliable occupancy probabilities (Brier score) among compared methods while running in real time on an embedded GPU.
|
| |
| 16:10-17:10, Paper FrPo7P.28 | |
| Multi-Robot Systems As Pedagogical Tools for Active Learning in Manufacturing Process Pedagogy |
|
| Farooq, Muhammad Umar | KAIST |
| Jang, Wonseok | KAIST |
| Suh, Mingyu | KAIST |
| Ban, Sahngjin | Korea Advanced Institute of Science and Technology |
| Jang, Young Jae | Korea Advanced Institute of Science and Technology |
Keywords: Robotic Applications, Autonomous Vehicle Systems
Abstract: Engineering pedagogy is increasingly adopting active learning methodologies to improve learning outcomes in modern classrooms. In some implementations, interdisciplinary modules are integrated to enhance cross-disciplinary knowledge while reinforcing students' understanding of the core course concepts. This work reports a case study of an active learning-based project that employs a multi-robot system in a manufacturing process course. The project was designed using a problem-based learning methodology, and a pseudo-shop-floor production line was developed on a testbed, with material handling coordinated by a multi-robot system. Enrolled students were divided into teams, where they operated the robots, controlled production, and calculated key performance indicators (KPIs) using the manufacturing process knowledge acquired during the course. Survey results indicated that students highly rated the multi-robot system-based project in terms of improving course learning and enabling the practical application of theoretical concepts. Instructor observations during interviews further suggested increased student motivation and improved conceptual understanding. Overall, the case study demonstrates that robots can be effectively integrated as pedagogical tools to bridge the gap between theory and practice in engineering pedagogy, potentially improving learning outcomes and student engagement.
|
| |
| 16:10-17:10, Paper FrPo7P.29 | |
| RUMBLE: Reinforcement Learning-Based Risk-Aware Unified Motor-Load Balancing and Locomotion Adaptation for Thermal Endurance in Quadruped Robots |
|
| Kim, Jangho | Daegu Gyeongbuk Institute of Science and Technology |
| Hong, Jinsong | DGIST |
| Oh, Sehoon | DGIST |
Keywords: Robotic Applications, Artificial Intelligence Systems, Industrial Applications of Control
Abstract: Quadruped robots are increasingly being considered for industrial tasks such as inspection, logistics, and payload transportation. In such tasks, heavy payloads, long operating times, and irregular terrain jointly impose repeated high-torque demands on electrically actuated motors, which can raise motor temperatures close to protection limits. Sudden activation of actuator protection caused by overheating can interrupt the task and may also create safety risks in industrial environments. This paper proposes a temperature-conditioned locomotion modulation framework that adjusts high-level gait and posture commands according to actuator temperature. Instead of uniformly scaling all motor commands or stopping the robot, the proposed Stage 2 module modulates variables such as body pitch, duty factor, and gait frequency to reduce current and torque usage of thermally critical actuators. The method is evaluated in simulation under global high-temperature, single-joint hot, and representative single-leg hot conditions. The results show that mean current and torque RMS decrease under high-temperature states, and that selected hot joints and a representative hot leg use less actuator effort than their all-normal references. These results suggest that the proposed framework has potential for extension to high-load legged-robot tasks such as long-duration payload transportation, industrial inspection, and disaster response.
|
| |
| 16:10-17:10, Paper FrPo7P.30 | |
| DCL-DAE: Dilated CNN and Bi-LSTM Based IMU Denoising Autoencoder for Robust Robot Manipulator Control |
|
| Heo, Ji-yoon | Sejong University |
| Woo, Hyunsoo | Sejong University |
Keywords: Sensors and Signal Processing, Robotic Applications, Artificial Intelligence Systems
Abstract: While IMU sensors are critical for precise robot manipulator control, diverse disturbance environments induce high-frequency, nonlinear noise that compromises trajectory stability. Conventional filtering and standard deep learning models often degrade temporal resolution or incur high computational costs. To address these challenges, we propose DCL-DAE (Dilated CNN-LSTM Denoising Autoencoder), which integrates dilated convolutions with Bi-LSTM network. By adopting dilated convolutions, the encoder expands the receptive field while preserving fine-grained time-series features. Furthermore, the Bi-LSTM latent space captures bidirectional temporal contexts to isolate and filter out diverse disturbance anomalies. Evaluated on the CASPER dataset under five distinct disturbance conditions, DCL-DAE achieved an average trajectory reconstruction improvement of 68.01%, with an error suppression rate of up to 89.38% in the z-axis acceleration. Qualitative validations via PyBullet physics simulation further confirmed that the proposed architecture robustly mitigates end-effector jitter, directly translating numerical denoising gains into smooth and stable physical control.
|
| |
| 16:10-17:10, Paper FrPo7P.31 | |
| Autonomous Soil Monitoring with Vision-Language Guided Robotic Sensor Insertion |
|
| Shin, Chanhee | Chungbuk National University |
| Kim, Jayoung | Chungbuk National University |
Keywords: Robotic Applications, Robot Mechanism and Control, Artificial Intelligence Systems
Abstract: This paper presents an autonomous in-situ soil moisture monitoring system that employs a 6-DoF manipulator mounted on a mobile robot platform. The system uses a vision-language model (VLM) for zero-shot wilted-plant detection and a three-stage vertical approach path that guides a soil moisture sensor into the target pot through sequential approach, insertion, and extraction stages. Conventional unconstrained inverse kinematics (IK)-based motion results in oblique entry trajectories, yielding a 24% collision rate and only 72% task success. By decoupling horizontal positioning from vertical insertion, the proposed three-stage path reduces the collision rate to 4% (an 83.3% reduction) and raises the success rate to 96%. Across 50 repeated trials, the system achieves a position repeatability of 2.4 mm (SD 0.9 mm) and a tilt error within 2.3°. The end-to-end pipeline completes each pot in 65.0 ± 9.2 s, reducing the total number of sensors required by 80%.
|
| |
| 16:10-17:10, Paper FrPo7P.32 | |
| Lightweight Semantic Map Merging Based on Overlapped Semantic Object Detection and Matching for Small Swarm Robot Systems |
|
| Kim, Geunhee | Kumoh National Institute of Technology |
| Lee, Heoncheol | Kumoh National Institute of Technology |
Keywords: Robotic Applications, Sensors and Signal Processing, Robot Vision
Abstract: As the utilization of multi-robot systems in indoor environments continues to expand, the importance of semantic map merging which integrates independent maps generated by individual robots into a comprehensive global map has increasingly gained prominence. However, map merging on small-scale robot platforms faces inherent challenges, including noise from object detection, uncertainties in communication data, and the limitations of constrained computational resources, making it difficult to simultaneously achieve both real-time performance and robustness. To address these limitations, this paper proposes a lightweight semantic map merging framework designed to achieve both computational efficiency and high registration accuracy. Experimental results using small-scale robots equipped with the Jetson Nano platform demonstrate that the proposed framework maintains high registration success rates and real-time processing efficiency, even within significantly constrained computational environments.
|
| |
| 16:10-17:10, Paper FrPo7P.33 | |
| A Unified Framework for Collision-Free Path Planning and Contact-Compliant Safety Control in Surgical Assistant Robots |
|
| Lee, Sanghoon | KAIST |
| Kim, Sungmin | Korea Advanced Institute of Science and Technology (KAIST) |
| Kim, JinYeol | Korea Advanced Institute of Science and Technology |
| Han, Seo Wook | Korean Advanced Institute of Science and Technology |
| Kim, Min Jun | KAIST |
Keywords: Robotic Applications, Biomedical Instruments and Systems, Robot Vision
Abstract: Surgical assistant robots can support surgeons in repetitive and physically constrained tasks, but their deployment in shared surgical workspaces requires safe operation around dynamic obstacles, delicate anatomical structures, and contact interactions. In such settings, robots must generate feasible motions toward surgical objectives while maintaining safe physical interaction under unexpected contact or human intervention. This paper presents a safety-aware planning and control framework that addresses both geometric and contact safety. The framework couples online obstacle perception and collision-aware motion planning with force-bounding low-level control. The planner adapts reference motion under obstacle, task, and robot constraints, while the controller regulates task-space interaction forces under disturbances and model uncertainty. Validation in a realistic surgical-assistance scenario demonstrates that the proposed framework can coordinate online motion adaptation and force-bounding execution in a contact-rich shared workspace.
|
| |
| 16:10-17:10, Paper FrPo7P.34 | |
| Low-Order Fluid Dynamics Estimation Encoder for Robotic Fluid Transport |
|
| Lee, Jun-Woo | Kookmin University |
| Kim, TaeHyun | Kookmin University |
| Kim, Sebeom | Kookmin University |
| Seo, Hyung-Tae | Kookmin University |
Keywords: Robotic Applications, Robot Vision, Artificial Intelligence Systems
Abstract: Robotic liquid transport requires the robot to account for fluid surface responses induced by container motion. Instead of reconstructing the full fluid state or directly estimating raw liquid properties, this paper focuses on estimating low-order response parameters relevant to downstream transport control. We propose a GRU-based Fluid Encoder that maps time-series fluid surface features to the damping ratio zeta and natural frequency omega_n. To introduce actuator-tracking dynamics into the excitation process, a MuJoCo-based virtual container system is used to convert commanded chirp motion into simulated container acceleration, which then drives the low-order sloshing response model. A chirp excitation spanning the target frequency range is adopted to improve parameter identifiability. The proposed encoder is further compared against two analytical frequency-domain baselines, referred to as Blind FRF and Privileged FRF. A control-relevance validation further indicates that the estimated parameters preserve information relevant to downstream motion adaptation.
|
| |
| 16:10-17:10, Paper FrPo7P.35 | |
| Inference-Time Temporal Max Pooling for Robust PMSM Drive-Module Fault Diagnosis under Sensor Relocation |
|
| Youn, Donggyu | UST(University of Science and Technology |
| Jeung, Deokgi | Korea Institute of Machinery and Materials |
| Sin, MinKi | Korea Institute of Machinery & Materials |
| Cho, Jang Ho | Korea Institute of Machinery & Materials |
Keywords: Sensors and Signal Processing, Robotic Applications, Artificial Intelligence Systems
Abstract: Sensor relocation alters vibration transmission paths, degrading vibration-based fault diagnosis of permanent magnet synchronous motor (PMSM) drive modules and increasing false alarms. This study evaluates inference-time temporal post-processing for robust multi-class diagnosis under sensor relocation and physically induced disturbances. A triaxial vibration dataset covering five diagnostic conditions and four sensor locations was used to train a lightweight 1D-CNN–BiLSTM model (20.20 M MACs per segment) at a single source location and evaluate it at three other locations, including two strictly unseen test locations. Raw inference, temporal majority voting (TMV), temporal moving average (TMA), and temporal max pooling (TMP) were compared. At 500 rpm with a representative window of (w=40), TMP achieved 95.28% and 99.85% accuracy with 0.00% false alarm rate at the two unseen locations, demonstrating improved cross-location reliability without target-location training data.
|
| |
| 16:10-17:10, Paper FrPo7P.36 | |
| Physics-Guided Acoustic Anomaly Detection with Environment Noise Modeling for Submarine Rotating Machinery |
|
| Kim, KwangSik | Inha University |
| Kim, Edam | Department of Naval Architecture and Ocean Engineering, Inha University |
| Kim, Yoo-Lim | Naval Ship System R&D Team, Hanwha Ocean Co. Ltd |
| Roh, Young-Ki | Naval Ship System R&D Team, Hanwha Ocean Co. Ltd |
| Lee, Won-Joon | Naval Ship System R&D Team, Hanwha Ocean Co. Ltd |
| Lee, JangHyun | Department of Naval Architecture and Ocean Engineering, Inha University |
Keywords: Sensors and Signal Processing, Artificial Intelligence Systems
Abstract: This study proposes an unsupervised anomaly detection framework for the early detection of abnormal operating conditions in submarine rotating machinery under limited-data environments. Because acoustic signals are contaminated by underwater ambient noise and onboard mechanical noise, robust performance against environmental noise is essential. A physics-based submarine noise model was developed by incorporating colored background noise, structure-borne resonance noise, band-limited auxiliary noise, tonal components, and sensor noise. Realistic noise-augmented datasets were generated under signal-to-noise ratio (SNR) conditions ranging from −6 to 0 dB. To interpret the physical characteristics of the acoustic signals, a structural dynamics model of the pump–mount–support structure was established. Modal analysis identified the natural frequencies and mode shapes, while harmonic response analysis revealed resonance-prone frequency regions and vibration amplification at different rotational speeds. For edge-computing applications, three unsupervised anomaly detection algorithms were evaluated: (i) a Gaussian Mixture Model (GMM) using statistical MFCC features, (ii) an Ensemble Autoencoder using statistical acoustic features, and (iii) a Conv1D-based Ensemble Autoencoder using log Mel-spectrogram sequences. Performance was evaluated using AUC, F1-score, and computational efficiency. The GMM achieved competitive performance with low computational cost, whereas the Conv1D-based model provided higher detection accuracy by exploiting temporal acoustic patterns at the expense of greater computational complexity. These results highlight the importance of selecting an appropriate anomaly detection algorithm by balancing detection performance and computational efficiency for resource-constrained edge-computing environments.
|
| |
| 16:10-17:10, Paper FrPo7P.37 | |
| VR-Based Teleoperation and Data Collection Pipeline for VLA Training on a Dual-Arm Manipulator |
|
| Hong, SuHyun | University of Seoul |
| Hwang, Myun Joong | University of Seoul |
Keywords: Robotic Applications, Artificial Intelligence Systems
Abstract: Vision-Language-Action (VLA) models have recently emerged as a promising approach to learning robot manipulation policies from visual observations and language instructions. However, deploying a VLA model on a specific real-world robot requires a demonstration dataset compatible with the robot's kinematic structure, sensor configuration, and action representation. This paper presents a VR-controller-based teleoperation and data collection pipeline for VLA training on a dual-arm platform. The system supports independent teleoperation of the left and right PiPER manipulators and records multi-view RGB images, dual-arm robot states, actions, and language instructions in the LeRobotDataset v3.0 format. The proposed pipeline was evaluated through SmolVLA fine-tuning and real-robot execution on a controlled bottle pick-and-place task.
|
| |
| 16:10-17:10, Paper FrPo7P.38 | |
| When UAVs Meet IoT: Digital Twin-Based Training System |
|
| Jung, Sungwook | KETI (Korea Electronics Technology Institute) |
| Jung, Won-Seok | Korea Electronics Technology Institute |
| Choi, Sung-Chan | Korea Electronics Technology Institute |
| Sung, Nak-Myoung | Korea Electronics Technology Institute |
| Ahn, Il-Yeop | Korea Electronics Technology Institute |
Keywords: Robotic Applications, Autonomous Vehicle Systems, Control Devices and Instruments
Abstract: This paper presents a digital twin-based training system that integrates physical and software-in-the-loop UAVs with a ground control system and a three-dimensional simulation environment. A digital twin relay engine built on the oneM2M-compliant Mobius platform provides a common middleware layer for UAV telemetry exchange, command delivery, and physical--virtual state synchronization. Platform-specific flight-control interfaces are isolated within the relay layer, allowing the ground control system and simulator to process physical and virtual UAV data through a unified communication structure. The implemented system was evaluated using both SITL-based and physical UAV platforms. The experimental demonstration confirmed end-to-end telemetry delivery, ground control system integration, and the reflection of physical UAV states in the virtual environment.
|
| |
| 16:10-17:10, Paper FrPo7P.39 | |
| Design Progress of an Articulated Robotic Arm for Low-Payload Maintenance Tasks in KSTAR |
|
| Lee, Dohee | Korea Institute of Fusion Energy |
| Kim, Hong-Tack | Korea Institute of Fusion Energy |
| Park, Young Min | Korea Institute of Fusion Energy |
| Hong, Kwon Hee | Korea Institute of Fusion Energy |
| Her, Namil | KFE |
| Choi, Jungsup | SEOULTECH UNIVERSITY |
| Kim, Jinhyun | Seoul National University of Science and Technology |
| Kim, Beom Seok | Seoul National University of Science and Technology |
| Moon, Jeong Whan | KNR System |
| Ryew, Sung Moo | KnR Systems Inc |
Keywords: Robotic Applications
Abstract: Maintenance of fusion experimental devices like KSTAR is challenging due to harsh conditions such as high vacuum and temperatures. Typically, long downtime is required to cool and reduce radiation levels before possible human access. To address this, we designed an articulated robotic arm to perform low-payload maintenance tasks inside the KSTAR device. This paper presents the design progress of an articulated robotic arm, including system configuration, hardware design, and structural analysis, as part of a stage-wise approach to developing the arm. The arm is connected to the shuttle device and stored in a cask, totaling 11 m and 13 DoF. We conduct a finite element method (FEM) analysis to ensure design reliability and safety under target load conditions. We use the proposed robot system to perform essential maintenance operations for KSTAR, including visual inspection, debris removal, and other simple tasks. Our robot arm system has the potential to reduce maintenance preparation time, minimize radiation exposure to personnel, and contribute to improving the fusion experiment’s operational uptime and efficiency
|
| |
| 16:10-17:10, Paper FrPo7P.40 | |
| A Hybrid Robotic Transfer Platform for Manual-Inspection Baggage Logistics |
|
| Jung, Jooik | Incheon International Airport Corporation |
| Cho, Nam-Hyun | Incheon International Airport Corporation |
| Weon, Ihnsik | Korea Institute of Industrial Technology |
Keywords: Robotic Applications, Industrial Applications of Control, Process Control Systems
Abstract: This paper presents a hybrid robotic transfer platform for manual-inspection baggage logistics. In many airport baggage handling systems, baggage selected for manual inspection is still transported by workers between the main baggage handling flow and physically separated inspection areas. To address such recurring operational conditions, this paper reframes the problem as a transferable platform design challenge. It proposes a modular architecture that integrates legacy baggage and process interfaces, centralized orchestration, interoperable autonomous mobile robots (AMRs), conveyor and station handoff modules, vertical transfer interfacing, and safety and traceability functions. The hybrid concept separates repetitive handoff functions from mobile transport functions, thereby reducing technical risk while preserving flexibility across diverse site conditions. The manuscript summarizes the platform concept, layered system architecture, implementation rationale, and verification scenarios for phased deployment in airport environments.
|
| |
| 16:10-17:10, Paper FrPo7P.41 | |
| Human-Guided Reinforcement Learning-Based Manipulation for Multi-Purpose Agricultural Robots |
|
| Kim, Changjo | Chonnam National University |
| Kim, Gangmin | Chonnam National University |
| Choi, Jangju | Chonnam National University |
| Kim, Sumin | Chonnam National University |
| Park, Yonghyun | Gwangju Institute of Science and Technology (GIST) |
| Son, Hyoung Il | Chonnam National University |
Keywords: Robotic Applications, Artificial Intelligence Systems, Robot Mechanism and Control
Abstract: This study proposes a human-guided, reinforcement learning-based manipulation method for dual-arm agricultural robots performing various farming tasks in complex environments. By combining imitation learning and reinforcement learning, the proposed method enables the system to learn efficient and stable manipulation strategies while reducing reliance on manually designed control rules. First, the system learns human behaviors related to farming tasks through behavior cloning (BC), and then refines the learned policy using the proximal policy optimization (PPO) algorithm to enhance robustness and adaptability across diverse agricultural environments. This enables the system to mimic complex dual-arm cooperative behaviors, such as performing a primary task with one arm while removing obstacles with the other, allowing efficient operation even in complex and irregular environments. Through this study, agricultural robots are expected to perform tasks in a human-like manner across complex and diverse agricultural environments.
|
| |
| 16:10-17:10, Paper FrPo7P.42 | |
| Tactile-Based Reinforcement Learning for Visibility Enhancing Branch Pushing in Agricultural Robot |
|
| Kim, Gangmin | Chonnam National University |
| Jo, Yuseung | Chonnam National University |
| Kim, Changjo | Chonnam National University |
| Park, Jiwoon | Chonnam University |
| Ahn, Yuhyun | Chonnam National University |
| Park, Yonghyun | Gwangju Institute of Science and Technology (GIST) |
| Son, Hyoung Il | Chonnam National University |
Keywords: Robotic Applications, Artificial Intelligence Systems, Robot Mechanism and Control
Abstract: Agricultural robots in dense crop environments often suffer from visibility limitations caused by occluding leaves and branches. This study presents a tactile-based reinforcement learning for visibility enhancing branch pushing in agricultural robots. A soft tactile sensor module attached to a three-finger gripper estimates contact cues, including contact center, contact-direction tendency, and moment tendency during interaction with a deformable branch. These cues are used in a reinforcement learning policy trained in a physics-based branch simulation environment. In a 200 mm pushing task, the tactile-based policy achieved higher final visibility, required fewer steps to success, and reduced the end-effector tracking RMSE from 311.82 mm to 43.70 mm compared with the non-tactile condition. The results indicate that tactile-based reinforcement learning improves pushing accuracy during occlusion-clearing manipulation.
|
| |
| 16:10-17:10, Paper FrPo7P.43 | |
| Emergency Vehicle Siren Recognition, Azimuth Estimation, and Distance Estimation Using Multichannel Acoustic Features |
|
| An, Hyogeon | Tech University of Korea |
| Park, Sangjun | Tech University of Korea |
| Park, Seongkeun | Tech University of Korea |
Keywords: Sensors and Signal Processing, Artificial Intelligence Systems, Autonomous Vehicle Systems
Abstract: This paper proposes a model that simultaneously performs event recognition, azimuth estimation, and distance estimation for emergency vehicle siren sounds using multichannel acoustic features. Conventional SELD-based models have been effectively used for sound event detection and direction-of-arrival estimation; however, distance information of the sound source should also be considered to determine the approach status of an emergency vehicle. Therefore, in this study, a Distance Head is added to a CST-Former-based SELD architecture, enabling a single model to jointly predict siren events, azimuth, and distance information. Mel-spectrograms and GCC-based multichannel acoustic features are used as input features, and the channel-, frequency-, and time-axis attention mechanisms of the CST-Former are employed to learn spatial information and time-frequency patterns from acoustic signals. For distance estimation, log-distance regression is applied to reduce the effect of scale differences in distance values, and a distance mask is used so that the loss is computed only for frames with valid distance labels. The applicability of the proposed model is verified by evaluating sound recognition performance, azimuth estimation error, and distance estimation error using multichannel siren data acquired in a real-world environment.
|
| |
| 16:10-17:10, Paper FrPo7P.44 | |
| Fall-Risk Information Report Generation System for On-Site Safety Diagnosis (I) |
|
| Hong, Sung Min | Kunsan National University |
| Kim, Hwa Seok | Kunsan National University |
| Kim, Sun Young | Kunsan National University |
Keywords: Navigation, Guidance and Control, Artificial Intelligence Systems, Sensors and Signal Processing
Abstract: Structural inspection in aging buildings and construction sites often involves hazardous areas where direct human access is difficult. To support safer preliminary inspection, this study proposes a mobile robot-based safety diagnosis system that detects surface damage candidates and fall-risk areas, records their spatial information, and generates structured inspection reports using a Vision-Language Model (VLM). Surface damage candidates are detected from robot-acquired RGB images using YOLOv8n-seg, while fall-risk regions are classified using the proposed RGB-D-based model that exploits both visual appearance and depth cues. Detected events are associated with robot poses and representative 3D points using FAST-LIO2-based mapping information. Experimental results show that the proposed RGB-D model achieves an accuracy of 0.956 and an F1-score of 0.953, outperforming YOLO-based classification models. The VLM-generated reports provide event location, visual evidence, risk type, and required follow-up verification, demonstrating the feasibility of robot-assisted preliminary safety inspection for expert diagnosis in complex and hazardous environments.
|
| |
| 16:10-17:10, Paper FrPo7P.45 | |
| Geometry-Consistent Spatial Danger Mapping for Floor-Opening Fall-Risk Assessment on Construction Sites (I) |
|
| Kim, Hwa Seok | Kunsan National University |
| Hong, Sung Min | Kunsan National University |
| Kim, Sun Young | Kunsan National University |
Keywords: Navigation, Guidance and Control, Artificial Intelligence Systems, Sensors and Signal Processing
Abstract: Floor openings on construction sites can lead to serious fall accidents when they are not recognized and controlled in advance. This study proposes a preliminary safety inspection framework that estimates floor-opening risk from multi-view RGB images and 3D geometric information. The proposed method combines SAM3-based floor segmentation, VGGT-based depth and camera pose estimation, and DINOv3/CLIP-based 3D feature aggregation to construct a geometry-consistent scene representation. Based on this representation, support-deficient regions inside the floor plane are detected as opening candidates, and a protection-aware spatial danger map is generated by considering cover-like and guard-like structures. A rule-based decision module determines the safety status using measured geometric evidence, while a vision-language model summarizes the inspection result with retrieved regulatory evidence. Experiments on real construction-site scenes show that the proposed framework can quantify hazardous floor areas and reduce false alarms compared with image-only VLM assessment.
|
| |
| 16:10-17:10, Paper FrPo7P.46 | |
| Actor–Critic Auxiliary Particle Filtering with a δ-GLMB Label Set for Robust Infrared Multi-Target Tracking (I) |
|
| Kang, Chang Ho | Sejong University |
| Choi, Ji Hun | Sejong University |
| Lee, Dong Hoon | Sejong University |
| Kim, Sun Young | Kunsan National University |
Keywords: Navigation, Guidance and Control, Artificial Intelligence Systems, Sensors and Signal Processing
Abstract: Infrared (IR) small-target multi-target tracking is dominated by missed detections, localization noise, and clutter, conditions under which tracking-by-detection pipelines degrade sharply. We present an Actor–Critic Auxiliary Particle Filter (ACPF) embedded as the single-target density of a full δ-Generalized Labeled Multi-Bernoulli (δ-GLMB) recursion. The Actor draws a measurement-guided two-stage auxiliary proposal, while the Critic adapts the process-noise scale online from normalized-innovation and effective-sample-size statistics, so the particle cloud widens its search exactly when measurements become uninformative. During missed-detection frames an appearance-histogram likelihood, compared by a debiased Sinkhorn optimal-transport divergence, coasts each track through occlusion. The δ-GLMB layer manages target number, labels, and data association with a Gibbs-sampled joint predict–update and Bayes-optimal state extraction. On an 825-run controlled-degradation campaign over IR sequences, a proposed variant attains the best MOTA on each of the four degraded detector fronts (on par with BoT-SORT under occlusion), outperforming controlled re-implementations of strong tracking-by-detection (ByteTrack, BoT-SORT) on a shared measurement stream, and substantially surpassing naive SORT, which collapses under clutter.
|
| |
| 16:10-17:10, Paper FrPo7P.47 | |
| Fault-Aware Fixed-Lag SE(2) Smoothing for Online Learned Residual Correction of Visual Odometry (I) |
|
| Kang, Chang Ho | Sejong University |
| Choi, Ji Hun | Sejong University |
| Yoon, Sung Jin | Sejong University |
| Kim, Sun Young | Kunsan National University |
Keywords: Navigation, Guidance and Control, Artificial Intelligence Systems, Sensors and Signal Processing
Abstract: We present an online learned residual-correction framework for visual odometry (VO) on the KITTI odometry benchmark and study how it can be extended toward bundle-adjustment (BA)-style smoothing without sacrificing its online, fault-aware character. Starting from an Error-State SE(2) particle filter whose correction is produced by an optimal-transport-based neural module (OTPF) acting on the frame-to-frame VO residuals of an ASpanFormer matching front-end, we add a fixed-lag SE(2) smoothing layer together with a fault-detection head and a BLACKOUT frame-skipping mechanism. Through a controlled comparison over 11 KITTI sequences, 3 random seeds, and raw / fault-injected scenarios — spanning a raw-residual baseline, the SE(2)-OTPF online filter, a fixed-lag SE(2) smoother, and four OTPF-coupled smoother variants — we find that a fixed-lag SE(2) residual smoother reduces the position-residual RMSE by up to about 85% under injected faults while preserving rotation, whereas tightly coupling the neural OTPF correction into the optimizer induces a yaw/rotation collapse. At a tuned gate the learned fault detector flags 92% of injected fault events at low false alarm (rising to 97% at a higher-recall gate) with near-zero latency, its only residual gap being the sub-noise onset of slow drift. We therefore position the method not as a replacement for full visual BA but as a fault-aware online residual-correction filter with an optional fixed-lag smoothing layer, and we report the accuracy, runtime (5-21 ms/frame), and fault-detection trade-offs across the design space.
|