| | |
Last updated on July 23, 2026. This conference program is tentative and subject to change
Technical Program for Tuesday August 18, 2026
| |
| TuAT1 |
201 |
| Data Analytics and Optimization for Smart Industry 1 |
Regular Session |
| |
| 10:30-10:50, Paper TuAT1.1 | |
| Dynamic Scheduling Optimization of M/PH/C Queue Based on Q-Learning: A Case Study of Reheating Furnace Control in Steel Industry |
|
| Liu, Aowen | Beijing Information Science and Technology University |
| Jia, Yanhe | Beijing Information Science and Technology University |
Keywords: Agent-Based Systems, Machine learning, Intelligent and Flexible Manufacturing
Abstract: The hot rolling process serves as a critical component in steel manufacturing. The scheduling of reheating furnaces has been influenced by stochastic factors. These influences pose ongoing challenges to production efficiency and cost management. This paper introduces a dynamic adaptive scheduling (DAS) policy grounded in reinforcement learning to mitigate issues related to queue congestion, equipment idleness, and the frequent star-stop operations. These problems arise from the random arrival of orders and variability in service times within the production system. The proposed method models the reheating furnaces scheduling system using an M/PH/C queue model and it is formulated as a Markov decision process (MDP). An optimization framework utilizing Q-learning is presented. This is achieved by formulating a unified reward function that encompasses queuing, equipment downtime, and start-stop transition costs. In the simulation experiment, the DAS policy was contrasted with conventional average, fixed, and double-threshold approaches. The results indicate that the proposed policy significantly improved critical metrics, such as average total cost, equipment switch frequency, queue length, and the number of equipment utilization. Over time, the optimization rate rose from 26.00% to 45.76%. The findings confirm the effectiveness of this approach in the reheating furnace scheduling system.
|
| |
| 10:50-11:10, Paper TuAT1.2 | |
| Toward a Holistic Framework for Data Quality Assessment: A Literature-Driven KPI Model and Unified Scoring Method for Industrial AI |
|
| Melliti, Yosr | Technische Hochschule Augsburg |
| Aybar, Ugur | Predictores.ai GmbH |
| Seebacher, Uwe | Predictores.ai GmbH |
| Legat, Christoph | Technical University of Applied Sciences Augsburg |
Keywords: AI-Based Methods
Abstract: Data-driven artificial intelligence (AI) is increasingly used in industrial environments, where the reliability of machine learning applications depends strongly on the quality of the underlying data. However, data quality is a multidimensional concept, and existing assessment approaches often remain fragmented, relying on isolated indicators without providing a consistent overall judgment of dataset suitability for AI-related use. This paper proposes a user-oriented framework for tabular data quality assessment in industrial AI contexts. The framework is designed as a guided two-step approach. First, it provides an interpretable overall benchmark of dataset quality. Second, it enables a more detailed assessment of individual quality dimensions in order to identify specific quality deficiencies and improvement potential. To support the overall assessment, the paper introduces a consolidated set of intrinsic data quality key performance indicators (KPIs) together with a normalization and aggregation method that renders heterogeneous KPIs comparable and combines them into a unified score. The proposed framework is evaluated on publicly available datasets. The results show that the unified assessment reveals structural weaknesses that remain hidden when only global dataset statistics are considered, while the detailed view supports the identification of dimension-specific quality deficiencies. The framework thereby provides a systematic basis for monitoring, benchmarking, and improving tabular data quality in industrial AI pipelines.
|
| |
| 11:10-11:30, Paper TuAT1.3 | |
| STMixer-P: Spatio-Temporal Mixer Model with Pseudo-Label Learning for High-Precision CIR Localization |
|
| Zhao, Guoliang | Xi'an Jiaotong University |
| Xu, Zhanbo | Xi'an Jiaotong University |
| Liu, Yaping | Xi'an Jiaotong University |
| Wu, Jiang | Xian Jiaotong University |
| Liu, Kun | Xi'an Jiaotong University |
| Guan, Xiaohong | Xi'an Jiaotong University |
Keywords: AI-Based Methods, Big-Data and Data Mining, Machine learning
Abstract: High-precision localization based on Channel Impulse Response (CIR) has recently attracted significant attention due to its low deployment cost and widespread infrastructure. However, existing methods still suffer from limited localization precision and weak cross-scenario generalization, mainly due to the lack of architectures tailored to the intrinsic spatio-temporal characteristics of CIR and the scarcity of labeled data in practical scenarios. To address these issues, we propose STMixer, a spatio-temporal Mixer model. STMixer jointly models the temporal channel characteristics of CIR signals and the spatial distribution of Access Points (APs) through dedicated channel and spatial attention modules together with a Mixer module, enabling more effective CIR feature representation. In addition, we introduce a knowledge-distillation-based pseudo-label learning framework to leverage abundant unlabeled data. The framework adopts an ensemble-based pseudo-label generation strategy and iterative self-training to progressively improve pseudo-label quality, thereby improving localization precision. Extensive experiments demonstrate that STMixer achieves superior localization precision in in-domain scenarios (0.0808 m localization error) and stronger cross-scenario generalization (3.7387 m) compared with existing methods. Furthermore, under limited labeled data conditions, the proposed STMixer-P generates high-quality pseudo labels that effectively expand the training set and reduce the localization error by 89.58%. These results highlight the potential of the proposed approach for practical high-precision localization.
|
| |
| 11:30-11:50, Paper TuAT1.4 | |
| Real-Time Partial Discharge Detection at the Edge Via Hybrid Time-Frequency Sequence Modeling |
|
| Zheng, Yuqi | Hunan University |
| Hong, Lerong | Hunan University |
| Cao, Hongzhuang | Hunan University |
| Fan, Zetong | Hunan University |
| Li, Rui | Hunan University |
Keywords: AI-Based Methods, Failure Detection and Recovery
Abstract: Reliable partial discharge detection is critical for ensuring the insulation integrity of automotive electrical machines. However, deploying highly accurate diagnostic algorithms on resource-constrained edge computing devices remains challenging due to severe industrial background noise. This paper proposes a lightweight hybrid diagnostic architecture that integrates physical signal processing with a data-driven sequence model. First, the stationary wavelet transform is utilized to explicitly decouple transient discharge energy from broadband noise, providing a robust physical prior. Subsequently, a gated recurrent unit network, enhanced by a numerically stabilized temporal attention mechanism, adaptively aggregates the isolated transient features across time steps. Experimental evaluations on a generalized industrial dataset demonstrate that the proposed framework achieves an overall accuracy of 99.78%, surpassing complex two-dimensional deep learning models while reducing the parameter footprint by 97.0%. Furthermore, hardware-in-the-loop tests on an edge microcontroller validate a total end-to-end processing latency of 14.69 milliseconds. The results confirm that the proposed system strictly satisfies the sub-cycle latency requirements of continuous testing environments, effectively achieving high-fidelity electrical diagnosis under severe computational constraints.
|
| |
| 11:50-12:10, Paper TuAT1.5 | |
| Data Dynamic Classification and Grading for Power Systems Based on Correlation Awareness and GRPO |
|
| Tao, Wenwei | Power Dispatching and Control Center of China Southern Power Grid |
| Li, Jialu | Power Dispatching and Control Center of China Southern Power Grid |
| Hou, Namin | Power Dispatching and Control Center of China Southern Power Grid |
| Wang, Hong | China Southern Power Grid Electric Power Research Institute Co., Ltd |
| Tan, Weitao | China Southern Power Grid Electric Power Research Institute Co., Ltd |
Keywords: AI-Based Methods, Learning and Adaptive Systems, Reinforcement
Abstract: Open power-market and operational data can expose sensitive grid states through hidden correlations, enabling inference attacks on topology-related states, congestion patterns, and commercial information. This paper proposes a physically informed, correlation-aware dynamic data classification and grading framework. Data classification assigns each item to a functional category, while grading determines its security level by jointly considering inference risk, availability cost, and compliance constraints. MIC is used to quantify empirical nonlinear dependencies as data-level inference-risk indicators, and an enhanced GRPO algorithm is developed to optimize adaptive grading policies. Experiments on a 39-category PJM market dataset show that the proposed method identifies high-risk features and achieves higher utility than DQN and PPO. The framework offers an adaptive mechanism for mitigating inference risks in evolving power-market data environments.
|
| |
| 12:10-12:30, Paper TuAT1.6 | |
| Enhancing Proprioceptive Reinforcement Learning-Based Quadrupedal Locomotion Controller Via Explicit-Implicit Estimation Network |
|
| Chang, Yu-Cheng | National Taiwan University |
| Ye, Yiru | National Taiwan University (NTU) |
| Lian, Feng-Li | National Taiwan University |
Keywords: Reinforcement, Optimization and Optimal Control, Motion Control
Abstract: Industrial inspection often involves safety risks, making legged robots highly suitable for replacing human labor. However, robust locomotion across diverse, unstructured terrains remains challenging due to estimation uncertainties. In this work, we propose a proprioceptive reinforcement learning (RL)-based locomotion framework enhanced by a novel explicit-implicit estimation network (EI-Net). The explicit network accurately estimates physically meaningful states, while the implicit network leverages a self-supervised learning approach to decorrelate features, achieving superior terrain categorization. Together, this architecture significantly improves dynamic stability. Simulation and real-world experiments show that the robot traverses stairs up to 60% of its body height and single steps up to 0.22 m, achieving an 80% success rate, outperforming other methods. Notably, the framework achieves zero-shot sim-to-real transfer and strong robustness, completing a 2.25km outdoor traversal with a 3 kg payload across unmodeled terrains.
|
| |
| TuAT2 |
202 |
| Frontier Technology of Industrial & Systems Engineering 1 |
Regular Session |
| |
| 10:30-10:50, Paper TuAT2.1 | |
| A MOEA/D-DRA Approach for Multi-Mode Resource-Constrained Aero-Engine Assembly Scheduling |
|
| Du, Huimin | Northeastern University |
| Guo, Qingxin | Northeastern University |
| Dong, Zhiming | Northeastern University |
Keywords: Assembly, Planning, Scheduling and Coordination, Optimization and Optimal Control
Abstract: Aiming at the challenges of complex processes, limited resources, and the need to balance efficiency and cost in current aero-engine assembly, the aero-engine assembly scheduling problem is abstracted as a Multi-mode Resource-Constrained Project Scheduling Problem (MRCPSP). A multi-objective optimization model is constructed considering two conflicting objectives: makespan and resource usage cost. On this basis, the Multi-Objective Evolutionary Algorithm based on Decomposition (MOEA/D) with Dynamic Resource Allocation (DRA) is introduced. The Tchebycheff decomposition method is used to transform the multi-objective problem into a set of single-objective sub-problems for co-evolutionary solution. A three-part encoding scheme including task sequence, mode selection, and resource capacity, along with corresponding genetic operators, is designed. Simulation analysis of an aero-engine component assembly instance shows that this method can effectively obtain a uniformly distributed Pareto front, providing a scientific decision-making basis for assembly shop managers to trade off between schedule and cost, realizing optimal allocation of aero-engine assembly resources.
|
| |
| 10:50-11:10, Paper TuAT2.2 | |
| Continuous-Time Priority-Based Search with Offline Collision Detection and Roadmap Annotation for Efficient Multi-AGV Coordination |
|
| Signorile, Federico | Politecnico Di Bari |
| Carli, Raffaele | Politecnico Di Bari |
| Bertogli, Alessandro | Elettric 80 S.p.A |
| Olmi, Roberto | Elettric80 SpA |
| Koenig, Sven | University of California, Irvine |
| Dotoli, Mariagrazia | Politecnico Di Bari |
Keywords: Autonomous Vehicle Navigation, Motion and Path Planning, Logistics
Abstract: This paper addresses the coordination of automated guided vehicles (AGVs) in industrial warehouse environments, where roadmaps are highly connected and redundant, making scalability and safety critical challenges. We introduce continuous-time priority based search with continuous-time conflicts (CPBS-CTC), a continuous-time extension of priority-based search (PBS) combined with an offline roadmap annotation that enables efficient collision detection. The contribution is two-fold. From a methodological perspective, the paper advances the multi-agent path finding (MAPF) literature by extending PBS to continuous-time domains. From an application perspective, the algorithm is integrated into an industrial-grade simulator that provides a realistic digital twin of warehouse layouts and AGV fleets, paving the way for practical deployment in real-world scenarios. We benchmark CPBS-CTC against state-of-the-art continuous-time MAPF solvers, including continuous-time conflict-based search (CCBS), an annotated CCBS variant, and prioritized planning baselines with and without restarts, on five real-world warehouse layouts. Results show a clear trade-off: CCBS guarantees optimal solution cost but fails to scale, whereas prioritized planning baselines achieve very fast runtimes but become unreliable in dense scenarios. Overall, CPBS-CTC provides the best compromise, achieving the highest success rates on several layouts while maintaining competitive runtime and solution quality.
|
| |
| 11:10-11:30, Paper TuAT2.3 | |
| Beyond Trade-Offs: Mashing Serverful and Serverless Cloud Resources for Budget-Constrained Workflow Execution |
|
| Zhou, Qixin | Chongqing University |
| Sun, Ruyi | Chongqing University |
| Zhou, MengChu | New Jersey Institute of Technology |
| Wu, Quanwang | Chongqing University |
Keywords: Cloud Computing For Automation, Planning, Scheduling and Coordination, Software, Middleware and Programming Environments
Abstract: Serverless computing is characterized by fine-grained billing and elastic scalability, which makes it highly suitable for workflow execution. Nevertheless, it is hindered by cold-start latency and rigid execution constraints. In comparison, conventional serverful cloud resources like virtual machines employ coarser-grained provisioning but offer lower unit computing costs. This paper explores the combination of serverful and serverless resources to leverage their complementary advantages for efficient and cost-effective workflow execution. We design a hybrid resource management framework that dynamically assigns workflow tasks across these two resource types. Furthermore, we present a Budget-constrained Workflow scheduling algorithm for Blended cloud environments (BWB), which aims to minimize makespan under user-defined budget constraints. Experiments are performed using realistic workflow applications in real-world cloud scenarios. By comparing BWB with state-of-the-art serverful, serverless, and hybrid scheduling methods, the results show that BWB reduces the average makespan by at least 16.9%, which effectively validates the cost-effectiveness of the blended cloud resource scheme for workflow execution.
|
| |
| 11:30-11:50, Paper TuAT2.4 | |
| Integrated Server Placement and Task Assignment in Multi-Access Edge Computing Via Column Generation |
|
| Ling, Guangsen | Northeastern University |
| Wang, Gongshu | Northeastern University |
Keywords: Cyber-physical Production Systems and Industry 4.0, Planning, Scheduling and Coordination, Task Planning
Abstract: In multi-access edge computing (MEC) networks, heterogeneous servers, uneven task distributions, and service preferences complicate server placement and task assignment. This paper studies a preference-aware edge server placement and task assignment problem that jointly determines placement locations, server types, and task assignments to minimize total latency and placement cost. A bi-objective MINLP model is formulated, and an ε-constraint-based column generation framework (ε-CG) is developed, where the pricing subproblem is solved by dynamic programming. Numerical results demonstrate the trade-off between service quality and placement cost, showing that the proposed method obtains high-quality Pareto solutions faster than Gurobi and NSGA-II.
|
| |
| 11:50-12:10, Paper TuAT2.5 | |
| Continuous Improvement in Container Manufacturing: A Bottleneck Analysis-Based Simulation Approach |
|
| Wang, Ziwei | Tsinghua University |
| Zhao, Yishen | Tsinghua University |
| Dong, Heng | Tsinghua University |
| Wang, Xiaojun | Shanghai Universal Logistics Technology Co., Ltd |
| Huang, Gang | Shanghai Universal Logistics Technology Co., Ltd |
| Chen, Yunlong | Dong Fang International Container (Qidong) Co., Ltd |
| Chen, Zhen | Dong Fang International Container (Qidong) Co., Ltd |
| Li, Jingshan | City University of Hong Kong |
| Wang, Feifan | Tsinghua University |
Keywords: Factory Automation, Intelligent and Flexible Manufacturing, Cyber-physical Production Systems and Industry 4.0
Abstract: Identification and mitigation of bottlenecks are fundamental for continuous improvement in complex manufacturing systems. In dry-container manufacturing for cargo ships, this task is challenging because production processes are strongly coupled with limited buffering, batch replenishment, synchronized movement, and complex control policies, while onsite trial-and-error is costly. Discrete-event simulation (DES) is widely adopted for predicting system performance and quantifying the effect of potential changes, and the arrow assignment method in production systems engineering (PSE) provides an effective way for bottleneck analysis. In this paper, a DES model using SIMUL8 is proposed to integrate the arrow assignment method in PSE to identify and mitigate bottlenecks for continuous improvements. The model is validated against actual production outputs with high accuracy. Specifically, in a case study of the general assembly and post-processing parts of a real-world container manufacturing line, the proposed model identifies the Welding Chain 1 and the Outer Coat Pre-Coating Booth as two bottlenecks. After their cycle times are reduced by 5%, the bottleneck shifts to the Corner Fitting Removal Station. A further 5% reduction of cycle time at this station increases the mean throughput from 5501.04 to 5533.77 containers over a two-week evaluation period. Such results show that the proposed model can provide a practical tool for bottleneck analysis and continuous improvement in complex container manufacturing systems.
|
| |
| 12:10-12:30, Paper TuAT2.6 | |
| Scheduling Method of Robotic Flexible Flow Shops Based on Petri Nets and Masked-PPO |
|
| Xiaojuan, Yin | Nantong University |
| Bo, Xu | Nantong University |
| Tianbing, Jiang | Nantong University |
| Huixia, Liu | Nantong University |
|
|
| |
| TuAT3 |
203 |
| Systems Technology of Industrial Automation and Control 1 |
Regular Session |
| |
| 10:30-10:50, Paper TuAT3.1 | |
| Performance-Adaptive Active Learning for Quality Control in Smartphone Mainboard Manufacturing |
|
| Yu, Ruizai | University of Chinese Academy of Sciences |
| Shen*, Zhen | Institute of Automation, Chinese Academy of Sciences |
| Xisong, Dong | Institute of Automation, Chinese Academy of Sciences |
| Wu, Huaiyu | State Key Laboratory for Management and Control of Complex Systems, Institute of Automation, Chinese Academy of Sciences |
| Hu, Bin | Institute of Automation, Chinese Academy of Sciences |
| Xiong, Gang | Institute of Automation, Chinese Academy of Sciences |
Keywords: AI-Based Methods, Failure Detection and Recovery, Reinforcement
Abstract: In highly integrated smartphone mainboard manufacturing, defect scarcity, expensive annotation, and appearance variation across production stages severely limit conventional inspection methods. To address these challenges, we propose Enhanced Tri-CEAL, a performance-adaptive active learning framework that integrates three key components: a performance-aware dynamic weighting strategy for adaptive query selection, knowledge-distillation-based collaboration with staged adaptive thresholding to balance pseudo-label quality and quantity, and a dynamically weighted cross-entropy loss that mitigates class imbalance. We evaluate the method on DAGM 2007 and KolektorSDD using sequential active-learning rounds organized as a stage-wise diagnostic protocol on static benchmarks. Enhanced Tri-CEAL achieves 97.8% accuracy on DAGM 2007 and maintains robust defect-recognition performance on KolektorSDD. Compared with full supervision, it reduces annotation cost by 78% while maintaining reliable detection performance.
|
| |
| 10:50-11:10, Paper TuAT3.2 | |
| Riemannian Manifold Optimization Based Magnetic Microrobots Control Subject to Torque Singularity and State Constraints |
|
| Zhang, Junjie | Jiangnan University |
| Liu, Yueyue | Jiangnan University |
| Fan, Qigao | Jiangnan University |
Keywords: Automation at Micro-Nano Scales, Motion Control
Abstract: Magnetic microrobots offer enormous potential for targeted medical interventions and microassembly; however, their operational reliability is frequently compromised by torque singularities and complex environmental constraints. This paper proposes a robust control framework that integrates Riemannian manifold optimization with Control Barrier Functions (CBF) to address this problem. By reparameterizing the state space on a Grassmann manifold and employing an L_1 norm-based cost function, the system provides intrinsic immunity to singular configurations where the magnetic moment aligns with the driving field. Unlike conventional Euclidean methods, this manifold-based approach constrains the Riemannian gradient, effectively preventing numerical divergence and ensuring bounded control efforts even near degenerate states. To bridge the gap between algorithmic optimization and physical execution, Tikhonov regularization is introduced to regulate coil current amplitude and derivatives, ensuring hardware safety and smooth state transitions.Furthermore, the framework incorporates a CBF-based reactive decision-making layer, enabling the microrobot to perform real-time local path replanning and dynamic obstacle avoidance. Theoretical analysis and experimental validation in complex maze environments demonstrate that the proposed method ensures global asymptotic stability, suppresses trajectory oscillations, and significantly enhances the robustness of microrobot navigation in safety-critical applications.
|
| |
| 11:10-11:30, Paper TuAT3.3 | |
| Research on Port AGV Path Tracking Control Based on Deep Reinforcement Learning |
|
| Qi, Xiaohang | Wuhan University of Technology |
| Li, Wenfeng | Wuhan University of Technology |
| Zhang, Qiang | Xinjiang University |
Keywords: Autonomous Vehicle Navigation, Reinforcement, Motion Control
Abstract: Abstract—To improve the path-tracking accuracy and motion stability of port Automated Guided Vehicles (AGVs), this paper proposes a centralized coupled control strategy for the lateral and longitudinal motion of AGVs based on deep reinforcement learning. The proposed strategy is developed based on a two-degree-of-freedom dynamic model of the AGV, where a deep reinforcement learning controller is constructed and the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm is adopted algorithm is adopted to obtain the optimal control actions for the front-wheel steering angle and longitudinal velocity. Through an end-to-end learning mechanism, the controller directly establishes a joint mapping between lateral and longitudinal actions from the overall vehicle dynamics, naturally adapting to the strong coupling characteristics of the system and avoiding the accuracy loss caused by traditional decoupled control methods. Simulation and real-vehicle experiments were conducted under typical port AGV operating conditions in the Matlab/Simulink simulation environment and on a self-developed port experimental platform. The simulation results show that the lateral error at the docking point reaches 0.011 m. In the real-vehicle experiments, the lateral error at the docking point is 0.03 m and the heading angle error is 0.03 rad. The experimental results demonstrate that the proposed control strategy can significantly improve the docking accuracy and cornering stability of AGVs, providing a new technical approach for addressing the poor steering stability of port AGVs.
|
| |
| 11:30-11:50, Paper TuAT3.4 | |
| Motion Planning and Locomotive Reduction for Twistable Snake Robots |
|
| Fan, Qingling | Shanghai Jiao Tong University |
| Ren, Zhongqiang | Shanghai Jiao Tong University |
Keywords: Biomimetics, Motion and Path Planning, Motion Control
Abstract: This paper studies motion planning and locomotive reduction of snake robots, a class of hyper-redundant mechanisms that can traverse complex environments with limbless locomotion. To coordinate the large number of degrees of freedom of snake robots, existing research has developed versatile gaits that are combined via locomotive reduction techniques to simplify the control of a snake robot to that of a simpler system such as a differential-drive car. While being effective in many situations, the terrain traversal speed of the robot is often limited due to its reliance on solely undulatory motion. We believe combining undulation and twisting motion has the potential to speed up terrain traversal and improve motion robustness on complex terrains. With that in mind, this paper develops a rolling vector approach for a twistable snake robot, which cannot only undulate its backbone curve as existing snake robots do, but also twist about its backbone curve infinitely to achieve wheel-like rolling when needed. Results show that, compared to the existing locomotive reduction method that relies solely on undulatory motion, ours is more robust across complex terrains and moves faster (up to a 66.5% improvement) due to twisting motion. We also verify the various motion modes of our rolling vector approach on a real robot.
|
| |
| 11:50-12:10, Paper TuAT3.5 | |
| Disturbance-Augmented Simulations for Improved Sim-To-Real Transfer of Compliant Reinforcement Learning-Based Robotic Control |
|
| Parnada, Anselmo | University of Birmingham |
| Lan, Feiying | University of Birmingham |
| Zhang, Ze | The University of Sheffield |
| Castellani, Marco | University of Birmingham |
| Liang, Chaozhi | University of Birmingham |
| Law, James | The University of Sheffield |
| Tiwari, Ashutosh | University of Sheffield |
| Wang, Yongjing | University of Birmingham |
Keywords: Calibration and Identification, Robust/Adaptive Control, Cyber-physical Production Systems and Industry 4.0
Abstract: Sim-to-real transfer enables scalable acquisition of compliant robotic skills for contact-rich tasks, yet dynamics mismatches between simulation and hardware hinder reliable deployment. Prior approaches assume disturbance-free training dynamics, neglecting endogenous disturbances from nominal-plant modelling errors that degrade real-world performance. This simplification limits transfer of high-compliance six-DoF skills, as prior successes relied on lower-compliance or three-DoF control. Consequently, current methods cannot support skill acquisition for high-dexterity tasks demanding high compliance. We address this gap with Disturbance Augmented Simulations for improved sim-to-real TRAnsfer (DASTRA), a general simulation design framework for structured disturbance modelling to produce disturbance-aware policies more robust to real-world modelling errors. Guided by DASTRA, we propose four disturbance models—inertial parameter mismatch, friction extensions, joint sensor bias, and policy delays—and evaluate their effect on sim-to-real transfer of a six-DoF peg-insertion skill. Results identify inertial mismatch, sensor bias, and policy delays as primary drivers of transfer robustness, providing immediately usable simulation enhancements for high-compliance manipulation tasks. Enhanced friction modelling did not improve performance, underscoring the need for careful disturbance model design. Collectively, these findings establish disturbance-augmented simulator design as essential for reliable transfer of high-dexterity, high-compliance skills. Beyond peg insertion, DASTRA can enable scalable deployment of compliant skills in challenging industrial settings such as assembly and disassembly, and motivates continued refinement to further close the sim-to-real dynamics gap
|
| |
| 12:10-12:30, Paper TuAT3.6 | |
| Dynamic Peg-In-Hole for Automated Refrigerant Recharging on Moving Production Lines |
|
| Bogucki, Dawid Emanuel | University of Pisa |
| Baracca, Marco | University of Pisa |
| Simonini, Giorgio | Università Di Pisa |
| Nottolini, Caterina | FT Srl Future Technologies |
| Ciaffarafà, Sara | FT Srl Future Technologies |
| Salaris, Paolo | University of Pisa |
Keywords: Collaborative Robots in Manufacturing, Compliant Assembly, Motion Control
Abstract: The adoption of collaborative robots in modern manufacturing systems for continuous in-motion assembly operations is steadily increasing. In this context, the peg-in-hole task plays a relevant role, although the most existing work assumes a static assembly scenario. However, in dynamic applications, the required tracking and alignment of the parts make the insertion stage more challenging, especially when tight clearances and additional constraining components are present. To address this, we propose a dynamic peg-in-hole control strategy consisting of a visual servoing tracking phase followed by an in-motion compliant search and insertion stage. We applied this strategy to an industrial refrigerant fluid recharging system. In this system, a specific injector, attached to a velocity-controlled robot, needs to be connected with a hydraulic quick coupling moving on an assembly line. The control strategy was validated both in simulation and on a physical robot, demonstrating its ability to successfully complete the task despite the challenges posed by strict timing constraints, complex mating part geometries, and a dynamic scenario.
|
| |
| TuAT4 |
204 |
| Enabling Technology of Industrial Intelligence 1 |
Regular Session |
| |
| 10:30-10:50, Paper TuAT4.1 | |
| Simulation-Based Multi-Fillet Evaluation of Woody Breast Poultry Fillets |
|
| Sen Mukherjee, Chirantan | University of Texas at Arlington |
| Yoon, Seung-Chul | U.S. Department of Agriculture |
| Beksi, William J. | University of Texas at Arlington |
Keywords: Agricultural Automation, Factory Automation, Computer Vision in Automation
Abstract: Woody breast (WB) is a myopathy in modern broiler chickens that causes the breast muscle to become unusually stiff and fibrous, leading to decreased meat quality and significant economic losses. State-of-the-art automated WB detection relies on a side-view imaging system to analyze the bending behavior of a single fillet as it falls off a conveyor belt. While highly accurate, this approach is constrained by its single-fillet field of view, creating throughput bottlenecks on commercial processing lines. In this paper, we address this limitation via a novel multi-fillet detection architecture utilizing a top down camera configuration. To validate our approach, we first develop a high-fidelity digital twin of an industrial conveyor system. Next, we synthesize a diverse dataset of 3D fillet meshes and model their viscoelastic bending dynamics using a physics-based simulation engine. Lastly, a continuous 2D shape deformation score is extracted from the top-down perspective as the simulated fillets traverse the roller precipice. Experimental results demonstrate that the top-down shape score effectively captures the contour changes of the fillets as it bends, providing a robust and scalable alternative to a side-view imaging system for simultaneous multi-fillet WB evaluation.
|
| |
| 10:50-11:10, Paper TuAT4.2 | |
| Fourier-Guided Dual-Domain Consistency Learning for Semi-Supervised Polyp Segmentation |
|
| Cong, Jingwen | National Frontiers Science Center for Industrial Intelligence and Systems Optimization, Northeastern University |
| Liu, Yue | National Frontiers Science Center for Industrial Intelligence and Systems Optimization, Northeastern University |
| Su, Lijie | National Frontiers Science Center for Industrial Intelligence and Systems Optimization, Northeastern University |
Keywords: AI and Machine Learning in Healthcare, AI-Based Methods, Machine learning
Abstract: Accurate segmentation of the colorectal polyp is crucial for the diagnosis and treatment of computer-assisted colonoscopy, as well as the early prevention of colorectal cancer.However, existing high-performance segmentation models typically rely on large amounts of pixel-level annotated data, which severely limits their broader clinical applicability. To alleviate the problem of annotation scarcity, this paper proposes a Fourier-guided dual-domain consistency semi-supervised polyp segmentation framework. Built upon a Teacher–Student architecture, the proposed method utilizes an EMA teacher model to generate soft pseudo-labels and employs a dynamic confidence filtering mechanism to select high-confidence regions for consistency learning. On this basis, a Fourier-domain consistency module is further introduced, which maps prediction probability maps into logarithmic amplitude spectrum space to enhance the model’s capability for high-frequency boundary structure modeling. Meanwhile, an adversarial consistency module is designed to enforce structural alignment between student predictions and teacher pseudo-labels at the global semantic distribution level through a generative adversarial network (GAN). To improve the overall stability of training, a stage-wise curriculum learning strategy is adopted, progressively achieving supervised warm-up, spatial-frequency collaborative optimization, and global adversarial alignment. Preliminary experiments are planned on the Kvasir-SEG, CVCClinicDB, and ETIS-Larib datasets to systematically evaluate the model’s boundary fidelity, structural integrity, and crossdataset generalization capability under low-annotation settings.
|
| |
| 11:10-11:30, Paper TuAT4.3 | |
| Benchmarking Prompt-Conditioned ROI Inference Via GT-Anchored Validation on Boundary Representation CAD Models |
|
| Lee, Jaeyoung | Korea Institute of Industrial Technology |
| Seo, Yechan | SUNGKYUNKWAN UNIVERSITY |
| Tae, Hyunchul | Korea Institute of Industrial Technology |
| Yoo, Young-Jun | Korea Institute of Industrial Technology |
| Kim, Moon-Jo | Korea Institute of Industrial Technology |
| Won, Hong-In | Korea Institute of Industrial Technology |
Keywords: AI-Based Methods, Computer Vision for Manufacturing, Agent-Based Systems
Abstract: Recent studies in language-guided industrial inspection have begun to infer Regions of Interest (ROIs) directly from natural-language prompts and CAD-native Boundary Representation (B-Rep). This shift is practically important because ROI errors propagate directly to downstream inspection planning: under-selection risks missed critical surfaces, whereas over-selection increases sensing redundancy, viewpoint count, and inspection cycle time. However, inferred ROIs are still typically treated as intermediate outputs, and it remains unclear whether they are actually consistent with benchmark-defined target regions. In this paper, we formulate the prompt-conditioned ROI inference as a measurable validation problem. We propose a GT-anchored ROI validation protocol that converts an inspection prompt into a structured query, constructs a benchmark-derived target ROI, and evaluates the inferred ROI along three complementary dimensions: semantic target selection, rendered-region agreement, and geometric diagnostics. The protocol is instantiated on MFInstSeg using family-specific GT construction rules for direct-feature, bottom-of-feature, surface-with-feature, ranking-surface, spatial-relation, and whole-model prompts. Experimental results show that prompt-conditioned ROI inference is generally reliable for feature-centric and relation-centric prompt families, but also reveal a critical mismatch case in which perfect coarse semantic selection does not translate into correct final ROI geometry. These findings show that ROI validation cannot be reduced to instance-level correctness alone and requires explicit comparison against benchmark-derived target regions before language-guided ROI modules are integrated into downstream industrial inspection planning.
|
| |
| 11:30-11:50, Paper TuAT4.4 | |
| Agentic AI-Enabled Semantic Commissioning of a Cognitive Digital Twin for Reconfigurable Manufacturing |
|
| Liu, Yangyang | University of Auckland |
| Xu, Xun | University of Auckland |
| Polzer, Jan | The University of Auckland |
Keywords: AI-Based Methods, Cyber-physical Production Systems and Industry 4.0, Computer Vision for Manufacturing
Abstract: Rapid bespoke commissioning of the Cognitive Digital Twin (CDT) is a major challenge in reconfigurable manufacturing. Traditional digital twin (DT) construction methods primarily focus on geometric reconstruction, often neglecting the deep semantic integration and functional interoperability necessary for autonomous reasoning. This paper proposes an agent-based, AI-driven workflow to automate end-to-end CDT debugging. The system utilises LangGraph as a multi-agent orchestration engine to achieve dual-path synthesis: the semantic path extracts technical specifications from unstructured documents using Retrieval Augmented Generation (RAG), while the functional path autonomously discovers and binds to real-time industrial telemetry data using Model Context Protocol (MCP). Experimental validation in a robotic machining cell demonstrates that the system achieves a mean average accuracy (mAP) of 97.2% in perception and reduces the deployment cycle from several weeks to an average of 2 hours, marking a paradigm shift from manual scripting to autonomous orchestration.
|
| |
| 11:50-12:10, Paper TuAT4.5 | |
| A Synthetic-Driven Vision System for Assembly Step Recognition |
|
| Zhang, Hui | ETH Zurich, Inspire AG |
| Lei, Xuanang | ETH Zurich |
| Wang, Rui | ETH Zurich |
| Ferchow, Julian | Inspire AG, ETH Zurich |
| Meboldt, Mirko | ETH Zurich |
Keywords: Assembly, Computer Vision for Manufacturing, Intelligent and Flexible Manufacturing
Abstract: Quality control in industrial assembly is essential, and real-time monitoring of the assembly process is crucial for preventing costly defects and ensuring production reliability. Vision-based automated inspection offers a powerful solution for such real-time monitoring. However, due to the specialized industrial components and processes, training these models typically relies on task-specific real-world data, which is costly and labor-intensive to collect and annotate. In this paper, we propose a system that automatically generates realistic assembly sequences and further trains real-time inspection models using the synthetic data. It can be efficiently applied to a given task within an hour, requiring only CAD models and simple step descriptions. Focusing on practical challenges, our system integrates a physics-based motion generation module to capture the variance of different human assembly, designs domain-randomized rendering to deal with the environmental complexity and variation, and employs an object-detection-based step recognition module for robust sim-to-real transfer, leading to 92.4% accuracy on a real-world assembly case with 46.7%, 15.8% and 61.2% performance improvement, respectively. Overall, our system provides a practical solution for industrial assembly inspection without requiring expensive real-world data collection and annotation, with the effectiveness validated on real industrial assembly tasks.
|
| |
| 12:10-12:30, Paper TuAT4.6 | |
| GAN-Augmented Crack Detection in Shadowed Infrastructure Environments Employing an Enhanced YOLO11 Model |
|
| Lyu, Chen | The University of Auckland |
| Kobayashi, Masahiro | The University of Auckland |
| Lynch, Angus | University of Auckland |
| Xu, Xun | University of Auckland |
| Liarokapis, Minas | National Technical University of Athens |
Keywords: Automation in Construction, Automation Technologies for Smart Cities, Computer Vision in Automation
Abstract: Crack detection is critical for infrastructure maintenance. However, existing public datasets lack images with complex backgrounds, such as shadows, limiting the model's generalization ability in real-world scenarios. To tackle this problem, this study proposes a lightweight and annotation-free framework: (1) enhancing YOLO11s with Convolutional Block Attention Module (CBAM) to improve model's feature extraction capability; (2) employing Wasserstein GAN (WGAN) to augment training data without manual annotation. The WGAN is trained to generate synthetic shadow masks for the fusion with shadow-free crack images for shadowed data augmentation. Model performance is evaluated on two self-built test sets: shadow-free test set and shadowed test set. Results show that YOLO11s-CBAM achieves superior accuracy and stability across all metrics (precision, recall, mAP@50, mAP@50-95). By investigating the integration ratio of synthetic shadow-crack images, we found that adding 30% synthetic data for data training can enhance the model's shadow robustness optimally. However, excessive synthetic data degrades the WGAN-based improvement efficiency. This research provides a lightweight and annotation-free solution for crack detection in shadowed environments, offering practical value for real-time and efficient infrastructure inspection applications.
|
| |
| TuAT5 |
205 |
| Emerging Technology of Automation System 1 |
Regular Session |
| |
| 10:30-10:50, Paper TuAT5.1 | |
| Autonomous Reactive Masonry Construction Using Collaborative Heterogeneous Aerial Robots with Experimental Demonstration |
|
| Stamatopoulos, Marios-Nektarios | Luleå University of Technology |
| Small, Elias | Luleå University of Technology |
| Velhal, Shridhar | Lulea Technical University |
| Banerjee, Avijit | Luleå University of Technology |
| Nikolakopoulos, George | Luleå University of Technology |
Keywords: Automation in Construction
Abstract: This article presents a fully autonomous aerial masonry construction framework using heterogeneous unmanned aerial vehicles (UAVs), supported by experimental validation. Two specialized UAVs were developed for the task: (i) a brick-carrier UAV equipped with a ball-joint actuation mechanism for precise brick manipulation, and (ii) an adhesion UAV integrating a servo-controlled valve and extruder nozzle for accurate adhesion application. The proposed framework employs a reactive mission planning unit that combines a dependency graph of the construction layout with a conflict graph to manage simultaneous task execution, while hierarchical state machines ensure robust operation and safe transitions during task execution. Dynamic task allocation allows real-time adaptation while minimum-jerk trajectory generation ensures smooth and precise UAV motion during brick pickup and placement. Additionally, the brick-carrier UAV employs an onboard vision system that estimates brick poses in real time using ArUco markers and a geometry-constrained least-squares optimization filter, enabling accurate alignment during construction. To the best of the authors’ knowledge, this work represents the first experimental demonstration of fully autonomous aerial masonry construction using heterogeneous UAVs, where one UAV precisely places the bricks while another autonomously applies adhesion material between them. The experimental results showcase the effectiveness of the proposed framework and demonstrate its potential to serve as a foundation for future developments in autonomous aerial robotic construction.
|
| |
| 10:50-11:10, Paper TuAT5.2 | |
| Autonomous Collision-Free Navigation Via Anisotropic Interfered Fluid Dynamical System and Look-Ahead Trap Perception Mechanism |
|
| He, Xinqi | Nanjing University of Aeronautics and Astronautics |
| Peng, Xiuhui | Nanjing University of Aeronautics and Astronautics |
| Wen, Guanghui | Southeast University |
Keywords: Autonomous Vehicle Navigation, Collision Avoidance
Abstract: This paper addresses the limitations of autonomous obstacle avoidance for agents operating in complex 2D environments by proposing an anisotropic interfered fluid dynamical system (A-IFDS) navigation architecture. An anisotropic modulation matrix based on relative velocity is introduced to directionally reshape obstacle influence according to relative motion, enabling more proactive avoidance of approaching dynamic obstacles. For stagnation in concave regions, a look-ahead trap perception mechanism (LTPM) is developed to detect potential traps within a limited field of view and trigger hybrid switching between A-IFDS navigation and trap escape. Model predictive control (MPC) is further employed to track the generated reference velocity while satisfying velocity and acceleration limits, ensuring executable motion under physical constraints. Simulations demonstrate that, compared with APF, DWA, A-IFDS without LTPM, and conventional IFDS, the proposed method improves reachability in concavetrap environments, initiates more proactive avoidance of dynamic obstacles, and reduces control effort under velocity and acceleration constraints.
|
| |
| 11:10-11:30, Paper TuAT5.3 | |
| Improving Ranging Accuracy of FMCW LiDAR Using Temporal Convolutional Network |
|
| Cheng, Lawrence R | University of Illinois at Urbana-Champaign |
| Hu, Hua | Photon Logic |
Keywords: Autonomous Vehicle Navigation, Deep Learning in Robotics and Automation, AI-Based Methods
Abstract: FMCW LiDAR has become a cornerstone sensing technology for autonomous vehicles, robotics, and intelligent systems. However, its ranging accuracy is fundamentally constrained by nonlinear chirp distortion and laser phase noise. This work presents an integrated framework that addresses both challenges within a unified processing pipeline. The proposed approach uses IQ demodulation to preprocess the received signal and leverages a temporal convolutional network (TCN) to estimate nonlinear phase errors and recover the beat frequency. By compensating these distortions simultaneously, the method achieves a substantial boost in ranging precision. Evaluations and testing show that the TCN‑based solution reduces relative ranging error by nearly an order of magnitude compared with the conventional FFT‑based processing commonly used in industrial FMCW LiDAR systems.
|
| |
| 11:30-11:50, Paper TuAT5.4 | |
| SPC-ODD: A Model-Agnostic Statistical Resilience Layer for State Estimators in Safety-Critical Autonomous Systems |
|
| Chen, Kuanlin | Chung Yuan Christian University |
| Ou, Cheng-En | Chung Yuan Christian University |
Keywords: Autonomous Vehicle Navigation, Failure Detection and Recovery, Probability and Statistical Methods
Abstract: Industrial autonomous mobile robots and cyber-physical manufacturing systems increasingly rely on safety-critical state estimation pipelines; under unmodeled sensor degradation, these systems face catastrophic divergence that invalidates Operational Design Domain guarantees on the factory floor. To address this, we propose SPC-ODD (Statistical Process Control for Operational Design Domain), a deterministic, pluggable engineering protection layer combining Statistical Process Control (SPC) with kinematically adaptive gating. Model-agnostic and non-intrusive to any filter architecture, the framework employs exponentially weighted moving averages (EWMA) and cumulative sum (CUSUM) charts to bound detection latency via analytical Average Run Length (ARL) theorems, with hyperparameters (h drift=6.0, k drift=0.5) derived from Siegmund's approximation. Evaluated across 2,000 rigorous Monte Carlo trials with highly dynamic kinematics and first-order Gauss-Markov (FOGM)-biased inertial measurement unit (IMU) physics, SPC-ODD empirically suppresses catastrophic divergence to zero within tested regimes and reduces the out-of-spec (OOS) ratio to 0.0% across all evaluated anomaly regimes. The framework executes with a computational overhead of less than 5 μs per step and explicitly aligns with the auditable safety requirements of ISO 21448 (SOTIF). Cross-platform validation on the EuRoC MH_01_easy micro aerial vehicle (MAV) sequence confirms detection consistency (Δt = 0.804 s) without re-tuning, demonstrating framework generalizability beyond ground AMR deployment.
|
| |
| 11:50-12:10, Paper TuAT5.5 | |
| Online Quantification of the Map-Environment Similarity to Estimate and Improve Localization Performance for Mobile Robots |
|
| Rosendahl Dam, Christian | University of Southern Denmark |
| Suvei, Stefan-Daniel | Mobile Industrial Robots |
| Bodenhagen, Leon | University of Southern Denmark |
Keywords: Autonomous Vehicle Navigation, Formal Methods in Robotics and Automation, Industrial and Service Robotics
Abstract: Autonomous mobile robots that localize using a map based approach struggle when the surrounding environment has changed significantly compared to the map. Systems that continuously update the map risk corrupting the map as a result of false data associations and wrong pose estimates. The performance of these systems can be improved using a metric that quantifies the difference between the observed environment and its reference map. This paper proposes a method to compute this value, named the Environment Similarity Index. The metric is implemented based on a 2D lidar and tested both in a simulated and a real environment. The experiments show that, given that the localization error is within normal range, the index can accurately represent the true map-environment difference, and that there is a very strong negative correlation between the value of the index and the expected pose error of Adaptive Monte-Carlo Localization. The index can therefore be used as a general indicator of when the pose error will be impacted by these changes. Finally, the experiments show how the index can be used actively to generate heatmaps of observed local map deviations and to improve localization performance in changing environments by dynamically adjusting the magnitude of the expected odometry noise.
|
| |
| 12:10-12:30, Paper TuAT5.6 | |
| Fast Adaptive Planning for Autonomous Tractor-Trailer Systems Via Iterative LQR Steering |
|
| Zhou, Tianyu | Purdue University |
| Wang, Yebin | Mitsubishi Electric Research Laboratories |
| Umat, Akhil | Mitsubishi Electric Automotive America |
| Tiwari, Astha | Mitsubishi Electric Automotive America, Inc |
| Di Cairano, Stefano | Mitsubishi Electric Research Laboratories |
Keywords: Autonomous Vehicle Navigation, Motion and Path Planning
Abstract: Aiming to tackle the computation challenges arising from the stringent positioning accuracy requirement as well as the uncertainties in system dynamics and constraints, this paper presents a fast adaptive planning method, Adaptive-iAGT, for tractor-trailer systems. Built on~cite{ZheWanCai24}, which grows the tree using pre-computed motion primitives (MPs) and performs A*-like search for kinodynamic planning, the proposed method bears three distinctive features. First, iterative LQR (iLQR) is deployed to solve the stabilization problem around the goal state during goal-reaching trials, which enhances the success rate of goal-reaching and significantly reduces planning time. Second, we propose a two-stage framework to treat system uncertainties. Specifically, path planning at the first stage determines a nominal trajectory based on nominal system parameters, and then adaptive motion planning computes the final trajectory based on the nominal trajectory and real system parameters by solving a sequence of iLQR tracking problems. This approach eliminates the need to pre-compute and store MPs for different sets of model parameters. Third, an analytical scaling method is introduced in motion planning to accommodate customizable velocity and control limits. This allows us to neglect these constraints when solving iLQR problems during goal-reaching and trajectory adaptation, and drastically improve their success rate. Simulations and field tests validate the effectiveness of the proposed algorithm. The proposed framework can be readily extended to other motion systems.
|
| |
| TuAT6 |
503 |
LLM-Based Application for Manufacture-Circulation Industrial System (MCIS)
1 |
Regular Session |
| |
| 10:30-10:50, Paper TuAT6.1 | |
| A Knowledge-Driven and Event-State-Aware AI-Agent Framework for Intelligent Crane Hoisting Decision-Making |
|
| Li, Laiyi | State Key Laboratory for Manufacturing Systems Engineering |
| Li, Yuke | State Key Laboratory for Manufacturing Systems Engineering |
| Caike, Zhenyu | State Key Laboratory for Manufacturing Systems Engineering |
| Yang, Maolin | State Key Laboratory for Manufacturing Systems Engineering |
| Jiang, Pingyu | State Key Laboratory for Manufacturing Systems Engineering |
Keywords: Autonomous Agents, Intelligent and Flexible Manufacturing, AI-Based Methods
Abstract: Intelligent decision-making for crane hoisting in complex environments is a significant challenge crucial for operational safety and efficiency. Typically, current operations rely on rigid logic or human experience but lack the cognitive flexibility for autonomous operations. Furthermore, the development of intelligent cranes is heavily constrained by unstandardized decision-making processes and insufficient high-quality domain knowledge sets. This paper constructs a knowledge- driven AI agent framework to address these issues. First, a Scene-Process-Action graphical model is established to formalize the decision space by mapping discrete scene elements to continuous processes. Subsequently, the schema model guides the mapping of event-state logic and domain input-process-output objects into structured prompts, which drive large language models to generate high-quality knowledge sets. Finally, a knowledge-driven and event-state-aware AI agent is constructed by embedding an event-state triggering mechanism into the cognitive core to realize real-time process perception and intelligent decision-making. A case study of the bridge crane involving graphical modeling and key decision points is conducted, successfully constructing decision knowledge and the corresponding AI agent. The results demonstrate that the proposed framework enables real-time process perception and reliable decision-making in dynamic industrial crane environments.
|
| |
| 10:50-11:10, Paper TuAT6.2 | |
| Prompt Multi-Agent Decision Transformer for Energy-Saving Heterogeneous Networks |
|
| Zhou, Yuxuan | Nanjing University of Science and Technology |
| Xia, Pengcheng | Nanjing University of Science and Technology |
| Ni, Yiyang | Jiangsu Second Normal University |
| Li, Jun | Southeast University |
|
|
| |
| 11:10-11:30, Paper TuAT6.3 | |
| Dynamic Ensemble Framework with Neural Meta-Learner for Carbon Dioxide Emission Forecasting in the Steel Industry |
|
| Liu, Weishuo | Northeastern University |
| Lang, Jin | Northeastern University |
| Zhang, Yanyan | Northeastern University |
| Zhao, Shengnan | Northeastern University |
Keywords: Big data Analytics for Large-scale Energy Systems, AI-Based Methods, Big-Data and Data Mining
Abstract: The steel industry is a major contributor to global carbon dioxide emissions, and forecasting its emissions is challenging due to the nonlinear and time-dependent nature of production processes. This study proposes a dynamic ensemble forecasting framework that integrates Support Vector Regression, Random Forest, and Long Short-Term Memory networks, with a Dynamic Ensemble Neural Network (DENN) serving as a meta-model to adaptively compute the weights of base models. Experiments conducted on real-world energy consumption and emission data from Daewoo Steel demonstrate that the proposed framework consistently outperforms individual models and conventional ensemble methods, achieving lower prediction errors and effectively capturing both shortterm ffuctuations and long-term trends. These ffndings indicate that the proposed framework enhances both the accuracy and robustness of carbon dioxide emission forecasting, offering a reliable data-driven approach for carbon management and sustainable development in the steel industry.
|
| |
| 11:30-11:50, Paper TuAT6.4 | |
| A Blockchain-Based Energy Trading Framework for Enhancing Prosumers’ Profitability in Green Manufacturing Microgrids |
|
| Hung, Min-Hsiung | Chinese Culture University |
| Liu, Yu-Han | Institute of Manufacturing Information and Systems, National Cheng Kung University |
| Lin, Yu-Chuan | Chinese Culture University |
| Chen, Chao-Chun | National Cheng Kung University |
| Tieng, Hao | National Cheng Kung University |
| Cheng, Fan-Tien | National Cheng Kung University |
Keywords: Cyber-physical Production Systems and Industry 4.0, Intelligent and Flexible Manufacturing, Renewable Energy Sources
Abstract: Traditional energy markets struggle to efficiently allocate surplus green energy because legacy power grids enforce static pricing models and delay settlements. To address this problem, we propose the Microgrid Two-stage Double Auction Scheme (MTDAS). This blockchain platform enables direct energy trading while protecting internal privacy and ensuring publicly verifiable settlements. MTDAS features dynamic pricing, order aggregation, and instant profit tracking. Our evaluations show that MTDAS outperforms recent blockchain energy auctions. Specifically, it boosts user profits by 31.66%, increases the number of successful trades by 38.98%, and reduces processing time by 65.17%. The proposed method offers a promising solution for the manufacturing industry to monetize renewable energy and maximize financial returns.
|
| |
| 11:50-12:10, Paper TuAT6.5 | |
| Generative Digital-Twin Prototyper: An LLM-Based Framework for Automated Process Design and Analysis Tool Integration |
|
| Lee, Geonchang | Sungkyunkwan University |
| Woo, Honguk | Sungkyunkwan University |
Keywords: Cyber-physical Production Systems and Industry 4.0, Intelligent and Flexible Manufacturing, Software, Middleware and Programming Environments
Abstract: Implementing digital twins, a core technology of smart industry, requires sophisticated manufacturing data and a simulation framework built upon it. However, real-world manufacturing faces a dual barrier: data inaccessibility that prevents non-experts from entering the field, and interface fragmentation between tools that burdens experts. In this paper, we propose GDP (Generative Digital-twin Prototyper), a framework that leverages large language models (LLMs) to simultaneously address both user groups. GDP enables non-experts to generate process data in a zero-shot manner without manufacturing knowledge, providing a virtual testbed applicable to Physical AI research including humanoid robots. For experts, GDP dramatically reduces interface construction costs through LLM-based automatic adapter generation and an Auto-Repair mechanism. Experimental results show that zero-shot process generation across 10 product families achieves an F1 of 88.9%, demonstrating the feasibility of data acquisition by non-experts. In tool integration experiments, Auto-Repair achieves a 100% execution pass rate while observing 16.2% “silent failures,” demonstrating the necessity of two-tier evaluation beyond execution-only assessment.
|
| |
| 12:10-12:30, Paper TuAT6.6 | |
| OptiGovernor: A Lightweight Model-Driven CPU Frequency Governor for Server Energy Efficiency |
|
| Zhu, Zirui | Xi'an Jiaotong University |
| Xu, Zhanbo | Xi'an Jiaotong University |
| Liu, Yaping | Xi'an Jiaotong University |
| Wu, Jiang | Xian Jiaotong University |
| Shen, Yuanjun | Xi'an Jiaotong University |
| Wu, Jiajun | Xi'an Jiaotong University |
| Wu, Junjie | Huawei Technologies Co., Ltd |
| Wang, Jiangtao | Huawei Technologies Co., Ltd |
Keywords: Energy and Environment-aware Automation, Optimization and Optimal Control, Model Learning for Control
Abstract: Modern data centers face rising energy costs, with CPUs as major contributors and existing DVFS policies often failing to exploit workload characteristics, adapt across applications, or respect performance constraints. This paper presents OptiGovernor, a hybrid DVFS management framework that combines offline energy efficiency modeling with online, lightweight control to improve server-level energy efficiency. In the offline phase, we build regression models from extensive measurements across frequencies and workloads, using hardware performance counters—including instruction throughput, cache and TLB behavior, and stall metrics—to predict an energy-efficiency score. At runtime, a user-space governor continuously samples these counters, evaluates the learned model for each candidate frequency, and selects the setting that maximizes predicted efficiency while maintaining stable control via hysteresis. We implement OptiGovernor on a Linux server platform and evaluate it using industry-standard energy-efficiency benchmarks, including BenchSEETM and SERTTM. Without any workload-specific tuning, OptiGovernor achieves up to 12.40% and 4.88% higher energy-efficiency scores than the schedutil and ondemand governors, respectively, demonstrating the value of holistic, workload-aware DVFS in production-like environments.
|
| |
| TuAT7 |
504 |
AI for Safety, Sustainability, and Automation in Smart Maritime Transport
and Logistics Systems |
Special Session |
| Organizer: Kim, Heeyoung | KAIST |
| Organizer: Kim, Sungil | Ulsan National Institute of Science and Technology |
| Organizer: Koo, Wonmo | Korea Advanced Institute of Science and Technology |
| Organizer: Kim, Dohee | Changwon National University |
| |
| 10:30-10:50, Paper TuAT7.1 | |
| AIS-Based Vessel Trajectory Prediction Using Memory-Augmented Neural Networks (I) |
|
| Koo, Wonmo | KAIST |
| Chang, Sanha | KAIST |
| Kim, Heeyoung | KAIST |
Keywords: Collision Avoidance, AI-Based Methods, Intelligent Transportation Systems
Abstract: Accurate vessel trajectory prediction is essential for safe and efficient maritime operations, enabling collision avoidance and supporting route optimization. Although memory-augmented neural networks have recently shown strong performance in pedestrian and road-vehicle trajectory prediction by selectively retrieving relevant information from an external memory, their potential for vessel trajectory prediction remains underexplored. This paper presents an empirical investigation of memory-based trajectory prediction using Automatic Identification System (AIS) data. Experiments on data from the Gulf of Mexico and the New York Bight demonstrate consistent and substantial performance gains over a range of deep learning baselines that do not incorporate an external memory.
|
| |
| 10:50-11:10, Paper TuAT7.2 | |
| Rarity-Gated Context Conditioning for Offline Imitation Learning-Based Maritime Anomaly Detection (I) |
|
| Kim, Yongmin | Ulsan National Institute of Science and Technology |
| Jeon, ByeongHoon | Ulsan National Institute of Science and Technology |
| Kim, Sungil | Ulsan National Institute of Science and Technology |
Keywords: AI-Based Methods, Failure Detection and Recovery, Logistics
Abstract: Contextual anomaly detection aims to identify abnormal behavior conditional on context variables, but practical deployments often face highly imbalanced context distributions where rare regimes can be critical information. Under such frequency bias, context-conditioned models can produce unstable decisions and excessive false alarms in rare contexts. We propose Rarity-Gated Feature-wise Linear Modulation (RGFiLM), a rarity-aware conditioning module that combines feature-wise modulation (i.e., context-conditioned scaling and shifting of hidden features) with a gate controlled by a data-driven rarity score. The rarity score is estimated from the empirical distribution of context variables and regulates how strongly context modulates intermediate representations: the gate becomes more decisive under rare contexts while remaining conservative under frequent contexts. We evaluate RGFiLM on maritime trajectory anomaly detection using AIS motion sequences with ERA5 environmental context in an environment-sensitive detour scenario. When instantiated in a sequential anomaly scoring pipeline, RGFiLM improves the precision--recall trade-off over context-agnostic and context-naive baselines, achieving higher F1/precision than standard Feature-wise Linear Modulation (FiLM) conditioning while maintaining comparable recall and reducing the false positive rate. These results suggest that explicitly accounting for context rarity is an effective approach for reducing false alarms in context-sensitive anomaly detection.
|
| |
| 11:10-11:30, Paper TuAT7.3 | |
| Design and Implementation of an LLM-Based Maritime Risk Scenario Generation System for Vessel Traffic Services (I) |
|
| Park, Taekhyun | Pusan National University |
| Jo, Sangmin | Pusan National University |
| Hong, SeongMoon | Pusan National University |
| Kim, Dohee | Changwon National University |
| Bae, Hyerim | Pusan National University |
Keywords: AI-Based Methods, Big-Data and Data Mining, Intelligent Transportation Systems
Abstract: Maritime accident prevention requires timely, interpretable risk assessment that can support operational decision-making in Vessel Traffic Service (VTS) environments. We propose an integrated framework that combines Large Language Models (LLMs) with machine-learning-based risk inference to unify accident risk probability estimation and dynamic scenario generation. Heterogeneous data sources—Automatic Identification System (AIS) trajectories, meteorological observations, historical accident statistics, and tribunal adjudication reports—are standardized through a two-path LLM pipeline and aligned on a uniform spatial grid. A gradient boosting classifier with SHAP (SHapley Additive exPlanations)-based explainability identifies key risk drivers, and isotonic regression calibration produces well-calibrated, continuous risk scores. For scenario generation, a hypergraph-based retrieval augmented generation (RAG) knowledge base encodes multi-way relationships among accident cases, causal factors, and maritime regulations; retrieved evidence is combined with calibrated risk scores in few-shot, Chain-of-Thought (CoT) prompts to generate interpretable accident scenarios with a cause–progression–outcome structure. A web-based dashboard with on-device text-to-speech delivers grid-level risk visualization, narrative risk scenarios, and actionable navigational advisories within a unified prediction–explanation–response interface for maritime safety management.
|
| |
| 11:30-11:50, Paper TuAT7.4 | |
| Function-Space Priors for Bayesian Neural ODEs with Application to Vessel Trajectory Prediction (I) |
|
| Lee, Jaeyeong | KAIST |
| Koo, Wonmo | KAIST |
| Kim, Heeyoung | KAIST |
Keywords: Collision Avoidance, Calibration and Identification, Machine learning
Abstract: Vessel trajectory prediction from Automatic Identification System (AIS) data is essential for maritime situational awareness, yet it remains challenging due to irregular sampling, missing reports, and complex dynamics. Beyond accurate point forecasts, maritime applications also demand well-calibrated uncertainty estimates for reliable decision-making. Bayesian Neural Ordinary Differential Equations (ODEs) offer a principled framework for continuous-time trajectory modeling with uncertainty quantification by placing a prior over the neural vector field parameters. However, the commonly used isotropic Gaussian weight prior fails to encode informative structural properties of vessel dynamics, such as smoothness and locality. Existing function-space Bayesian neural network methods address this limitation for static mappings, but do not transfer directly to Neural ODEs, where the primary quantity of interest is the trajectory rather than the vector field itself. In principle, one could place a Gaussian process (GP) prior directly over ODE solutions, but this requires propagating distributions through a nonlinear ODE solver, which is analytically intractable. To address this challenge, we adopt a practical approach that imposes a GP-kernel-based prior directly on the vector field evaluated at a finite set of measurement points. Specifically, we augment the standard weight-space variational objective with a kernel-based regularizer that penalizes deviations of the vector field from the structure implied by a GP prior. We evaluate the proposed method on real-world AIS datasets and compare it against weight-space Bayesian Neural ODEs and GP-based ODE baselines. Experimental results demonstrate improvements in both predictive accuracy and uncertainty calibration.
|
| |
| 11:50-12:10, Paper TuAT7.5 | |
| Multi-Field Hybrid Retrieval-Augmented Generation for Maritime Accident Root Cause Analysis (I) |
|
| Kim, Seongjin | Ulsan National Institute of Science and Technology |
| Kim, Yongmin | Ulsan National Institute of Science and Technology |
| Kim, Sungil | Ulsan National Institute of Science and Technology |
Keywords: AI-Based Methods, Data fusion, Collision Avoidance
Abstract: Maritime accident adjudication reports contain critical tribunal findings for root cause analysis (RCA), yet retrieving relevant precedents and drafting consistent reports from decades of records remains labor-intensive. This paper proposes a multi-field hybrid retrieval-augmented generation (RAG) framework for automated maritime RCA, utilizing a comprehensive dataset of 13,329 Korea Maritime Safety Tribunal (KMST) reports (1971–2025). We transform raw adjudications into a structured knowledge base of "incident cards," indexing three distinct fields—Summary, Causes, and Disposition—alongside a hierarchical L1/L2 cause taxonomy. Our retrieval strategy employs a field-aware hybrid approach, fusing sparse and dense rankings via Reciprocal Rank Fusion (RRF). Given the lack of large-scale expert relevance labels, we evaluate retrieval performance using ceiling-normalized recall and nDCG based on a metadata-derived proxy relevance score. Experimental results demonstrate that our proposed retrieval significantly outperforms baseline methods, improving NormRecall@100 from 0.18 to 0.55. Furthermore, grounding the generator on the retrieved precedents enhances RCA generation quality over an LLM-only baseline, increasing the LLM-as-a-judge score from 3.34 to 3.72. These findings suggest that field-aware RAG can substantially streamline maritime safety investigation workflows by enabling faster precedent search and more consistent, evidence-based RCA drafting.
|
| |
| 12:10-12:30, Paper TuAT7.6 | |
| Diffusion-Based Accident Data Augmentation for Ship Collision Identification Using Relative Maritime Traffic Map (I) |
|
| Kim, Somyeong | Pusan National University |
| Kim, Dohee | Changwon National University |
| Bae, Hyerim | Pusan National University |
Keywords: AI-Based Methods, Big-Data and Data Mining, Data fusion
Abstract: This study proposes a diffusion-based data augmentation framework to address the severe class imbalance in Ship-to-Ship (S2S) collision identification. While AIS data provides abundant normal samples, accident-related encounter patterns are extremely scarce. To address this issue, we employ diffusion models, including DDPM, LDM, and NCSN++ with Conditional Learning, to generate high-fidelity Relative Maritime Traffic Maps (RMTM). Quality evaluations using KID, EMD, and spatial correlation metrics showed that Conditional DDPM achieved the highest fidelity, particularly at a 10 km radius (KID: 0.0001, EMD: 0.0121). Experimental results across multiple observation radii (5, 10, and 20 km) demonstrated that the proposed augmentation significantly outperforms traditional class-weighting and undersampling strategies. Notably, at a 20 km radius, MobileNet achieved an F1-score of 0.9532 and a Recall of 0.9909. These findings confirm that diffusion-based augmentation effectively restores the statistical distribution of rare accident data, providing a reliable foundation for enhancing maritime safety and autonomous navigation systems.
|
| |
| TuAT8 |
501 |
| Virtual Session 1 |
Regular Session |
| |
| 10:30-10:50, Paper TuAT8.1 | |
| Local String Stability-Driven Reinforcement Learning for Adaptive Swarm Coordination |
|
| Wei, Yusi | Durham University |
| E, Wenke | Durham University |
| Su, Yu-Hsiang | Durham University |
| Atapour-Abarghouei, Amir | Durham University |
| Arvin, Farshad | Durham University |
| Hu, Junyan | Durham University |
Keywords: Robot Networks, Autonomous Agents, Control Architectures and Programming
Abstract: This paper presents a framework that integrates input-to-state string stability (ISSS) theory with reinforcement learning to achieve stable, scalable, and adaptive swarm coordination. First, a virtual string perception decomposes N-agent interactions into parallel robot-following subtasks by synthesizing local virtual neighbors through safety-weighted aggregation of sensed positions. Second, an actor learns adaptive proportional-derivative gain vectors from local observations of the virtual formation structure, enabling real-time gain adaptation without global knowledge. Third, an ISSS-driven reward function aligns the reinforcement learning objective with Lyapunov decrease conditions via potential-based reward shaping, allowing standard learning algorithms to discover stabilizing policies. Simulation experiments demonstrate sublinear convergence scaling and superior disturbance amortization compared to baselines; real-robot validation on Mona robots confirms stable formation coordination.
|
| |
| 10:50-11:10, Paper TuAT8.2 | |
| Efficient Autoscaling of Cloud Applications Using MPC: A Complementarity Constraints Formulation |
|
| Masti, Daniele | Gran Sasso Science Institute |
| Smarra, Francesco | University of L'Aquila |
Keywords: Planning, Scheduling and Coordination, Optimization and Optimal Control, Cloud Computing For Automation
Abstract: Cloud applications based on microservices architecture require autoscaling mechanisms that can react to time-varying workloads while accounting for provisioning delays and discrete resource capacities. Model predictive control (MPC) approaches based on queueing-network models provide a principled way to handle these transient effects, but incorporating saturation phenomena, discrete capacity levels, and actuation logic typically leads to mixed logical dynamical (MLD) models and mixed-integer optimization problems that can be challenging to solve in real-time. This paper proposes an MPC formulation for autoscaling of cloud applications modeled as fluid queueing networks, where the hybrid dynamics induced by service saturation are represented using stage-wise complementarity constraints. By exploiting the piecewise-affine structure of the queueing dynamics, the resulting prediction model admits a linear time-invariant state update with a stable transition matrix, while nonsmooth effects are captured locally through complementarity constraints, thus enabling a faster solution than standard MLD-based formulations. The proposed approach is evaluated in simulation on autoscaling scenarios of increasing difficulty.
|
| |
| 11:10-11:30, Paper TuAT8.3 | |
| Accelerating Large-Scale Bundle Adjustment for LiDAR Mapping Via Parallel Computing |
|
| Cai, Yixi | KTH Royal Institute of Technology |
| Li, Rundong | University of Hong Kong |
| Xie, Yuhan | The University of Hong Kong |
| Zhang, Qingwen | KTH Royal Institute of Technology |
| Jensfelt, Patric | KTH - Royal Institute of Technology |
| Zhang, Fu | University of Hong Kong |
Keywords: Sensor Fusion
Abstract: LiDAR bundle adjustment is widely utilized in mapping to construct globally consistent point cloud maps. In this paper, we propose the first fully parallel computing framework to accelerate LiDAR bundle adjustment for large-scale mapping, incorporating three key techniques. First, we design an adaptive, asynchronous data loading strategy to efficiently process large-scale point cloud datasets on memory-constrained GPUs. Secondly, we present a novel bottom-up voxelization method for extracting planar features, enabling fully parallelized pre-processing. Thirdly, we build upon a majorization-minimization formulation to accelerate compute-intensive tasks in the optimization via parallel computation, including the computation of residuals, Jacobian and Hessian matrices, and a parallel increment solver. To support our design, we provide both theoretical and experimental analysis of the time complexity of our approach. Extensive benchmarking on large-scale public datasets across various computational platforms validates the robustness and adaptability of our approach, achieving up to a tenfold improvement in computational efficiency while preserving mapping accuracy comparable to state-of-the-art methods. To benefit future research, the implementation code will be made publicly available on GitHub upon acceptance.
|
| |
| 11:30-11:50, Paper TuAT8.4 | |
| Reinforcement Learning-Based Control of a Snake Robot for Navigating Narrow Spaces between Walls |
|
| Sasaki, Hiroki | Tokyo University of Agriculture and Technology |
| Ariizumi, Ryo | Tokyo University of Agriculture and Technology |
| Tanaka, Motoyasu | The Univ. of Electro-Communications |
Keywords: Deep Learning in Robotics and Automation, Motion Control, Industrial and Service Robotics
Abstract: Snake robots can adapt to various environments thanks to their elongated shape and high degree of freedom, similar to their biological counterparts. Therefore, they are promising solutions for inspecting narrow spaces between buildings where human access is limited. In such confined environments, manual control by a human operator is extremely difficult due to limited visibility from the outside; thus, autonomous locomotion is required. While several motions for snake robots moving within narrow spaces between walls are proposed and verified in physical simulations, they have yet to be validated using an actual robot. In this study, we evaluate the effectiveness of these specific motions―straight and turning motions―using an actual robot in a real-world environment. Furthermore, we propose a path-following control method for a snake robot navigating within narrow spaces between walls by using reinforcement learning to adjust its motion parameters. This study expands the operational envelope of snake robots, advancing the feasibility of autonomous navigation within narrow spaces between walls.
|
| |
| 11:50-12:10, Paper TuAT8.5 | |
| A Progress-Aware Leader-Follower Midair Docking System for Dual-Drone Aerial Manipulation |
|
| Cai, Yifan | University College London |
| Tan, Jan Ming Kevin | University College London |
| Li, Xiangqi | University College London |
| Jin, Chenzhe | University College London |
| Kemsaram, Narsimlu | University of Malaya (UM) |
| Modugno, Valerio | University College London |
Keywords: Planning, Scheduling and Coordination, Motion Control, Assembly
Abstract: Reliable midair docking between small unmanned aerial vehicles (UAVs) is essential for modular aerial cooperation and manipulation, but it requires precise relative-pose control and repeatable platform under tight thrust and payload constraints. We present a dual-drone docking platform where two quadrotors operate in a leader-follower formation and dock using a lightweight modular frame with passive magnetic latching. A progress-aware mission supervisor manages phase transitions: approach, alignment, capture, and settle. This platform integrates a complete hardware-software stack (ROS 2 with Crazyflie/PX4 interfaces) and synchronized logging for benchmark evaluation. We evaluate the platform in simulation and real-world experiments using quantitative metrics such as formation error, baseline, and yaw consistency, docking success rate, time-to-dock, and failure-mode statistics. The platform enables statistically grounded comparison of docking supervision and synchronization strategies and provides a practical testbed for modular aerial cooperation and repeatable midair aerial manipulation.
|
| |
| 12:10-12:30, Paper TuAT8.6 | |
| Dexterous Control of an 11-DOF Redundant Robot for CT-Guided Needle Insertion with Task-Oriented Weighted Policies |
|
| Zhang, Peihan | University of California San Diego |
| Chen, Derek | University of California, San Diego |
| Duriseti, Ishan | University of California San Diego |
| Richter, Florian | University of California, San Diego |
| Chiu, Zoe | Cornell University |
| Bohley, Moira | University of California San Diego |
| Yip, Michael C. | University of California, San Diego |
Keywords: Medical Robots and Systems, Motion Control
Abstract: Computed tomography (CT)-guided needle biopsies are critical for diagnosing a range of conditions, including lung cancer, but present challenges such as limited in-bore space, prolonged procedure times, and radiation exposure. Robotic assistance offers a promising solution by improving needle trajectory accuracy, reducing radiation exposure, and enabling real-time adjustments. In our previous work, we introduced a robotic platform designed for accurate needle insertion within the confined CT bore. However, its performance in clinical settings is restricted by limited dexterity and a constrained workspace. In this study, we present an 11-degree-of-freedom (DOF) robotic system that integrates a 6-DOF robotic base with an improved 5-DOF cable-driven end-effector, yielding a significantly expanded workspace and enhanced dexterity. To leverage the hyper-redundant degrees of freedom, we introduce a weighted inverse kinematics controller, along with a null-space control strategy to optimize maneuverability and dexterity. By using a task-oriented weight matrix as a hyperparameter, the system provides a two-stage priority scheme fit for both large-scale movement and fine in-bore adjustments. In clinically relevant simulated scenarios, the system demonstrates a consistent 97% reachability rate across various human models. In addition, the task-oriented weight-matrix policy is extensively explored in five representative subtasks seen during needle biopsy through both simulation and real-world experiments, demonstrating superior tracking accuracy and enhanced manipulability for CT-guided procedures.
|
| |
| TuAT9 |
502 |
Large Language Models in Spatial and Temporal Decision Automation:
Representation, Reliability, and Real-World Utility |
Special Session |
| Organizer: Li, Ziyue | Technical University of Munich |
| Organizer: Zhao, Meng | Lehigh University |
| |
| 10:30-10:50, Paper TuAT9.1 | |
| SQLBench: A Comprehensive Evaluation for Text-To-SQL Capabilities of Large Language Models (I) |
|
| Zhang, Bin | Institute of Automation, Chinese Academy of Sciences |
| Ye, Yuxiao | Hong Kong University of Science and Technology |
| Du, Guoqing | SenseTime Research |
| Hu, Xiaoru | SenseTime Research |
| Li, Zhishuai | SenseTime Research |
| Liu, Chi Harold | Beijing Institute of Technology |
| Xu, Zhiwei | Shandong University |
| Fan, Guoliang | Institute of Automation, Chinese Academy of Sciences |
| Zhao, Rui | SenseTime Research |
| Li, Ziyue | Technical University of Munich |
| Mao, Hangyu | Institute of Microelectronics, Chinese Academy of Sciences |
Keywords: AI-Based Methods, Agent-Based Systems, Autonomous Agents
Abstract: Large Language Models (LLMs) have emerged as a powerful tool in advancing the Text-to-SQL task, significantly outperforming traditional methods.Nevertheless, as a nascent research field, there is still no consensus on the optimal prompt templates and design frameworks. Additionally, existing benchmarks inadequately explore the performance of LLMs across the various sub-tasks of the Text-to-SQL process, which hinders the assessment of LLMs' cognitive capabilities and the optimization of LLM-based solutions.To address the aforementioned issues, we firstly construct a new dataset designed to mitigate the risk of overfitting in LLMs. Then we formulate five evaluation tasks to comprehensively assess the performance of diverse methods across various LLMs throughout the Text-to-SQL process.Our study highlights the performance disparities among LLMs and proposes optimal in-context learning solutions tailored to each task. These findings offer valuable insights for facilitating the development of LLM-based Text-to-SQL systems.
|
| |
| 10:50-11:10, Paper TuAT9.2 | |
| Controlling Large Language Model-Based Agents for Large-Scale Decision-Making: An Actor-Critic Approach (I) |
|
| Zhang, Bin | Institute of Automation, Chinese Academy of Sciences |
| Mao, Hangyu | CAS |
| Ruan, Jingqing | Meituan |
| Wen, Ying | Shanghai Jiao Tong University |
| Li, Yang | Shanghai Jiao Tong Univ |
| Zhang, Shao | Shanghai Jiao Tong University |
| Xu, Zhiwei | Shandong University |
| Li, Dapeng | Li Auto Inc |
| Li, Ziyue | Technical University of Munich |
| Li, Lijuan | Institute of Automation, Chinese Academy of Sciences |
| Fan, Guoliang | Institute of Automation, Chinese Academy of Sciences |
Keywords: Agent-Based Systems, Autonomous Agents, Task Planning
Abstract: The remarkable progress in Large Language Models (LLMs) opens up new avenues for addressing planning and decision-making problems in Multi-Agent Systems (MAS). However, as the number of agents increases, the issues of hallucination in LLMs and coordination in MAS have become increasingly prominent. Additionally, the efficient utilization of tokens emerges as a critical consideration when employing LLMs to facilitate the interactions among a substantial number of agents. In this paper, we develop a modular framework called LLaMAC to mitigate these challenges. LLaMAC implements a value distribution encoding similar to that found in the human brain, utilizing internal and external feedback mechanisms to facilitate collaboration and iterative reasoning among its modules. Through evaluations involving system resource allocation and robot grid transportation, we demonstrate the considerable advantages afforded by our proposed approach.
|
| |
| 11:10-11:30, Paper TuAT9.3 | |
| Data Pre-Play for Incremental Fault Diagnosis in Industrial Process (I) |
|
| Zhang, Xuerui | Technical University of Munich |
| Zhuo, Yue | Cornell University |
| Niu, Haiming | The CHN Energy Zhishen Control Technology Co., Ltd |
| Zhang, Zhigang | The CHN Energy Zhishen Control Technology Co., Ltd |
| Qian, Jinchuan | Zhejiang University of Science and Technology |
| Li, Ziyue | Technical University of Munich |
| Song, Zhihuan | Zhejiang University |
| Zhang, Xinmin | Zhejiang University |
Keywords: Diagnosis and Prognostics, Failure Detection and Recovery, AI-Based Methods
Abstract: Novel faults arise sequentially in dynamically changing industrial conditions, and a practical data-driven fault diagnosis model should recognize new fault classes without forgetting old ones. This scenario becomes more challenging when the storage and computation resources are insufficient in industrial production. However, the study of incremental learning on fault diagnosis has received inadequate attention. To handle this issue, this work proposes a novel data pre-play method that aims at training a base model with high distinguishability. The proposed method not only simplifies the deployment of incremental updates but also gains improved representation capacities and reserved embedding space for new tasks to mitigate forgetting. The effectiveness of the proposed method was verified in an industrial benchmark process, and the application results demonstrate its superiority compared to other baseline methods.
|
| |
| 11:30-11:50, Paper TuAT9.4 | |
| Fishing for Answers: Exploring One-Shot vs. Iterative Retrieval Strategies for Retrieval Augmented Generation (I) |
|
| Huifeng, Lin | SenseTime Research |
| Liang, Jintao | Beijing University of Posts and Telecommunications |
| Wu, You | Hong Kong University of Science and Technology |
| Zhao, Rui | SenseTime Research |
| Zeng, Zhuoqi | Hainan Bielefeld University of Applied Sciences |
| Li, Ziyue | Technical University of Munich |
Keywords: Agent-Based Systems, Autonomous Agents, Task Planning
Abstract: Retrieval-Augmented Generation (RAG) based on Large Language Models (LLMs) is a powerful solution to understand and query the industry's closed-source documents. However, basic RAG often struggles with complex QA tasks in legal and regulatory domains, particularly when dealing with numerous government documents. The traditional top-k strategy frequently misses golden chunks, leading to incomplete or inaccurate answers. To address these retrieval bottlenecks, we explore two strategies to improve evidence coverage and answer quality. The first is a One-SHOT retrieval method that adaptively selects chunks based on a token budget, allowing as much relevant content as possible to be included within the LLM's context window. Additionally, we design modules to further filter and refine the chunks. The second is an iterative retrieval strategy built on a Reasoning Agentic RAG framework, where a reasoning LLM dynamically issues search queries, evaluates retrieved results, and progressively refines the context over multiple turns. We identify query drift and retrieval laziness issues and further design two modules to tackle them. Through extensive experiments on a dataset of government documents, we aim to offer practical insights and guidance for real-world applications in legal and regulatory domains.
|
| |
| 11:50-12:10, Paper TuAT9.5 | |
| One Image Is All You Need: Agentic One-Shot Image Generation Via Text-Based World Models for Long-Tail Spatial Perception (I) |
|
| Zeng, Keqin | Tsinghua University |
| Su, Shuting | SenseTime Research |
| Lin, Shihao | Sun Yat-Sen University |
| Li, Ziyue | Technical University of Munich, Heilbronn Data Science Center, Munich Data Science Institute |
| Zhao, Rui | SenseTime Research |
Keywords: Computer Vision in Automation, AI-Based Methods, Autonomous Vehicle Navigation
Abstract: Reliable spatial decision automation, such as autonomous driving and maritime surveillance, critically depends on robust visual perception. However, real-world spatiotemporal data exhibits severe heterogeneity, often manifesting as extreme long-tail distributions for safety-critical scenarios. This data scarcity induces dataset shifts that degrade detection performance and pose safety risks. While synthetic data generation offers a potential solution, existing generative approaches, such as diffusion models and Generative Adversarial Networks (GANs), often lack explicit spatial grounding and structural constraints, resulting in spatial and physical inconsistencies in generated scenes. To address these challenges, we introduce WMGen-v1, an agentic text-based world model framework for long-tail spatial data generation. WMGen-v1 employs a Large Vision-Language Model (LVLM) to extract spatial relations from a single reference image, while a Large Language Model (LLM) injects prior world knowledge to perform logical scene reasoning. Subsequently, conditioned on the structured semantic representations produced by this reasoning process, a diffusion model generates diverse and physically grounded long-tail training data. Experiments on internal industrial datasets, ROADWork, and LaRS benchmarks demonstrate that WMGen-v1 significantly outperforms baseline approaches. Notably, detectors trained solely on our synthetic data achieve performance comparable to those trained on real data, effectively mitigating long-tail data scarcity and improving the reliability of spatial automation systems.
|
| |
| 12:10-12:30, Paper TuAT9.6 | |
| Multi-Modal Contrastive Learning for Implicit Earth Embeddings Via Location Tying (I) |
|
| Hecht, Jonathan | Computational Methods Lab, HafenCity University Hamburg; Institute of Geodesy and Geoinformation, University of Bonn, Germany |
| Arzoumanidis, Lukas | Computational Methods Lab, HafenCity University Hamburg |
| Li, Ziyue | Dept. of Operations & Technology, Technical University of Munich; Heilbronn Data Science Center; Munich Data Science Institute |
| Dehbi, Youness | Computational Methods Lab, HafenCity University Hamburg |
Keywords: Machine learning, Data fusion, Environment Monitoring and Management
Abstract: Spatial prediction tasks are often limited by a lack of high-quality labelled ground-truth observations. To overcome this challenge, self-supervised pre-training is a possible solution, with contrastive learning dominant for location encoders. Those approaches usually align geographic coordinates with just one additional modality. We propose two multimodal contrastive learning architectures: Multimodal Embedding via Location Tying (MELT) and Sequential Alternating Location Training (SALT). These architectures expand this framework beyond two modalities by utilising unpaired geospatial data. Both methods are technically viable and match the performance of the strongest two-modality baseline (SATCLIP) across four downstream tasks. However, increasing the number of modalities does not consistently improve performance, suggesting that the chosen location encoder is the main limitation—the contrastive objective reaches its peak early, regardless of modality diversity or pre-training volume. MELT provides more stable training than SALT and presents a stronger foundation for future scaling.
|
| |
| TuAT10 |
Convention Hall A |
| Best Paper Award Session 1 |
|
| |
| 10:30-10:50, Paper TuAT10.1 | |
| Scalable Production Scheduling: Linear Complexity Via Unified Homogeneous Graphs |
|
| Hoss, Jonathan | Rosenheim University of Applied Sciences |
| Link, Moritz | TH Rosenheim |
| Klarmann, Noah | Rosenheim University of Applied Sciences |
Keywords: Planning, Scheduling and Coordination, Reinforcement, Intelligent and Flexible Manufacturing
Abstract: Efficiently solving the Job Shop Scheduling Problem in real-world industrial applications requires policies that are both computationally lean and topologically robust. While Reinforcement Learning has shown potential in automating dispatching rules, existing models often struggle with a scalability bottleneck caused by quadratic graph complexity or the architectural overhead of heterogeneous layers. We introduce a unified graph framework that employs feature-based homogenization to project distinct node roles into a shared latent space. This allows a standard homogeneous Graph Isomorphism Network to capture complex resource contention with linear complexity, ensuring low-latency inference for large-scale industrial applications. Our empirical results demonstrate that our framework achieves state-of-the-art performance while exhibiting consistent zero-shot generalization. We identify the job-to-machine ratio as the primary driver of policy effectiveness, rather than absolute problem size. Based on this, we propose a hypothesis of structural saturation, demonstrating that policies trained on critically congested instances learn scale-invariant resolution strategies. Agents trained at this saturation point internalize invariant conflict-resolution logic, allowing them to treat massive rectangular instances as a sequential concatenation of saturated sub-problems. This approach eliminates the need for expensive scale-specific retraining and prevents overfitting to statistical shortcuts, providing a robust and efficient pathway for deploying RL solutions in dynamic production environments.
|
| |
| 10:50-11:10, Paper TuAT10.2 | |
| XFlowMP: Task-Conditioned Motion Fields for Generative Robot Planning with Schrödinger Bridges |
|
| Nguyen, Khang | Mohamed Bin Zayed University of Artificial Intelligence |
| Vu, Minh Nhat | VinUni |
Keywords: AI-Based Methods, Machine learning
Abstract: Generative robotic motion planning requires not only the synthesis of smooth and collision-free trajectories but also feasibility across diverse tasks and dynamic constraints. Prior planning methods, both traditional and generative, often struggle to incorporate high-level semantics with low-level constraints, especially the nexus between task configurations and motion controllability. In this work, we present XFlowMP, a task-conditioned generative motion planner that models robot trajectory evolution as entropic flows bridging stochastic noises and expert demonstrations via Schrodinger bridges given the inquiry task configuration. Specifically, our method leverages Schrodinger bridges as a conditional flow matching coupled with a score function to learn motion fields with high-order dynamics while encoding start-goal configurations, enabling the generation of collision-free and dynamically-feasible motions. Through evaluations, XFlowMP achieves up to 53.79% lower maximum mean discrepancy, 36.36% smoother motions, and 39.88% lower energy consumption when compared to the next-best baseline on the RobotPointMass benchmark and also reduces short-horizon planning time by 11.72%. On long-horizon motions in the LASA Handwriting dataset, our method maintains the trajectories with 1.26% lower maximum mean discrepancy, 3.96% smoother, and 31.97% lower energy. We further demonstrate the practicality of our method on the Kinova Gen3 manipulator, executing planning motions and confirming its robustness in real-world settings.
|
| |
| 11:10-11:30, Paper TuAT10.3 | |
| Zero-Shot Generalization from Motion Demonstrations to New Tasks |
|
| Freitag, Kilian Tamino | Chalmers University of Technology |
| Combrink, Alvin | Chalmers University of Technology |
| Figueroa, Nadia | University of Pennsylvania |
Keywords: Learning and Adaptive Systems, Motion and Path Planning, Machine learning
Abstract: Learning motion policies from expert demonstrations is an essential paradigm in modern robotics. While end-to-end models aim for broad generalization, they require large datasets and computationally heavy inference. Conversely, learning dynamical systems (DS) provides fast, reactive, and provably stable control from very few demonstrations. However, existing DS learning methods typically model isolated tasks and struggle to reuse demonstrations for novel behaviors. In this work, we formalize the problem of combining isolated demonstrations within a shared workspace to enable generalization to unseen tasks. The Gaussian Graph is introduced, which reinterprets spatial components of learned motion primitives as discrete vertices with connections to one another. This formulation allows us to bridge continuous control with discrete graph search. We propose two frameworks leveraging this graph: Stitching, for constructing time-invariant DSs, and Chaining, giving a sequence-based DS for complex motions while retaining convergence guarantees. Simulations and real-robot experiments show that these methods successfully generalize to new tasks where baseline methods fail.
|
| |
| 11:30-11:50, Paper TuAT10.4 | |
| Pareto-Guided Learning of Ordered Structures for Multi-Occupant Human-In-The-Loop HVAC Control from Behavior-Derived Feedback |
|
| Wu, Hongyi | Tsinghua University |
| Chen, Xi | Tsinghua University |
| Guan, Xiaohong | Xi'an Jiaotong University |
Keywords: Energy and Environment-aware Automation, Human Factors and Human-in-the-Loop, Reinforcement
Abstract: Smart HVAC control in multi-occupant environments presents a significant challenge in balancing conflicting individual preferences while maintaining collective comfort. While Human-in-the-Loop strategies have shown promise by incorporating real-time feedback, traditional methods often rely on explicit voting or are limited to settings with a single thermal preference profile. This paper proposes a Pareto-guided extension to implicit Q-learning under behavior-derived feedback that incorporates the structural information of human thermal preferences. The framework integrates reward correction and state augmentation to enable effective offline policy learning, allowing the agent to receive explicit directional guidance in uncomfortable regions. Extensive experiments under various conditions validate our approach through comprehensive simulations, with a mean relative improvement of 25.9% in absolute thermal discomfort compared to pure implicit Q-learning, demonstrating superior performance in maximizing occupant comfort.
|
| |
| 11:30-11:50, Paper TuAT10.4 | |
| Design and Implementation of a Human-Skin-Inspired Tactile Fingertip |
|
| Shen, Yike | Zhejiang University |
| Bao, Junyi | Zhejiang University |
| Qi, Yanhui | Zhejiang University |
| Wang, Ze | Zhejiang University |
| Chen, Jiming | Zhejiang University |
| Li, Gaofeng | Zhejiang University |
Keywords: Force and Tactile Sensing, Biomimetics, Sensor-based Control
Abstract: Tactile sensing is fundamental to dexterous robotic manipulation and human–robot interaction. However, existing tactile sensors face significant challenges in achieving human-skin-like perception over large areas with high resolution, posing a critical bottleneck in robotics. In this work, we developed a super-resolution biomimetic tactile fingertip featuring a curved sensing surface, extending previous planar-based techniques to complex 3D geometries. The system integrates 16 highly sensitive force sensors strategically distributed on a hemispherical-cylindrical skeleton—mimicking the dimensions of an Allegro Hand fingertip, and is encapsulated in a soft silicone layer to ensure compliant contact. We constructed a custom-built 5-DoF force measurement platform to precisely calibrate individual sensors and acquire ground-truth normal force data. Furthermore, we designed a 3D tactile reconstruction algorithm capable of decoding spatial contact force distributions from sparse sensor readings. Experimental results demonstrate that the system achieves high-precision localization across complex 3D geometries. By leveraging the bio-inspired mechanical diffusion within the soft layer, our algorithm yields a mean positional accuracy of approximately 2.7 mm. Compared to the average 8.5 mm physical pitch of the sparse sensor array, this represents a 4-fold enhancement in spatial resolution, effectively bridging the gap between sparse sensing hardware and high-density tactile perception.
|
| |
| 12:10-12:30, Paper TuAT10.6 | |
| Evolutionary Pseudo-Label Policy Search for Semi-Supervised Hot Rolling Quality Prediction |
|
| An, Jiqing | Northeastern University |
| Xu, Te | Northeastern University of China |
| Wu, Jian | Northeastern University |
Keywords: Zero-Defect Manufacturing, Intelligent and Flexible Manufacturing, Machine learning
Abstract: Reliable labels for hot-rolled steel properties come from destructive tests, so they are sparse, costly, and delayed, whereas process measurements are abundant. Using those unlabeled records in industrial tabular regression is difficult because multi-grade production introduces heterogeneity in both the feature and target spaces and the production stream also drifts over time. Across tensile strength, yield strength, and elongation, fixed pseudo-label heuristics behave differently from one target and label budget to another. We examine the problem through a Data Analytics and Optimization (DAO) lens. On the data-analytics side, pseudo-labeling becomes markedly more stable once features are aligned using the chronological training-period feature pool and once teacher and student are warm-started in a purely supervised stage. On the optimization side, we introduce Evolutionary Pseudo-Label Policy Search for Regression (EPPSR), which uses multi-objective genetic programming to evolve interpretable dual-tree policies for pseudo-label admission and weighting from six dynamic training signals. Experiments on 42,982 coils from 78 steel grades show that no fixed heuristic dominates across targets, whereas EPPSR is the strongest SSL method on the multi-grade benchmark and remains close to the best supervised baseline. Single-grade studies further show that adaptive pseudo-labeling matters most when labels are extremely scarce and heterogeneity is stronger.
|
| |
| TuBT1 |
201 |
| Data Analytics and Optimization for Smart Industry 2 |
Regular Session |
| |
| 14:00-14:20, Paper TuBT1.1 | |
| Protein Secondary Structure Prediction Based on Improved Generative Adversarial Networks |
|
| Li, Fei | Northeastern University |
| Xu, Meiling | Northeastern University |
Keywords: AI-Based Methods, Machine learning, AI and Machine Learning in Healthcare
Abstract: Predicting protein secondary structure is a fundamental task in computational biology, providing an important bridge from amino-acid sequence to higher-level structural organization. However, accurate Q8 prediction remains challenging because residue dependencies exist at multiple scales and are not easy to capture within a single modeling framework. To address this issue, we develop an improved generative adversarial network for protein secondary structure prediction. A MultiScale-Block extracts sequence patterns under different receptive fields, while an EnhancedResBlock further refines features and strengthens the modeling of long-range residue interactions. On this basis, we design OPSO-GAN, which incorporates an optimized particle swarm optimization strategy into GAN training through a two-level scheme: a super-swarm searches global hyperparameters, and a sub-swarm refines local model parameters. Preliminary experiments on CB513, CASP10, and CASP11 indicate stable and competitive Q8 prediction performance, and initial ablation results suggest that both the architectural modifications and the hierarchical OPSO strategy contribute to the observed improvement.
|
| |
| 14:20-14:40, Paper TuBT1.2 | |
| Deep Switching Koopman Networks: Discovery of Switching Dynamic Regimes in Industrial Systems |
|
| Zhu, Zhenyi | The Hong Kong University of Science and Technology (Guangzhou) |
| Zhang, Lingzheng | The Hong Kong University of Science and Technology (Guangzhou) |
| Shao, Chen | Karlsruhe Institute of Technology |
| Wang, Wenjia | National University of Singapore |
| Tsung, Fugee | HKUST |
Keywords: AI-Based Methods, Machine learning, Cyber-physical Production Systems and Industry 4.0
Abstract: Reliable modeling of non-stationary multivariate time series, particularly those arising in industrial monitoring, requires capturing the dynamics that govern sequential observations. Existing CNN- and Transformer-based classifiers often rely on stationary discriminative patterns and therefore do not explicitly represent transitions among operating regimes. To address this limitation, we introduce the Deep Switching Koopman Network (DSKN), which integrates Koopman operator theory with a Mixture of Experts (MoE) architecture to represent switching nonlinear dynamics. DSKN lifts observations into a latent space and combines two core designs: a spectrum-aware router that identifies frequency-domain regime signatures, and a Koopman Expert Bank that approximates local linear evolution. A composite objective regularizes the latent trajectories through classification, reconstruction, and dynamical consistency losses. Empirical evaluations on representative UEA multivariate benchmarks show that DSKN achieves strong classification performance while providing operator-theoretic interpretability through learned Koopman spectra and transparent transition logic. The source code will be released on GitHub.
|
| |
| 14:40-15:00, Paper TuBT1.3 | |
| Data Analytics-Based Evolutionary Algorithm for Structure Topology Optimization with Graph-Based Encoding |
|
| Pei, Hongxu | Northeastern University |
| Tang, Lixin | Northeastern University |
| Wang, Xianpeng | Northeastern University |
| Su, Lijie | Northeastern University |
Keywords: AI-Based Methods, Machine learning, Data fusion
Abstract: This paper proposes a novel graph-based encoding and population initialization strategy tailored for structural topology optimization. While gradient-based methods are widely employed, their practical utility is often hindered by high sensitivity to hyperparameters, the persistence of intermediate densities, and a propensity for stagnating in local optima. Despite the emergence of non-gradient algorithms, research that intrinsically integrates fundamental topology optimization principles remains scarce. The encoding strategy presented herein provides an explicit representation of structural topological features, which, when integrated into a Differential Evolution framework, enables direct structural search at the topological level. Furthermore, Delaunay triangulation is embedded into the initialization and decoding phases to enhance initial solution quality. Additionally, clustering analysis is utilized to preserve the multi-modal characteristics of the optimization process. Experimental results demonstrate that the proposed methodology effectively reduces invalid crossover and mutation operations, thereby facilitating the discovery of high-potential configurations and significantly improving global search efficiency.
|
| |
| 15:00-15:20, Paper TuBT1.4 | |
| Multi-Objective Optimization of BP Neural Networks for Wafer Quality Prediction Using NSGA-II |
|
| Ma, Kerong | Northeastern University |
| Song, Xiangman | Northeastern University |
| Xu, Te | Northeastern University of China |
| Wang, Kun | Northeastern University |
| Liu, Guoyuan | Northeastern University |
Keywords: AI-Based Methods, Machine learning, Semiconductor Manufacturing
Abstract: As the complexity of semiconductor fabrication continues to increase, wafer production now generates extensive process-monitoring data. However, these data commonly include numerous variables, duplicated information, and weak model transparency. To handle these problems, this paper builds a wafer-quality prediction framework in which NSGA-II is used to tune a BP neural network. Feature screening is first carried out to compress the input space. NSGA-II then searches the network architecture and major training parameters. Finally, SHAP analysis explains the trained model and extracts the process variables most related to wafer quality. The experimental results indicate that this method improves prediction accuracy and provides interpretable support for semiconductor process optimization.
|
| |
| 15:20-15:40, Paper TuBT1.5 | |
| An Improved SA Algorithm Based on the Symmetric Group for Permutation-Related Combinatorial Optimization Problems |
|
| Yao, Quanzheng | National Frontiers Science Center for Industrial Intelligence and Systems Optimization, Northeastern University |
| Dong, Yun | Liaoning Key Laboratory of Manufacturing System and Logistics Optimization |
| Zhao, Ren | Liaoning Engineering Laboratory of Data Analytics and Optimization for Smart Industry |
| Sun, Defeng | National Frontiers Science Center for Industrial Intelligence and Systems Optimization, Northeastern University |
Keywords: AI-Based Methods, Swarms, Optimization and Optimal Control
Abstract: In this paper, an improved simulated annealing (SA) algorithm based on the symmetric group (G-SA) is proposed. The algorithm can solve the permutation-related combinatorial optimization problems (COPs) with symmetric characteristics, such as traveling salesman problem (TSP). The algorithm partitions the permutation solution space into even and odd subspaces. Based on this partition, the initialization and neighborhood search strategies of G-SA are designed. As a result, G-SA searches for better solutions within only half of the solution space. This design improves the optimization efficiency of the algorithm. In the experiment, we select 15 TSP benchmark instances and use the classical SA as the baseline algorithm to evaluate the performance of G-SA. The experimental results show that the proposed G-SA achieves superior performance compared with the classical SA algorithm.
|
| |
| 15:40-16:00, Paper TuBT1.6 | |
| A Deep Reinforcement Learning-Based Adaptive Fireworks Algorithm for Spatiotemporal Resource-Constrained Aircraft Pulsating Assembly Line Balancing |
|
| Zhang, Zhihui | Tongji University |
| Qiao, Fei | Tongji University |
| Liu, Juan | Tongji University |
| Li, Wei | China Changan Automobile Group Co., Ltd |
| Ma, Yumin | Tongji University |
| Ai, Jiakang | Tongji University |
| Shi, Weina | China Changan Automobile Group Co., Ltd |
| Huang, Xinyu | Tongji University |
| Chen, Lingtao | Tongji University |
Keywords: Assembly, Hybrid Strategy of Intelligent Manufacturing, Planning, Scheduling and Coordination
Abstract: This paper addresses the highly complex aircraft pulsating assembly line balancing problem, strictly restricted by extreme spatiotemporal constraints, including rigid precedence for parallel operations, stringent limits on the total number of specialized workers and narrow physical area capacities. Traditional exact solvers suffer from combinatorial explosion when facing such constraints. Meanwhile, conventional meta-heuristics frequently struggle with feasibility due to blind searches, and existing reinforcement learning methods relying on offline pre-training tend to suffer from generalization bottlenecks when confronted with unknown topological structures. To overcome these challenges, a Meta-Controlled Online Adaptive Fireworks Algorithm (MCOA-FWA) is proposed. This online framework utilizes a deep Q-network as a meta-controller to dynamically schedule heterogeneous topological actions without offline pre-training. Furthermore, a Partition-Guided Spatiotemporal Collaborative Decoder (PG-STCD) based on dual-chromosome encoding is developed to decouple spatiotemporal conflicts. Extensive experiments on real-world industrial instances with up to 3182 operations demonstrate that MCOA-FWA effectively avoids deadlocks and guarantees physical feasibility. Compared to baseline algorithms, the proposed method achieves a significant performance improvement in takt time compression and robustness under tight computational budgets.
|
| |
| TuBT2 |
202 |
| Frontier Technology of Industrial & Systems Engineering 2 |
Regular Session |
| |
| 14:00-14:20, Paper TuBT2.1 | |
| Quay Crane Scheduling at an Indented Berth under an Alternative Unidirectional Movement Strategy |
|
| Chang, Yiran | Northeastern University |
| Sun, Defeng | Northeastern University |
Keywords: Logistics
Abstract: With the growth of containerized trade and the deployment of larger vessels, improving seaside operational efficiency has become increasingly important for container terminals. Indented berths enhance handling capacity by allowing quay cranes to operate on both sides of a vessel, but also introduce more complex interference relationships. This paper studies the quay crane scheduling problem under an indented berth configuration with a unidirectional movement strategy. A mixed-integer programming formulation is developed, integrating assignment, movement, reachability, and interference decisions. The proposed formulation strengthens cross-side safety constraints, introduces an effective-time representation, and removes symmetric solutions. Computational results show that the model can efficiently obtain optimal solutions and significantly improve computational performance compared with the baseline formulation.
|
| |
| 14:20-14:40, Paper TuBT2.2 | |
| Efficient Relaxation Models and Algorithms for Multi-Plant Production-Delivery Problems |
|
| Li, Feng | Huazhong University of Science and Technology |
| Yang, Xianyan | Wuhan University of Science and Technology |
| Deng, Zhenhao | Huazhong University of Science and Technology |
Keywords: Logistics, Inventory Management, Intelligent Transportation Systems
Abstract: We study an integrated production and delivery problem arising in make-to-order manufacturing with multiple plants and third-party logistics providers offering multiple shipping modes. Orders carry committed delivery dates, production is capacitated at the plant-day level, and transportation costs may be nonlinear in both transit time and shipment quantity. We first exploit the cost structure to establish a compact integer programming model in which every shipment uses the slowest feasible mode. We then characterize complexity for two shipping-cost regimes: the linear case is solvable in polynomial time, whereas the nonlinear case is strongly NP-hard. To obtain tight bounds, we develop three Dantzig-Wolfe reformulations based on plant patterns, customer patterns, and an extended formulation that couples the two via production-consistency and shipment-consistency linking constraints, and we prove that the extended formulation dominates the other two formulations at the LP-relaxation level. Computational experiments show that the extended formulation yields substantially stronger LP relaxation bounds and that the associated column-generation heuristic produces high-quality solutions.
|
| |
| 14:40-15:00, Paper TuBT2.3 | |
| TERRAN: A Transformer-Based Electric Vehicle Routing Agent for Real-Time Adaptive Navigation (I) |
|
| Tang, Maojie | University of California, Riverside |
| Yu, Nanpeng | University of California, Riverside |
| Karamouzas, Ioannis | University of California, Riverside |
| Ye, Zuzhao | University of California, Riverside |
Keywords: Logistics, Plug-in Electric Vehicles, Reinforcement
Abstract: The Electric Vehicle Routing Problem with Time Windows (EVRP-TW) poses significant challenges for sustainable logistics due to its tight coupling of spatial, temporal, and energy constraints. Classical optimization methods face trade-offs: exact solvers like CPLEX ensure optimality but require prohibitive runtimes, while metaheuristics like Variable Neighborhood Search struggle with feasibility under complex constraints. We propose TERRAN, a transformer-based reinforcement learning framework for real-time and scalable EVRP-TW optimization. TERRAN integrates three key components: 1) Future-Feasibility Pruning (FFP), which proactively eliminates energy-infeasible actions by verifying reachability to charging stations or depots before each move; 2) Staged Reward Scheduling, which progres-sively transitions from dense auxiliary signals to task-aligned rewards to guide the agent from achieving feasibility to mini- mizing cost; and 3) An End-to-End Transformer-Based RL Agent Tailored for EVRP-TW, which directly integrates EV-specific constraints—including battery consumption, charging decisions, and delivery time windows—into the policy network and decoding process, enabling unified, post-processing-free optimization across varying instance scales. Experiments on Solomon benchmark instances with 5–100 customers demonstrate that TERRAN achieves 100% feasibility across all problem scales. It matches CPLEX’s optimality on 5–customer instances, achieves up to 170,000× speedups with solutions within 1.5% of optimal on 15–customer instances, and delivers feasible solutions for 100–customer instances in 0.47 s, where CPLEX fails on over 80% of cases within 1 hour. These results establish TERRAN as a practical and scalable solution for real-time electric vehicle routing
|
| |
| 15:00-15:20, Paper TuBT2.4 | |
| A Generalized Parallel Sampling Attention Model for Real-Time Planning in Warehouse Mobile Manipulation Robots |
|
| Wei, Jiawei | Fudan University |
| Zhu, Ruilin | Hunan University of Finance and Economics |
| Liu, Caitian | Fudan University |
| Ren, Pingye | Fudan University |
| Chen, Xiong | Fudan University |
Keywords: Motion and Path Planning, Inventory Management, Optimization and Optimal Control
Abstract: Driven by the demand for high-density storage and high-frequency warehouse operations, mobile manipulation robots are increasingly adopted for inventory inspection tasks. However, under continuous chassis motion, dense and partially observable task points make real-time planning computationally demanding, especially when joint configuration choices strongly affect transition costs. To address this, we propose G-PSAM, a Generalized Parallel Sampling Attention Model that performs cross-level encoding and redundant decoding to jointly optimize task sequencing and configuration assignment. We further develop an efficient post-processing scheme that hierarchically searches and recomposes redundant network outputs, achieving a better trade-off between runtime and solution quality. Building on G-PSAM, we enhance the Redundant Cooperative Control (RCC) strategy to improve arm-chassis parallelism and reduce idle time during operation. Simulation and real-robot experiments demonstrate that our method outperforms neural optimization baselines, while the enhanced RCC strategy achieves robust and efficient on-robot execution.
|
| |
| 15:20-15:40, Paper TuBT2.5 | |
| Optimal Switching Control of Production–Inventory Systems between Production and Setup |
|
| Lv, Yuanda | Northeastern Univercity |
| Wang, Gongshu | Northeastern University |
Keywords: Optimization and Optimal Control, Inventory Management
Abstract: This paper studies the optimal inventory control of manufacturing systems operating in a dual-cycle production mode. Firstly, we postulate that the production rate exhibits periodic variations and each period is composed of two production cycles. Secondly, we characterize such a production-inventory system as a time-dependent switched system and describe the system by using dynamic equations, presenting the objective function to be optimized - a linear quadratic performance index functional. Thirdly, we prove mathematically that the optimal production strategy within a period can be derived by finding a proper stock level at the end of the first cycle and solving the common optimal solution separately over the consecutive two cycles. Finally, two production examples are presented to demonstrate validity of the proposed theoretical results.
|
| |
| 15:40-16:00, Paper TuBT2.6 | |
| A Column Generation Based Diving Algorithm for Slab Assignment Problem Considering Order Completeness |
|
| Zhang, Yuxuan | The National Frontiers Science Center for Industrial Intelligence and Systems Optimization, Northeastern University, China |
| Meng, Ying | Northeastern University |
| Tang, Lixin | Northeastern University |
Keywords: Optimization and Optimal Control, Planning, Scheduling and Coordination, Intelligent and Flexible Manufacturing
Abstract: We address the slab assignment problem arising from the steelmaking industry. A Slab is the first solid intermediate product in the steel production process, both as an output of steelmaking and as a raw material for hot rolling. In steel plants, a common operational task involves assigning slabs to customer orders as raw material for subsequent production. Slabs that are not linked to any specific order are referred to as unassigned slabs. During production, a large number of unassigned slabs are inevitably generated, leading to significant material waste and high inventory costs. Therefore, it is essential to assign them to unfilled orders, in order to enhance slab utilization and reduce overall production costs. In this paper, we investigate a slab assignment problem that maximizes slab utilization and customer satisfaction. To characterize the problem, a binary integer linear programming model is formulated. To obtain high-quality solutions effectively, we develop a diving heuristic based on column generation, accelerated by a column generation stabilization strategy. The performance of the proposed algorithm is validated through computational experiments on slab assignment instances. The results show that the proposed algorithm outperforms the general MIP solvers for both small and large scale problems.
|
| |
| TuBT3 |
203 |
| Systems Technology of Industrial Automation and Control 2 |
Regular Session |
| |
| 14:00-14:20, Paper TuBT3.1 | |
| Design of a Dual-Biped Climbing Robot System with Modular Grippers for Operation on Curved Pipeline Surfaces |
|
| Li, Zikang | University of Macau |
| Yin, Yuan | University of Macau |
| Zhang, Weijian | University of Macau |
| Xu, Qingsong | University of Macau |
Keywords: Mechanism Design in Meso, Micro and Nano Scale, Mechatronics in Meso, Micro and Nano Scale
Abstract: Wall-climbing robots enhance safety and efficiency in high-altitude environments like buildings, pipelines, and ships. However, most existing systems are single-robot platforms limited to simple tasks, lacking the capability for complex manipulation. This paper presents a novel Dual-Biped Curved-surface Climbing Robot (DCCR) system to address this gap. The system integrates two distinct biped robots: one with 5 degrees of freedom (5-DOF) and one with 3 degrees of freedom (3-DOF). Both robots are equipped with adaptive suction modules that feature variable-curvature design, enabling stable attachment and locomotion on both flat and curved surfaces. A core innovation is the use of detachable modular grippers. This allows the DCCR system to transition from a mobile locomotion mode to a stationary dual-manipulator configuration, thereby enabling complex object manipulation. The design of the robots, suction modules, and grippers is presented. Experimental results demonstrate the system’s performance in complex and coordinated operations, including object gripping and tool handling tasks on pipeline surfaces, validating its potential for cooperative manipulation applications.
|
| |
| 14:20-14:40, Paper TuBT3.2 | |
| Adaptive Method to Reduce Needle Tip Positioning Error During Reinsertion of Ultrafine Needles Considering Tissue Elasticity |
|
| Takemura, Koki | Waseda University |
| Tamura, Shotaro | Waseda University |
| Li, Yibai | Waseda University |
| Iwata, Hiroyasu | Waseda University |
Keywords: Medical Robots and Systems, Modelling, Simulation and Optimization in Healthcare, Sensor-based Control
Abstract: Computed tomography guided robotic needle insertion using fine-gauge needles offers a minimally invasive approach to interventional treatments. However, positional deviations of the needle tip from the target frequently necessitate reinsertion, which typically requires complete needle withdrawal, thereby increasing procedure time and patient burden. To address this issue, we integrated an automated reinsertion strategy into a robotic system that utilizes intratissue bending without full withdrawal. Fine-gauge needles are prone to deflect due to tissue resistance; thus, we developed a needle deformation model that predicts tip displacement by considering these interaction forces. By using a force sensor to monitor the forces during the bending process, the model calculates correction amounts that are adaptive to variations in tissue elasticity. An experimental validation using two PVC phantoms with distinct Young’s moduli demonstrated that a 5.0 mm target correction was consistently achieved within a 3.0 mm tolerance. These results demonstrate the feasibility of the force-sensor-based model for small automated reinsertion tasks, while its applicability to larger corrections and stiffer tissues remains limited.
|
| |
| 14:40-15:00, Paper TuBT3.3 | |
| Sensor-Based Reward Learning from Video Labels for Tumble Motion Control in a Household Dryer |
|
| Lee, Jinwoo | Seoul National University |
| Kang, Chanseok | LG Electronics |
| Bae, Guntae | LG Electronics |
Keywords: Model Learning for Control, Sensor-based Control
Abstract: Household dryers aim to maintain cataracting laundry motion for efficient drying. Applying reinforcement learning to this task in hindered by the challenge of defining reward function which represents internal motion quality. We propose a reinforcement learning framework that utilize learning-based reward function to improve drying efficiency. We utilize synchronized video with sensor data to train the reward function that maps sensor data to motion quality and evaluate the policy trained with the learning-based reward in multiple laundry load compositions. We demonstrated reinforcement learning in dryers with learning-based reward outperforms expert-designed baseline by 2.04% average and 2.86% best-case in drying performance. The results show that visual supervision is successfully transferred to a deployable sensor-only reward model, as well as the potential for applying reinforcement learning on real-world home appliances.
|
| |
| 15:00-15:20, Paper TuBT3.4 | |
| An Imaging-Constrained Path Planning Pipeline for Fine Robotic Visual Inspection of Rotary Parts |
|
| Diamond, Sophie | École De Technologie Supérieure |
| Tovar, Ricardo | École De Technologie Supérieure ÉTS |
| Dupoiron, Guillaume | École De Technologie Supérieure |
| Roberge, Jean-Philippe | École De Technologie Supérieure |
Keywords: Motion and Path Planning
Abstract: Visual surface inspection is a critical quality-assurance step for high-value aerospace hardware, yet the detection of barely perceptible defects (for example, subtle marks remaining after shot-peening) is still largely performed manually. A key reason is that defect contrast is highly dependent on illumination and viewing directions: the same flaw may be invisible under most angles and only emerge under a narrow range of light incidence. This paper presents an inspection-planning pipeline designed to automate such angle-sensitive inspections on rotary landing-gear components. The method first generates camera viewpoints that intentionally favor geometries known to enhance the visibility of low-contrast surface anomalies, then converts these viewpoints into a feasible robotic inspection trajectory. Planning is performed for a 7-DOF inspection cell composed of a 6-DOF serial robot coupled with a 1-DOF external rotational axis, allowing the system to exploit part reorientation while maintaining kinematic feasibility and practical reachability constraints. By integrating visibility-driven viewpoint generation with robot-plus-positioner path planning, the proposed approach provides a complete, application-oriented solution toward reliable and repeatable inspection of lightly textured aerospace surfaces.
|
| |
| 15:20-15:40, Paper TuBT3.5 | |
| Active Gaussian Splatting for Autonomous Reconstruction Using D-Optimal Next Best View Planning |
|
| Prashar, Kanav | Arizona State University |
| Torre, Carlos | Arizona State University |
| Staggers Jr, Rodney | Arizona State University |
| Vedantha Desikan, Bharath | Arizona State University |
| Das, Jnaneshwar | Arizona State University |
Keywords: Motion and Path Planning, Autonomous Agents, Reactive and Sensor-Based Planning
Abstract: Autonomous 3D reconstruction of geological features is a critical capability for planetary surface exploration, where communication latency prohibits manual view selection. Conventional approaches survey the environment exhaustively before selecting informative viewpoints, incurring unnecessary observation overhead. We propose an active Gaussian Splatting framework in which a drone autonomously selects the Next Best View (NBV) to maximise reconstruction quality of a target rock. Our approach assumes that an initial survey yields a high-fidelity point cloud of the target, from which a mesh model is derived and used to initialise 100{,}000 3D Gaussians on the surface. Gaussian positions are frozen at mesh-surface locations, and only appearance parameters --- colour, scale, rotation, and opacity --- are learned online from incrementally acquired images. Four evenly-spaced seed images initialise the model, after which, at each step, a diagonal Fisher Information Hessian is accumulated over the current training set and the candidate viewpoint that maximises the D-optimal information gain is selected as the next observation. We further provide a theoretical justification showing that convergence between steps is a emph{mathematical necessity} for the Taylor linearisation underlying the Fisher criterion, not merely an engineering heuristic.
|
| |
| 15:40-16:00, Paper TuBT3.6 | |
| Optimal Constrained Trajectory Planning for Redundant Manipulators Via Volumetric Obstacles Modeling and Analytical Gradient Fields |
|
| Mastromarino, Fabio | Politecnico Di Bari |
| Carli, Raffaele | Politecnico Di Bari |
| Longo, Nicola | COMAU SpA |
| Dotoli, Mariagrazia | Politecnico Di Bari |
Keywords: Motion and Path Planning, Collision Avoidance, Industrial and Service Robotics
Abstract: Trajectory planning for redundant manipulators in precision manufacturing tasks, such as continuous dispensing or automated inspection, requires the simultaneous satisfaction of strict operational constraints and real-time obstacle avoidance. Traditional planning approaches struggle in these scenarios: dense meshes are computationally demanding, while spherical approximations are overly conservative for elongated tools in cluttered spaces. To overcome these limitations, this paper proposes a computationally efficient trajectory planning framework implemented in ROS 2. By dynamically modeling the workspace with 3D Oriented Bounding Boxes and the manipulator as a kinematic chain of spatial cylinders, the framework achieves a tight, non-conservative geometric approximation. This novel pairing enables a closed-form distance formulation, yielding exact and continuous gradients for real-time null-space redundancy resolution without numerical differentiation. Furthermore, to ensure feasibility in fully constrained scenarios, a Dynamic Task Relaxation mechanism is introduced to temporarily induce artificial redundancy. The planning strategy further separates free-space motion from a constrained Cartesian execution phase, ensuring strict compliance with process requirements. The effectiveness of the proposed framework is validated through a representative industrial case study featuring both natively redundant and non-redundant robotic platforms, demonstrating improved computational efficiency, enhanced task reliability, and robust performance in highly constrained environments.
|
| |
| TuBT4 |
204 |
| Enabling Technology of Industrial Intelligence 2 |
Regular Session |
| |
| 14:00-14:20, Paper TuBT4.1 | |
| Shedding Light on Walls: Photometric Stereo for Autonomous Defect Detection |
|
| Lovato, Alessio | Istituto Italiano Di Tecnologia |
| Romiti, Edoardo | Istituto Italiano Di Tecnologia |
| Muratore, Luca | Istituto Italiano Di Tecnologia |
| Tsagarakis, Nikos | Istituto Italiano Di Tecnologia |
Keywords: Automation in Construction, Computer Vision in Automation
Abstract: Wall finishing is a multi-stage construction process involving putty application, spreading, sanding, and painting, each presenting distinct automation challenges. While robotic solutions for spray painting and putty application are already commercially available, sanding remains insufficiently addressed: identifying which regions require intervention demands surface inspection, yet defects on dried putty are difficult to perceive due to the near-textureless, low-contrast appearance of the surface. Inspired by the way human operators often use illumination to reveal surface irregularities and identify areas that require sanding, we propose a defect-detection approach based on photometric stereo, which uses four images acquired under different illumination directions to recover geometric information about the wall surface. The recovered geometric cues including the mean curvature, relative height, and albedo are encoded into a standard three-channel image that can be processed directly by off-the-shelf detection models such as YOLO and Mask R-CNN, without architectural modification, allowing them to exploit surface geometry rather than relying only on RGB appearance. Experimental results on both the training dataset and additional images outside the training distribution show that the proposed preprocessing improves defect detection performance over a single-image RGB baseline, increasing both precision and recall for Mask R-CNN and substantially increasing recall for YOLO.
|
| |
| 14:20-14:40, Paper TuBT4.2 | |
| Lang2Lift: A Language-Guided Autonomous Forklift System for Outdoor Industrial Pallet Handling |
|
| Hoang Nguyen, Huy | Austrian Institute of Technology |
| Huemer, Johannes | AIT Austrian Institute of Technology GmbH |
| Murschitz, Markus | AIT Austrian Institute of Technology GmbH |
| Glück, Tobias | AIT Austrian Institute of Technology GmbH |
| Vu, Minh Nhat | VinUni |
| Kugi, Andreas | TU Wien |
Keywords: Automation in Construction, Computer Vision in Automation, AI-Based Methods
Abstract: Automating pallet handling in outdoor logistics and construction environments remains challenging due to unstructured scenes, variable pallet configurations, and changing environmental conditions. We present textit{Lang2Lift}, a language-guided autonomous forklift system for practical pallet pick-up in real-world outdoor settings. The system allows operators to specify target pallets using natural language, enabling flexible selection among multiple candidates with different loads and spatial arrangements. Lang2Lift combines foundation-model-based perception, 6D pose estimation, geometric refinement, and forklift motion control in a closed-loop pipeline for autonomous fork insertion and lift. We deploy and evaluate the system on the ADAPT autonomous outdoor forklift platform across cluttered scenes, varying lighting conditions, and diverse payload configurations. Experiments on a prompt-conditioned dataset of 129 images and 387 prompt-image pairs, along with real-world single- and multi-pallet trials, demonstrate robust target pallet segmentation, tolerance-aware pose estimation, and reliable end-to-end autonomous pickup. Timing and failure analyses further reveal practical deployment trade-offs and system limitations. Video demonstrations are available at https://eric-nguyen1402.github.io/lang2lift.github.io /.
|
| |
| 14:40-15:00, Paper TuBT4.3 | |
| Omnidirectional Pyramidal Visual Map from 3D Gaussian Splatting for Intelligent Vehicle Localization |
|
| Hu, Zhaozheng | Wuhan University of Technology |
| Wu, Mingjun | Wuhan University of Technology |
| Wu, Qing | ChanagXing Ocean Laboratory |
| Pan, Jinming | Wuhan University of Technology |
| Feng, Feng | Wuhan University of Technology |
| Meng, Jie | Wuhan University of Technology |
Keywords: Autonomous Vehicle Navigation, Computer Vision for Transportation, Probability and Statistical Methods
Abstract: In this paper, we propose an Omnidirectional Pyramidal Visual Map (Omni-PVM) from 3D Gaussian Splatting (3DGS) to enhance intelligent vehicle localization. The proposed Omni-PVM consists of ORB holistic features offline extracted from pyramidal images of synthetic views that are generated from 3DGS given omnidirectional camera poses. Furthermore, we formulate the matching of Omni-PVM as a Hidden Markov Model (HMM) problem with the query image sequence as the observations. By modeling the transition from vehicle motion model and the emission from ORB matching, the optimal camera pose (i.e., vehicle localization) is readily solved with Viterbi algorithm. Finally, the localization results from Omni-PVM are fed into a factor graph together with those from Visual Odometry (VO) to accomplish vehicle localization. The proposed method has been validated in real underground parking lots with low illumination. Experimental results demonstrate that the method can achieve accurate, robust, and fast localization, especially in those challenging scenarios of sharp-turning curved areas. Code is available at https://github.com/Windbreak3r/Omni-PVM.
|
| |
| 15:00-15:20, Paper TuBT4.4 | |
| A New Implementation of NeoSLAM and a Comparative Evaluation with RatSLAM |
|
| Borges, Joao Victor | Universidade Federal Do Rio De Janeiro |
| Coelho, Fabio Junio | Universidade Federal Do Rio De Janeiro |
| Padrao, Paulo | Providence College |
| Fuentes, Jose | Florida International University |
| Costa, Ramon | Federal University of Rio De Janeiro |
| Hsu, Liu | COPPE-UFRJ |
| Bobadilla, Leonardo | Florida International University |
Keywords: Autonomous Vehicle Navigation, Software, Middleware and Programming Environments, Learning and Adaptive Systems
Abstract: This paper presents a new implementation of the NeoSLAM algorithm. The proposed version is a complete rewrite of NeoSLAM into a modular architecture using modern frameworks that, together, enable real-time execution with minimal discarding of input data. This work also provides a comparative evaluation between NeoSLAM and RatSLAM across three datasets under varying environmental conditions. The experimental results highlight differences in mapping consistency and trajectory reconstruction, demonstrating the effectiveness and practical applicability of the proposed ROS 2-based implementation. The results indicate that the new NeoSLAM outperforms the original in terms of processing throughput for real-time applications and achieves comparable performance to RatSLAM in terms of map reconstruction across the evaluated datasets.
|
| |
| 15:20-15:40, Paper TuBT4.5 | |
| 3D Scanning and Point Cloud Processing for Automated Ladder Stitching with a Cobot |
|
| Disimino, Giuseppe | Polytechnic University of Bari |
| Mangini, Agostino Marcello | Politecnico Di Bari |
| Roccotelli, Michele | Polytechnic University of Bari |
| Colucci, Marco | Polytechnic University of Bari |
| Tubito, Domenico | Applica Srl |
| Fanti, Maria Pia | Polytechnic University of Bari |
Keywords: Collaborative Robots in Manufacturing, Computer Vision for Manufacturing, Cloud Computing For Automation
Abstract: Ladder stitching is a hand-crafted invisible seam technique used in haute couture that, so far, has not been successfully replicated by any robotic system. A key challenge in automating this task is defining the stitching trajectory without manual operator teaching, particularly because fabric is non-rigid and cannot be modeled with standard CAD tools. This paper presents a fully automated pipeline that enables a collaborative robot to determine the stitching macro-trajectory directly from a 3D scan of the fabric, without operator intervention. The pipeline covers fabric scanning, point cloud registration and filtering, tension verification through planarity and curvature checks, stitching zone detection via surface normal analysis, and trajectory smoothing. Experimental validation on cotton and synthetic fabrics demonstrates that the generated paths provide a highly accurate spatial reference, with physical execution tests achieving a Mean Absolute Error well within the 4 mm application tolerance. However, maximum deviations observed on curved geometries highlight the limitations of standard inverse kinematics under unmodeled compliances. To bridge this gap, a crucial next step involves learning a data-driven error model capable of mapping the planned Cartesian trajectory points directly into the compensated joint configurations required to accurately reach the physical target on the fabric.
|
| |
| 15:40-16:00, Paper TuBT4.6 | |
| A Front-To-Back Registration Algorithm Based on Optical Flow for Industrial Double-Sided Inkjet Printing Machines |
|
| Wang, Yihan | Northeast University |
| Jiang, Daqi | Northeastern University |
Keywords: Computer Vision for Manufacturing, Factory Automation, Intelligent and Flexible Manufacturing
Abstract: This paper proposes S2G-RAFT, an optical-flow-based registration method for pixel-level front-to-back alignment in industrial double-sided inkjet printing on non-rigid fabrics. The method estimates a dense deviation field between the first-side printed image and the original design template, providing geometric guidance for second-side compensation. To improve robustness under complex fabric textures and subtle deformation, S2G-RAFT introduces a Geometric Gating Engine into the RAFT update module to adaptively reweight correlation features. In addition, a high-fidelity simulation dataset is constructed using warp-and-weft texture modeling and Gaussian Random Field deformation. Experiments on simulated fabric datasets show that S2G-RAFT reduces registration error while maintaining practical inference efficiency. These results demonstrate its potential for high-precision front-to-back alignment in digital textile printing.
|
| |
| TuBT5 |
205 |
| Emerging Technology of Automation System 2 |
Regular Session |
| |
| 14:00-14:20, Paper TuBT5.1 | |
| Multi-Embodiment Robotic Retargeting Via Guided Diffusion Model |
|
| Cao, Zhefeng | Hong Kong University of Science and Technology |
| Liu, Ben | Southern University of Science and Technology |
| Yang, Shunpeng | Hong Kong University of Science and Technology |
| Li, Sen | Hong Kong University of Science and Technology |
| Zhang, Wei | Southern University of Science and Technology |
| Chen, Hua | Zhejiang University |
Keywords: Deep Learning in Robotics and Automation
Abstract: The retargeting of diverse and kinematically feasible reference trajectories from human demonstrations facilitates the transfer of human skills to robots, thereby simplifying the learning of complex tasks. However, existing motion retargeting pipelines are heavily tied to specific skeletons with individually hand-crafted settings, which limits scalability across heterogeneous embodiments. In this work, our unified framework transfers given reference motions to multiple heterogeneous robots via a transformer-based diffusion model. We represent robotic skeletons as graph structures, enabling effective extraction of topological and geometrical properties via customized attention mechanisms. Due to the scarcity of high-quality motion datasets for diverse robots, we guide the diffusion model with energy-based retargeting losses. Experiments across heterogeneous robots validate our kinematically feasible retargeting framework on controllable kinematic structures, unseen embodiments, diverse motion domains, and sim-to-sim whole-body tracking tasks.
|
| |
| 14:20-14:40, Paper TuBT5.2 | |
| Threading Optimization for Vision-Language-Action Model Inference in Low-Cost Smart Agricultural Manipulation |
|
| Truongcao, Keith | Drexel University |
| Nhu, Christopher | Drexel University |
| An, Zijian | Drexel University |
| Nguyen, Phong | Drexel University |
| Cai, Siwei | Drexel University |
| Zhou, Lifeng | Drexel University |
Keywords: Deep Learning in Robotics and Automation, Agricultural Automation, Manipulation Planning
Abstract: Vision-Language Action (VLA) models face challenges like slow inference speed and poor fine-grained motion control, limiting their industrial adoption. While the Real-Time Action Chunking (RTAC) algorithm has been proposed to address these bottlenecks, bridging the gap between the algorithm provided in pseudocode to a stable, real-world deployment on a low-cost robotic arm remains a challenge. In this work, we present a complete system-level implementation of RTAC tailored for a low-cost robotic manipulation system. We advance beyond the original high-level pseudocode by optimizing the threading implementation for the policy inference and control pipeline, reducing end-to-end latency and improving responsiveness without modifying the underlying policy. We evaluate this system on tasks involving the manipulation of agricultural produce, specifically garlic bulbs and walnuts. Experimental results demonstrate that our custom threading implementation significantly improves control stability and speed compared to the base implementation of RTAC.
|
| |
| 14:40-15:00, Paper TuBT5.3 | |
| Learning to Unscrew the Cap Naturally: Skill Decomposition and Policy Distillation for Robot Twisting |
|
| Feng, Yanbin | Zhejiang University |
| Li, He | Zhejiang University |
| Han, Fuzhang | Zhejiang Humanoid Robotic Innovation Center Co |
| Xie, Liang | Zhejiang Humanoid Robotic Innovation Center Co |
| Zhou, Zhongxiang | Zhejiang Humanoid Robotic Innovation Center Co |
| Xiong, Rong | Zhejiang University |
Keywords: Deep Learning in Robotics and Automation, AI-Based Methods, Manipulation Planning
Abstract: Dexterous twisting is a fundamental yet challenging capability for anthropomorphic manipulation due to its high-dimensional action space, strong spatio-temporal coupling, and complex contact dynamics. While reinforcement learning (RL) often struggles with convergence under such conditions, learning from demonstration (LfD) requires extensive data collection and suffers from human-to-robot retargeting errors. This paper proposes SDPD, a task-oriented learning method that integrates skill decomposition (SD-RL) and policy distillation (PD). The SD-RL module decomposes the twisting task into three sequential sub-skills: Grasp, Rotate, and Release, and learns physically consistent policies under tactile guidance. Subsequently, an offline policy distillation procedure based on TD3+BC consolidates stage-wise expert policies into full-process policies with reduced dependence on hard phase switching. Simulation results show that SDPD improves manipulation efficiency and long-horizon twisting performance while maintaining coordinated contact-rich motions compared with representative baselines. Across the tested cap-size range, the proposed method achieves high-performance execution over 89.1% of the range, improving the coverage by 51.4 percentage points over the baseline. Sim-to-real deployment on the Shadow Hand platform further demonstrates stable real-world performance: the system can continuously twist the bottle cap for more than 25 full turns over a 20-minute period while maintaining a motion success rate above 85%.
|
| |
| 15:00-15:20, Paper TuBT5.4 | |
| Aerial Inspection Behaviors Via RL-Based Quadrotor Control for Under-Canopy Forest Environments |
|
| Lagos Suarez, Fausto Mauricio | Luleå University of Technology |
| Saradagi, Akshit | Luleå University of Technology, Luleå, Sweden |
| Sumathy, Vidya | Luleå University of Technology |
| Sankaranarayanan, Viswa Narayanan | Lulea University of Techonology |
| Nikolakopoulos, George | Luleå University of Technology |
Keywords: Deep Learning in Robotics and Automation, Autonomous Vehicle Navigation, Motion and Path Planning
Abstract: This paper addresses the problem of using a deep Reinforcement Learning (RL)-based low-level Quadrotor controller within an autonomous Quadrotor navigation stack for aerial inspection missions in under-canopy forest environments. Specifically, the article presents an end-to-end (mapping states to RPMs) Quadrotor control policy that achieves inspection view-pose tracking (simultaneous position and yaw reference tracking), which is crucial for various target inspection behaviors and point-to-point navigation in forests. To ensure safe and reliable deployment of the end-to-end RL controller in long-range missions, this article utilizes a higher navigation guidance layer comprising of a Traveling Salesman Problem planner (TSP) and a Rapidly-exploring Random Tree Star (RRT*) planner. Over a known map of a forest and a set of user-specified inspection regions, the TSP planner finds the optimal visitation sequence. Between two target regions, collision-free paths that respect the tracking limitations of the lower end-to-end RL policy are generated by an RRT* planner. Through five target inspection scenarios, this article demonstrates that an RL-based motor-level stabilizing controller, supported by a navigation guidance layer, can be used effectively as the low-level inspection execution module for under-canopy forest inspection missions.
|
| |
| 15:20-15:40, Paper TuBT5.5 | |
| Learning Collision Awareness from Egocentric View Using Diffusion Policy and Local Map Representation |
|
| Le, Dinh Dang Khoa | The University of Technology Sydney |
| Jayasuriya, Maleen | University of Canberra |
| Hu, Gibson | University of Technology, Sydney |
| Liu, Dikai | University of Technology, Sydney |
Keywords: Deep Learning in Robotics and Automation, Collision Avoidance, Motion and Path Planning
Abstract: Diffusion-based imitation learning has achieved remarkable success in robotic manipulation. Yet diffusion policies suffer from a limitation: they lack persistent spatial memory when operating with egocentric vision, leading to potential collisions with obstacles. This paper presents a framework that integrates diffusion policies with local map representation methods, which can serve as an external spatial memory, enabling collision-aware motion generation. Occupancy grid and distance field methods are investigated and evaluated across 2D navigation and 3D manipulation tasks. The results show that the diffusion policy with local distance fields consistently outperforms the original diffusion policy that uses images or the diffusion policy with an occupancy grid. The results also show that using 20-60% of the whole workspace map can achieve a similar success rate to the case of using the whole map, while maintaining computational efficiency.
|
| |
| 15:40-16:00, Paper TuBT5.6 | |
| Adaptive Graph Structure Learning with Node Transformation in Spatial-Temporal Graph Networks for Eating Intention Prediction |
|
| Zhang, Xuan | Hongkong Polytechnic University |
| Yuan, Ye | The Hong Kong Polytechnic University |
| Huang, Hailong | The Hong Kong Polytechnic University |
Keywords: Deep Learning in Robotics and Automation, Human Factors and Human-in-the-Loop, Learning and Adaptive Systems
Abstract: Accurately identifying eating intentions is the cornerstone of autonomous assistant robots in human-robot interaction (HRI). This work presents an eating intention recognition approach that integrates facial landmarks and head pose information. The proposed method employs a novel facial graph module to fuse the pose and landmarks information to extract spatiotemporal facial features and organize them as a graph. Then, a node attribute transform attention mechanism is responsible for transforming heterogeneous nodes with different attributes to homologous nodes, making a better capture of the distinct characteristics of nodes with different attributes. Finally, an adaptive adjacency matrix based similarity attention mechanism provides edge weights between nodes to represent the unique connection relationships for further embedding learning. The proposed method is evaluated on a self-collected dataset and compared with state-of-the-art methods. Experimental results demonstrate the superior performance of the proposed approach, with a 96.30% accuracy rate in recognizing eating intentions, higher than that of other approaches.
|
| |
| TuBT6 |
503 |
LLM-Based Application for Manufacture-Circulation Industrial System (MCIS)
2 |
Regular Session |
| |
| 14:00-14:20, Paper TuBT6.1 | |
| Multi-Timescale Workload-Thermal Coordinated Control for AI Data Centers Via Offline Multi-Agent Reinforcement Learning |
|
| Jia, Junyuan | Xi'an Jiaotong University |
| Xu, Zhanbo | Xi'an Jiaotong University |
| Liu, Jinhui | Xi'an Jiaotong University |
| Liu, Yaping | Xi'an Jiaotong University |
| Wu, Jiang | Xian Jiaotong University |
| Shen, Yuanjun | Xi'an Jiaotong University |
| Guan, Xiaohong | Xi'an Jiaotong University |
Keywords: Energy and Environment-aware Automation, Reinforcement, Planning, Scheduling and Coordination
Abstract: In AI data centers, IT equipment and cooling systems are the dominant energy consumers, jointly accounting for approximately 90% of the total energy footprint. While workload-thermal coordinated control offers energy-saving potential, effectively integrating these subsystems is challenged by the timescale discrepancy between rapid workload-side heat generation and slow thermal dynamics. Existing multi-timescale coordination architectures often decouple agents in ways that break Markov decision process (MDP) stationarity, hindering stable data-driven policy learning. To overcome this non-stationarity while minimizing energy consumption under server thermal constraints, we model joint workload allocation and air handling units regulation as a multi-timescale, cooperative multi-agent Dec-POMDP. We propose an offline centralized training with decentralized execution (CTDE) framework based on Implicit Q-learning (IQL). This framework incorporates a multi-timescale experience replay mechanism to safely derive control policies from historical data. Validated via an EnergyPlus-based simulator using real-world traces, our approach ensures a zero thermal violation rate, decreases cooling energy consumption by 19.4%, and improves power usage effectiveness (PUE) from 1.448 to 1.353 compared to conventional baselines.
|
| |
| 14:20-14:40, Paper TuBT6.2 | |
| Digital Waitron: A Human-Centric Digital Twin Framework for Just-In-Time Service in High-Contact Dining |
|
| Liang, Zhe | The Hong Kong Polytechnic University |
| Li, Ming | The Hong Kong Polytechnic University |
| Huang, George Q. | The Hong Kong Polytechnic University |
Keywords: Human-Centered Automation, Human Factors and Human-in-the-Loop, Data fusion
Abstract: High-contact dining depends on just-in-time (JIT) service to maintain service rhythm, responsiveness, and customer experience. In practice, persistent labor shortages and increasing waiter-to-table ratios make it difficult for staff to continuously perceive subtle and fast-changing table-side needs. This challenge is intensified by dynamic, occluded, and cluttered restaurant environments, where reliable service-state perception is difficult in practice. This paper presents Digital Waitron, a human-centric digital twin framework for assistive JIT service support in high-contact dining. Unlike industrial digital twins that often emphasize high-fidelity physical simulation, this service-oriented twin focuses on continuously synchronizing table-side service states and closing the loop through waiter-mediated service actions. Rather than automating service delivery, the framework augments waiters’ perception and decision-making through interpretable table-state understanding. By converting low-cost multimodal observations into service-relevant state cues, it extends waitstaff’s perceptual boundary and supports timely assistance to customers. Methodologically, the system combines modality-specific event extraction from weight and vision streams with lightweight multimodal alignment and explicit state reasoning to infer dining progression and generate assistive service support. A restaurant prototype and a preliminary event-level evaluation on real dining recordings show that the proposed method improves plate-event recovery over weight-only, vision-only, and simple temporal fusion baselines. The study provides a deployment-oriented reference for perceptual augmentation and hospitality digitalization in high-contact dining environments.
|
| |
| 14:40-15:00, Paper TuBT6.3 | |
| A Dual-Temporal Hybrid Model for Robust Power Estimation of Industrial Robotic Systems |
|
| Kim, Chani | University of Ulsan |
| Kang, Eun-Young | University of Southern Denmark |
| Heredia, Juan | University of Southern Denmark |
| Kim, Byeong-woo | University of Ulsan |
Keywords: Hybrid Strategy of Intelligent Manufacturing, Power and Energy Systems automation
Abstract: As the integration of robots in manufacturing continues to expand, accurately estimating robot power consumption has become crucial for improving energy efficiency. Conventional physics-based estimation methods often struggle to yield precise power predictions due to the complexity of modeling non-linear factors. To address these limitations, this paper proposes a Dual-Temporal(DT) hybrid model framework that integrates a physics-based model with a data-driven approach. The proposed method first estimates initial joint currents based on physical laws to reflect robotic dynamics. A data-driven model then processes these estimates to account for non-linear characteristics, which are difficult to model analytically. To capture both local and global temporal dependencies, we developed a DT Hybrid model by integrating Long Short-Term Memory (LSTM) and Gated Recurrent Units (GRU). Experimental results on unseen data demonstrate that the proposed DT hybrid model achieves superior accuracy and generalization performance. These findings suggest that the DT hybrid approach effectively mitigates the modeling constraints of pure physical methods, providing a robust framework for power consumption analysis in industrial robotics.
|
| |
| 15:00-15:20, Paper TuBT6.4 | |
| A Novel Rolling-Window and Production-Maintenance-Based Machine Learning Approach to Virtual Metrology in Semiconductor Manufacturing (I) |
|
| Fan, Shu-Kai S. | National Taipei University of Technology |
Keywords: Machine learning, Manufacturing, Maintenance and Supply Chains, Semiconductor Manufacturing
Abstract: Virtual metrology enables semiconductor manufacturers to continuously monitor and predict the quality and performance of semiconductor production during the manufacturing process. This kind of real-time monitoring tool is essentially important for maintaining process stability and ensuring consistent product quality. In this paper, a new rolling-window and production-maintenance-based approach is proposed on a weekly basis to establish the virtual metrology system. The proposed thickening and dewatering framework serves as the central core of the virtual metrology model, enabling real-time information and feedback to aid in making timely adjustments to the processing tool within machinery limitations. A meticulous approach to addressing each component within this innovative framework has been instrumental in achieving exceptional VM performance. The training dataset is routinely updated as the production maintenance is taking place to facilitate model retraining. A practical challenge to the domain engineer under certain circumstances where the process parameter considered in the virtual metrology model is realistically run out of specification is also rigorously investigated in this research. For the evaluation purposes, a real-world industrial case of the chemical vapor deposition process in semiconductor manufacturing is presented to illustrate the proposed approach. The proposed virtual metrology approach shows a substantial improvement in mean absolute error with 39.99%, 26.26%, and 30.75% for three different recipes, respectively, as compared to the base model currently implemented on the operation site.
|
| |
| 15:20-15:40, Paper TuBT6.5 | |
| Mechanism Design and Kinematic Analysis of a Confined Mechanism for Minimally Invasive Surgery |
|
| Soghomonyan, Hayk | Zhejiang University |
| Gunawan, Ferrell Xavier | Zhejiang University ZJU-UIUC |
| Marpaung, Nico S P | Zhejiang University |
| Yang, Liangjing | Zhejiang University |
Keywords: Mechanism Design in Meso, Micro and Nano Scale, Product Design, Development and Prototyping, Medical Robots and Systems
Abstract: The remote center of motion (RCM) is a mechanism that constrains the end effector to point at a fixed point in space while the joints and links move. This kinematic constraint is critical in minimally invasive surgeries (MIS) to prevent tissue damage and maintain access to the surgical site through a small incision. Current control-based RCM robots pose several safety issues, including over-reliance on complex control algorithms, and high manufacturing cost. In response to these issues, we adopt a design-based approach, reducing the reliance on software so as to provide a safer, mechanically simpler and potentially more cost-effective solution. The proposed design has several new features compared to the current state-of-the-art: (1) relatively easier to control, (2) cost effective and easily manufacturable, (3) lightweight, (4) compact, (5) fewer number of actuators required, (6) has the potential to be upgraded to a version with controllable radius of the link motion and RCM. Through systematic mathematical modeling, kinematic analysis, and physical prototyping, we demonstrated joint trajectories that maintain the RCM and validated the design concept with the implementation of a joint-controlled physical prototype showing sub-millimeter errors. Observation shows that the proposed RCM mechanism design is simpler, safer, more precise and affordable, while offering promising potential for advancements in various fields.
|
| |
| 15:40-16:00, Paper TuBT6.6 | |
| Automatic Cross-Sectional Layout Generation and Stiffness Optimization for Medical Continuum Robots |
|
| Ning, Maizhen | Duke Kunshan University |
| Zhang, Wendu | The Chinese University of Hong Kong |
| Hu, Bintao | Xi'an Jiaotong-Liverpool University |
| Huang, Yuanrui | Xi'an Jiaotong-Liverpool University |
Keywords: Medical Robots and Systems, AI and Machine Learning in Healthcare, AI-Based Methods
Abstract: Continuum robots have become prominent in medical applications due to their flexibility; they can integrate medical instruments to enable surgical operations. However, designing the cross-sectional layouts of these robots is time-consuming, especially when fitting various tool sizes within limited space. To address this challenge, we propose a framework that automates the generation of continuum robot cross-sectional layouts based on user-defined requirements. We formalize tool channel relationships as geometric constraints, enabling the automatic creation of error-free layouts that satisfy all mathematical and physical constraints. Additionally, our approach introduces the cross-sectional moment of inertia as a constraint, enabling stiffness customization for diverse applications through optimized layout design. Both simulations and experiments demonstrate the effectiveness of the proposed framework. These innovations hold significant implications for the development of continuum robots and have the potential to be used in clinical applications.
|
| |
| TuBT7 |
504 |
| AI-Driven Decision-Making and Control in Intelligent Transportation Systems |
Special Session |
| Organizer: Ma, Benedict Jun | The Hong Kong University of Science and Technology (Guangzhou) |
| Organizer: Kuo, Yong-Hong | The University of Hong Kong |
| Organizer: Xu, Guanhao | Oak Ridge National Laboratory |
| Organizer: Zhao, Zhouqiao | MIT |
| Organizer: Liu, Yishun | Central South University |
| Organizer: Yang, Hai | The Hong Kong University of Science and Technology |
| |
| 14:00-14:20, Paper TuBT7.1 | |
| Hierarchical Decision-Making and Control for Connected and Autonomous Vehicle Platoons at Unsignalized Intersections (I) |
|
| Gu, Songtao | The Hong Kong University of Science and Technology (Guangzhou) |
| Ma, Benedict Jun | The Hong Kong University of Science and Technology (Guangzhou) |
Keywords: Intelligent Transportation Systems, Planning, Scheduling and Coordination, Autonomous Vehicle Navigation
Abstract: Efficient and safe coordination of connected and automated vehicle (CAV) platoons at unsignalized intersections remains a critical challenge for smart city logistics. Existing hierarchical frameworks often suffer from sequential scheduling inefficiencies, non-convergent low-level control loops, stop-line singular deadlocks, and ambiguous experimental validation. To address these deficiencies, this paper proposes a tractable bi-level architecture that strictly decouples strategic scheduling from tactical trajectory regulation. The strategic layer employs an event-triggered mixed-integer linear programming scheduler with a batch execution bonus to compute conflict-free crossing time windows. The tactical layer integrates a fuel-optimal quintic polynomial planner with real-time spatiotemporal projection feedback to track the assigned windows smoothly. Meanwhile, internal platoon members execute a localized proportional-derivative-plus-integral (PD+I) controller, augmented with anti-windup saturation clamps and brick-wall kinematic safety filters. Extensive simulations are conducted in the SUMO platform using a realistic multi-lane intersection topology under hyper-saturated traffic demand profiles. Benchmarked against conventional first-come-first-served and actuated signal control, the proposed model achieves a peak macroscopic throughput of 4493 veh/h, a 100% non-stop passage rate, an average reduction of 41.2% in control delays, and savings of up to 47.1% in cumulative fuel consumption and CO 2 emissions, while maintaining zero operational deadlocks across all high-volume traffic gradients. These results highlight the proposed framework's effectiveness in enhancing intersection throughput, energy efficiency, and operational robustness for CAV platoons.
|
| |
| 14:20-14:40, Paper TuBT7.2 | |
| A Cohesive Adaptive Hybrid Framework for Short-Term Railway Traction Load Forecasting: Integrating Hierarchical Decomposition with Adaptive Input Steps (I) |
|
| Ma, Qian | Xiangtan University |
| Lv, Runwei | Xiangtan University |
| Yang, Hanlei | Xiangtan University |
Keywords: Intelligent Transportation Systems, Machine learning, Power and Energy Systems automation
Abstract: Short-term forecasting of electrified railway traction loads is critical for power system stability but remains challenging due to the load's inherent non-stationarity, high-frequency volatility, and complex dependencies. To address the limitations of conventional models in capturing these characteristics, this paper proposes a novel, self-adaptive forecasting framework. First, a Hierarchical PSO-VMD strategy is introduced, which adaptively optimizes parameters and specifically re-decomposes high-frequency nested oscillations based on sample entropy. Second, to solve the "Temporal Receptive Field Mismatch", we propose a PSO-driven Adaptive Input Steps (PSO-AIS) strategy. This mechanism utilizes a coarse-to-fine search to dynamically align the model's sequence lookback observation window with the time-varying inertia of traction loads. Third, a self-evolving hybrid SSA-TCN-LSTM predictor is established; the SSA globally optimizes hyperparameters to balance local feature extraction and long-term dependency modeling. Finally, an adaptive robust weighted fusion mechanism optimizes final reconstruction weights. Experimental results on real-world substation data demonstrate that the proposed framework significantly outperforms baseline models in predictive accuracy and general stability.
|
| |
| 14:40-15:00, Paper TuBT7.3 | |
| Physics-Guided Distributional Shaping for Runtime Risk Assessment in Autonomous Driving (I) |
|
| Liu, Shuai | Xi'an Jiaotong University |
| Yan, Chao-Bo | Xi'an Jiaotong University |
| Hu, Jianchen | Xi'an Jiaotong University |
| Li, Dachuan | Southern University of Science and Technology |
Keywords: Intelligent Transportation Systems, Autonomous Vehicle Navigation
Abstract: Reliable risk assessment in interactive traffic remains a fundamental challenge, as rule-based methods rely on rigid safety constraints while pure statistical models lack explicit kinematic boundaries. We propose a physics-guided distributional shaping framework that reweights a learned interaction density with a bounded kinematic energy through Gibbs tilting, and approximates the shaped density with a closed-form Gaussian mixture model (GMM) for efficient online scoring. Unlike runtime rule correction, the proposed framework embeds the safety bias directly into the scoring distribution offline, while a distance-aware activation term further reduces false alarms from distant non-threatening agents. In zero-shot transfer from nuPlan to Lyft under a Time-to-Collision (TTC)-based high-risk proxy, the proposed method achieves 0.9649 receiver operating characteristic area under the curve (ROC-AUC) and 77.98% true positive rate (TPR) at 5% false positive rate (FPR), outperforming rule-based and purely statistical baselines in pairwise risk ranking.
|
| |
| 15:00-15:20, Paper TuBT7.4 | |
| Variable Speed Limit Control Strategy for Freeways Based on CAV-Leading Mixed Vehicle Groups (I) |
|
| Jin, Kongning | Hunan Provincial Key Laboratory of Energy Saving Control and Safety Monitoring for Rail Transportation |
| Xing, Lu | Central South University |
| Xiong, Kaibin | Central South University |
| Cao, Yijun | Changsha University of Science & Technology |
| Liu, Yishun | Central South University |
| Li, Mingzhen | Central South University |
Keywords: Intelligent Transportation Systems, Autonomous Vehicle Navigation, Process Control
Abstract: Connected and automated vehicles (CAVs) offer significant potential for improving traffic control at freeway bottlenecks. This study proposes a novel variable speed limit (VSL) control strategy for mixed traffic composed of CAVs and human-driven vehicles (HDVs), using CAV-leading mixed vehicle groups as the minimum controllable unit.
|
| |
| 15:20-15:40, Paper TuBT7.5 | |
| Vessel Traffic Flow Prediction on Sparse Data Via Spatio-Temporal Graph Neural Networks with a Learnable Tweedie Head (I) |
|
| Lee, Kyeongjun | KAIST |
| Kim, Heeyoung | KAIST |
Keywords: Intelligent Transportation Systems, Probability and Statistical Methods, Machine learning
Abstract: Accurate vessel traffic flow prediction is crucial for smart port operations and navigational safety. However, maritime traffic flow data are often highly sparse with intermittent bursts, making robust forecasting challenging. Under such conditions, conventional spatio-temporal graph neural networks (ST-GNNs) can degrade toward conservative near-zero predictions and fail to capture non-zero activity. Although zero-inflated negative binomial (ZINB) models partially address excess zeros, their two-part formulation can still remain conservative around abrupt transitions. To address these issues, we propose a model-agnostic learnable Tweedie head that can be attached as a plug-and-play output module to arbitrary ST-GNN backbones. Instead of likelihood-based Tweedie training, which typically requires surrogate objectives, our approach optimizes the closed-form Tweedie unit deviance and predicts the mean for point forecasting while learning a node-level variance power to capture heterogeneous variability across port areas. Experiments on a maritime traffic graph constructed from real-world AIS data in the Port of Los Angeles and Long Beach show that the proposed head consistently improves RMSE across multiple ST-GNN backbones, especially on non-zero events, leading to more reliable forecasts for practical maritime traffic control.
|
| |
| 15:40-16:00, Paper TuBT7.6 | |
| Class-Based Storage Policy in Robotic Cellular Warehousing Systems (I) |
|
| Ma, Benedict Jun | The Hong Kong University of Science and Technology (Guangzhou) |
Keywords: Logistics, Inventory Management, Manufacturing, Maintenance and Supply Chains
Abstract: Robotic Cellular Warehousing Systems (RCWS) provide a flexible and scalable automation solution for modern order fulfillment operations. In such systems, robot travel distance during order picking is a key determinant of throughput and operational efficiency. This paper analytically investigates the impact of storage policies on robot traveling in RCWS, with a focus on comparing random storage and class-based storage. We develop a tractable model of the system operating under a pick-by-batch policy and derive closed-form expressions for the expected robot travel distance under both storage policies. The analysis explicitly accounts for warehouse layout, robot capacity, and demand heterogeneity across products. We show that while class-based storage weakly dominates random storage in the horizontal travel component, its overall performance critically depends on how demand classes are spatially arranged relative to the unloading port. To further illustrate these analytical insights, we conduct numerical experiments under specific demand distributions. The results reveal non-monotonic effects of the number of storage classes and highlight strong interaction effects among demand skewness, storage granularity, and robot capacity. These findings provide clear conditions under which class-based storage outperforms random storage and offer practical guidance for storage design in RCWS.
|
| |
| TuBT8 |
501 |
| Virtual Session 2 |
Regular Session |
| |
| 14:00-14:20, Paper TuBT8.1 | |
| Mechanical Kinematic Synchronization for Multi-Axis End-Effector Alignment: An Industry-Focused Case Study in Robotic Material Handling |
|
| Yadav, Santosh | Amazon Robotics |
Keywords: Foundations of Automation, Motion Control, Simulation and Animation
Abstract: Multi-axis end-effectors used in industrial material- handling tasks often require coordinated linear and rotational motion to maintain part orientation during lifting, transferring, or loading operations. In many systems, this synchronization is achieved through software compensation and high-bandwidth control loops, which can become unstable when mechanical tolerances, compliance, or asymmetric loading introduce non- linear behavior. This paper presents a mechanical-design-driven approach to kinematic synchronization for a dual-axis end- effector performing rotational alignment during vertical transla- tion. A redesigned mechanism architecture introduces geometric coupling, constraint allocation, and stiffness-balanced linkages that enforce a predictable kinematic relationship between axes without relying on active control. A comprehensive experimental study across three mechanism variants, four load conditions, and five travel distances (over 2,400 motion cycles) demonstrates a 13.3× reduction in angular drift, an 11.7× reduction in phase error, a 12× improvement in angular repeatability, and an 82% reduction in required control bandwidth. Finite element analysis (FEA) validates the structural integrity and stiffness distribution required for reliable synchronization. The findings highlight the value of embedding synchronization directly into mechanism de- sign to improve reliability, repeatability, and integration efficiency in industrial robotic systems. As a work-in-progress developed independently, this research outlines ongoing durability testing, scaling efforts, and opportunities for industrial deployment. Index Terms—Mechanical synchronization, end-effector, mate- rial handling, kinematics, FEA, industrial robotics.
|
| |
| 14:20-14:40, Paper TuBT8.2 | |
| GraspSense: Physically Grounded Grasp and Grip Planning for a Dexterous Robotic Hand Via Language-Guided Perception and Force Maps |
|
| Semenyakina, Elizaveta | Skolkovo Institute of Science and Technology |
| Snegirev, Ivan | Skolkovo Institute of Science and Technology |
| Lezina, Mariya | Skolkovo Institute of Science and Technology |
| Altamirano Cabrera, Miguel | Skolkovo Institute of Science and Technology (Skoltech), Moscow, Russia |
| Gulyamova, Safina | Skolkovo Institute of Science and Technology |
| Tsetserukou, Dzmitry | Skolkovo Institute of Science and Technology |
Keywords: Motion Control, Robust/Adaptive Control, Motion and Path Planning
Abstract: Dexterous robotic manipulation requires more than geometrically valid grasps: it demands physically grounded contact strategies that account for the spatially non-uniform mechanical properties of the object. However, existing grasp planners typically treat the surface as structurally homogeneous, even though contact in a weak region can damage the object despite a geometrically perfect grasp. We present a pipeline for grasp selection and force regulation in a five-fingered robotic hand, based on a map of locally admissible contact loads. From an operator command, the system identifies the target object, reconstructs its 3D geometry using SAM3D, and imports the model into Isaac Sim. A physics-informed geometric analysis then computes a force map encoding the maximum lateral contact force admissible at each surface location without deformation. Grasp candidates are filtered by geometric validity and task-goal consistency. When multiple candidates are comparable under classical metrics, they are re-ranked using a force-map-aware criterion that favors grasps with contacts in mechanically admissible regions. An adaptive impedance controller scales the stiffness of each finger according to the locally admissible force at the contact point, enabling safe and reliable grasp execution. Validation on paper, plastic, and glass cups demonstrates that the proposed approach consistently selects structurally stronger contact regions and maintains grip forces within safe bounds. Consequently, the work reframes dexterous manipulation from a purely geometric problem into a physically grounded, joint planning problem of grasp selection and grip execution for future humanoid systems.
|
| |
| 14:40-15:00, Paper TuBT8.3 | |
| GRACE: Gradient-Free Robot Action Generation Via Combined Diffusion-MPPI Posterior Estimation |
|
| Park, Leesai | Kyung Hee University |
| Hong, Jiho | Kyung Hee University |
| Kim, Sanghyun | Kyung Hee University |
Keywords: Optimization and Optimal Control, Motion Control, Robust/Adaptive Control
Abstract: Diffusion-based planning captures multimodal action distributions from demonstration data, but existing guidance methods require differentiable task objectives and dynamics, excluding the nondifferentiable constraints common in robot deployment. We introduce GRACE (Gradient-free Robot Action generation via Combined diffusion-MPPI posterior Estimation), a framework that uses Model Predictive Path Integral (MPPI) control as a sampling-based guidance mechanism for the diffusion reverse process. By exploiting the common score-based structure shared by diffusion denoising and MPPI, we combine the two updates into a posterior reverse kernel without requiring differentiability of the objective or dynamics, and show that standard gradient guidance is a local first-order special case of this correction. We validate GRACE on a 7-DoF manipulator goal-reaching task, where MPPI guidance enforces obstacle avoidance via binary collision constraints without retraining the diffusion prior.
|
| |
| 15:00-15:20, Paper TuBT8.4 | |
| Decentralized LLM-Driven Coordination of Acoustic Robots for Contactless Object Manipulation |
|
| Wang, Yingying | University College London |
| Kemsaram, Narsimlu | University of Malaya |
| Subramanian, Sriram | University College London |
Keywords: Planning, Scheduling and Coordination, Agent-Based Systems, AI-Based Methods
Abstract: Natural language interfaces can significantly reduce the complexity of interacting with multi-robot systems, especially in automation settings where non-expert users must issue high-level commands. In parallel, acoustic manipulation using ultrasonic phased arrays enables contactless object handling for contamination-sensitive applications such as healthcare, laboratory automation, and precision transport. However, the integration of large language models (LLMs) with distributed acoustic mobile robots remains largely unexplored. This paper presents a decentralized framework for natural language-driven coordination of acoustic robots for contactless object manipulation. The proposed framework converts spoken instructions into executable multi-robot task plans through a pipeline that combines Whisper-based speech recognition, LLM-driven semantic parsing, structured JSON task representation, and distributed scheduling. The JSON schema encodes robot assignments, temporal dependencies, spatial constraints, and synchronization requirements for sequential, parallel, and synchronized execution. To support reliable coordination, we introduce a distributed synchronization protocol within a ROS 2-based multi-agent architecture. The system is implemented on two TurtleBot3-based acoustic robots, each equipped with an ultrasonic phased array for contactless object transport. The framework was evaluated in three task scenarios: sequential execution, parallel multi-robot transport, and synchronized cooperative manipulation. The system achieved task success rates of 96% for sequential tasks, 86% for parallel execution, and 70% for synchronized collaborative transport.
|
| |
| 15:20-15:40, Paper TuBT8.5 | |
| Kinetic Minimalism: Benchmark Agility and Contact-Rich Whole-Body Interaction in an 8-DOF Wheeled Quadruped |
|
| Pan, Yuxin | Duke University |
Keywords: Industrial and Service Robotics, Actuation and Joint Mechanisms, Deep Learning in Robotics and Automation
Abstract: High-degree-of-freedom quadrupeds exhibit impressive agility, but their mechanical complexity, sim-to-real fragility, and cost often limit practical industrial deployment. This work investigates whether a drastically reduced-actuation wheeled quadruped can retain benchmark-level dynamic capability while simplifying the hardware and control stack. We present an 8-degree-of-freedom (DoF) platform with four torque-controlled hip joints and four velocity-driven wheels, and show how missing steering and knee joints can be compensated through footprint-modulated differential drive, whole-body pitching, and torque-based proprioception. The platform is validated through three layers of evidence: dynamic analysis showing ballistic jumping capability, a single end-to-end reinforcement learning (RL) policy that solves a reconstructed Barkour agility benchmark in Isaac Sim, and physical experiments showcasing contact-rich whole-body interactions. The fully learned end-to-end policy achieves performance comparable to an expert-demonstration residual policy. The sim-to-real gap is bridged by characterizing the torque-controlled hips, which also provide proprioceptive signatures of terrain interaction. Robust real-world demonstrations include dynamic jumping, stair traversal, bracing, and environment-aided whole-body non-prehensile interaction. These results establish reduced-DoF wheeled quadrupeds as capable, reproducible, and practical platforms for agile industrial robotics.
|
| |
| 15:40-16:00, Paper TuBT8.6 | |
| Latent Dynamics–Aware OOD Monitoring for Trajectory Prediction with Provable Guarantees |
|
| Guo, Tongfei | Northeastern University |
| Su, Lili | Northeastern University |
Keywords: Autonomous Vehicle Navigation, Formal Methods in Robotics and Automation, Probability and Statistical Methods
Abstract: In safety-critical Cyber-Physical Systems (CPS), accurate trajectory prediction serves as vital guidance for downstream planning and control. Although deep learning models provide high-fidelity forecasts in the validation dataset, their reliability degrades in out-of-distribution (OOD) scenarios arising from environmental uncertainty or rare traffic behaviors in real-world implementation. Detecting OOD events is particularly challenging because of the evolving traffic conditions and changes in interaction patterns. At the same time, formal guarantees on detection delay and false-alarm rates are critically needed due to the safety-critical nature of autonomous driving. Following recent work [1], we reframe OOD monitoring for trajectory prediction as the quickest changepoint detection (QCD) problem, which provides a principled statistical framework with well-established theory. We observe that the real-world evolution of prediction errors on in-distribution (ID) data can be well modeled by a Hidden Markov Model (HMM). Leveraging this structure, we extend the recent cumulative Maximum Mean Discrepancy approach to our setting. The resulting method avoids requiring detailed prior knowledge of the post-change distribution while admitting provable guarantees on detection delay and false-alarm rate. We validate on three real-world driving datasets, demonstrating reductions in detection delay while remaining robust to heavy-tailed distributions and unknown post-change conditions.
|
| |
| TuBT9 |
502 |
| Emerging Data Science in Manufacturing 1 |
Special Session |
| Organizer: Lee, Chia-Yen | National Taiwan University |
| Organizer: Hsu, Chia-Yu | National Tsing Hua University |
| Organizer: Lugaresi, Giovanni | KU Leuven |
| Organizer: Fan, Shu-Kai S. | National Taipei University of Technology |
| Organizer: Jang, Young Jae | Korea Advanced Institute of Science and Technology |
| Organizer: Hung, Yu-Hsin | National Yang Ming Chiao Tung University |
| Organizer: Lin, Chin-Yi | University of Texas at El Paso |
| Organizer: Tsai, Tsung-Han | National Taipei University of Business |
| Organizer: Ma, Kang-Ting | National Dong Hwa University |
| |
| 14:00-14:20, Paper TuBT9.1 | |
| Modeling and Control of Fatigue-Driven Performance in Human-Centric Production Systems (I) |
|
| Lin, Junjie | KU Leuven |
| Lugaresi, Giovanni | KU Leuven |
Keywords: Human-Centered Automation, Discrete Event Dynamic Automation Systems, Behavior-Based Systems
Abstract: In traditional manufacturing systems, human performance is often modeled as a static service rate. However, worker performance is state-dependent and varies with fatigue accumulation and recovery. Ignoring this divergence can distort estimates of key performance metrics and increase the risk of undesirable outcomes such as delays and unexpected congestion. Motivated by this gap, this paper studies how rest policies can mitigate the effects of fatigue-driven stochastic processing times in labor-intensive operations and develops a framework that can be easily coupled with existing production planning and control framework. Jobs carry random attributes (e.g., load/weight) that influence both the expected service duration and the fatigue accumulation rate. We evaluate three break policies: (i) threshold-triggered breaks, (ii) periodic breaks, and (iii) rest-to-target breaks. Using a simulation-based evaluation framework, we compare policies under the same fixed workload and identify policy-specific parameter settings that balance throughput and average fatigue. Results show that actuating controlled rest can substantially shift the throughput–fatigue trade-off without changing system structure, providing an explainable baseline for extending fatigue-aware policy design to richer settings.
|
| |
| 14:20-14:40, Paper TuBT9.2 | |
| A Digital Twin Framework for Dynamic Routing of Overhead Hoist Transport in Semiconductor Fabs (I) |
|
| Japhne, Ferdinandz | Samsung Electronics |
| Park, Joonghoo | KAIST |
| Jang, Young Jae | Korea Advanced Institute of Science and Technology |
Keywords: Semiconductor Manufacturing, Discrete Event Dynamic Automation Systems, Reinforcement
Abstract: Digital twin technology is gaining significant traction in manufacturing, particularly within semiconductor fabrication, owing to its capacity to model complex operational dynamics. These simulation capabilities empower factories to anticipate and manage operational shifts prior to physical deployment. Nevertheless, current literature predominantly limits digital twin implementation to preliminary design stages, overlooking its utility for proactive decision-making in active production environments. Concurrently, existing autonomous decision-making algorithms often fail to adapt to dynamic external disruptions. To address these limitations, this study presents a framework that employs digital twin simulations to dynamically adjust decision-making processes in response to operational changes. We validate the proposed framework's effectiveness using a realistic semiconductor fab simulation. The results indicate that operational variations critically influence autonomous system performance, reinforcing the vital role of digital twin in supporting resilient and adaptive manufacturing processes.
|
| |
| 14:40-15:00, Paper TuBT9.3 | |
| An LLM-Based Recommendation System for U.S. Import Tariff (HTS) Code Classification of Machinery and Electrical Equipment (I) |
|
| Ma, Kang-Ting | National Dong Hwa University |
| Cheng, Ya-Hui | National Dong Hwa University |
| Chou, Che-Wei | Feng Chia University |
| Wang, You-Qin | National Dong Hwa University |
Keywords: AI-Based Methods, Manufacturing, Maintenance and Supply Chains, Sustainable Production and Service Automation
Abstract: In an era of increasing global trade volatility and frequent tariff changes, the United States remains a key market for Taiwan’s export-oriented small and medium-sized enterprises (SMEs). Accurate classification under the Harmonized Tariff Schedule of the United States (HTSUS) is essential for compliance and risk mitigation. However, SMEs often struggle with classification consistency and auditability due to limited expertise, staffing constraints, and the complexity of HTSUS, which also hinders digital transformation and automation. This study proposes a Vision Chain-of-Thought (Vision-CoT) framework for HTSUS classification directly from product imagery. The framework integrates Multimodal Large Language Models (MLLMs) with Retrieval-Augmented Generation (RAG) to support structured, stepwise reasoning. The system analyzes visual features while dynamically retrieving relevant regulatory provisions from an official HTSUS knowledge base. By incorporating explicit regulatory citations, the framework ensures interpretability, traceability, and alignment with legal requirements. The system was evaluated using Top-k coverage metrics and G-Eval assessments. Results show strong performance, achieving 82.72% Top-1 coverage at the 6-digit level and 81.48% at the 8-digit level. With a high-threshold quality gate based on weighted G-Eval scores, the system reached 71.6% accuracy with a low false-positive rate, indicating reliable recommendations when quality criteria are met. Positioned as a decision-support tool, the framework assists customs professionals by handling routine tasks and providing auditable reasoning. Overall, it offers a scalable and cost-effective solution for SMEs, improving classification consistency, transparency, and resilience in trade compliance.
|
| |
| 15:00-15:20, Paper TuBT9.4 | |
| Optimization of Motion Parameters for the Selective Laser Melting System (I) |
|
| Chen, Shyh-Leh | National Chung Cheng Univeristy |
| Ho, Yi-Chien | National Chung Cheng Univeristy |
Keywords: Additive Manufacturing, Optimization and Optimal Control
Abstract: Additive Manufacturing (AM) or Selective Laser Melting (SLM), in particular, has been used in industries in a vast number of applications since the 1980s [1, 2]. It has become common in various manufacturing processes, ranging from car engines, ships, and building constructions. In AM applications, getting better accuracy and surface roughness as the main criteria with time efficient are challenging tasks. This study will investigate the optimization of the motion parameters in SLM experimentally. Following the given result of this work, the machine operators will be able to understand the effects of the motion parameters, and speed up the initial setting with time efficiency.
|
| |
| 15:20-15:40, Paper TuBT9.5 | |
| Attribution Via Inconsistency Diagnosis for Virtual Metrology Maintenance (I) |
|
| Li, Chi | National Cheng Kung University |
| Hsieh, Yu-Ming | National Cheng Kung University |
| Liu, Yi-Hsin | National Cheng Kung University |
| Cheng, Fan-Tien | National Cheng Kung University |
Keywords: Factory Automation, Intelligent and Flexible Manufacturing, Machine learning
Abstract: Virtual metrology (VM) models in semiconductor manufacturing may degrade over time, yet poor online predictions are difficult to interpret because prediction error alone cannot distinguish model inadequacy from insufficient feature representation. To address this issue, we propose Attribution via Inconsistency Diagnosis (AID), a framework for VM maintenance. At the core of AID is the Data Variation Index (DVI), which quantifies local feature–response inconsistency by measuring whether similar feature vectors yield dissimilar responses. DVI is implemented as a normalized local leave-one-out error computed from neighboring reference wafers in a weighted feature space. AID jointly diagnoses failures using prediction error and DVI, assigning them to three maintenance actions. High error with low DVI suggests model-driven failure and calls for model optimization. High error with high DVI indicates inadequate feature representation and motivates DVI-guided feature reselection. If performance remains unsatisfactory after reselection, AID escalates the case for engineering investigation by providing the issue wafer with high-variation reference wafers for root-cause analysis. To instantiate AID, we adopt a k-nearest-neighbor-based DVI (k=50) and a surrogate-assisted genetic algorithm (SAGA) for feature reselection. Experiments on a semiconductor VM dataset using five independent 80/20 holdout splits—excluding test data from all training and selection processes—show SAGA-based reselection reduces test MAE by 14.3%, compared to 11.5% for a DVI-only GA variant. These results position AID not as a predictive model, but as a diagnostic framework that distinguishes model-improvable from feature-deficient failures and translates VM degradation into actionable maintenance strategies.
|
| |
| 15:40-16:00, Paper TuBT9.6 | |
| Multi-Agent Reinforcement Learning with Communication for Machine Maintenance in Flexible Flow Shops (I) |
|
| Hung, Yu-Hsin | National Yang Ming Chiao Tung University |
| Sun, Min-En | National Taiwan University |
| Lee, Chia-Yen | National Taiwan University |
Keywords: Manufacturing, Maintenance and Supply Chains, Reinforcement, Planning, Scheduling and Coordination
Abstract: Flexible flow shop (FFS) systems face preventive maintenance challenges due to machine deterioration, stochastic failures, and limited maintenance resources. Conventional methods often struggle with scalability, while decentralized multi-agent reinforcement learning (MARL) may suffer from poor coordination. This study proposes a communication-based MARL framework where machine agents exchange local information to improve cooperative maintenance decisions. CommNet and TarMAC are integrated with Multi-Agent Proximal Policy Optimization (MAPPO) and tested in a simulated FFS environment. Results show improved learning stability and higher long-term rewards, demonstrating the value of explicit communication for maintenance coordination in complex manufacturing systems.
|
| |
| TuBT10 |
Convention Hall A |
| Best Paper Award Session 2 |
|
| |
| 14:00-14:20, Paper TuBT10.1 | |
| TurboADMM: A Structure-Exploiting Parallel Solver for Multi-Agent Trajectory Optimization |
|
| Chen, Yucheng | Independent Researcher |
Keywords: Optimization and Optimal Control, Collision Avoidance, Robot Networks
Abstract: Multi-agent trajectory optimization with dense interaction networks requires solving large coupled QPs at control rates, yet existing solvers don’t simultaneously exploit temporal structure, agent decomposition, and iteration similarity. Generalpurpose QP solvers (e.g., OSQP, MOSEK) typically solve multiagent problems monolithically, resulting in poor scalability as the number of agents grows. Structure-exploiting solvers (e.g. HPIPM) leverage temporal structure through Riccati recursion but can be vulnerable to dense coupling constraints. We introduce TurboADMM, a specialized single-machine QP solver that achieves empirically near-linear scaling with respect to agent count on our benchmarks. The solver is built through systematic co-design of three complementary components. First, ADMM decomposition creates per-agent subproblems that can be solved in parallel. Second, a Riccati-based warmstart exploits temporal structure to provide high-quality primal–dual initialization for each agent’s QP. Third, parametric QP hotstart in qpOASES reuses similar KKT factorizations across successive ADMM iterations.1 Ablation studies demonstrate the multiplicative value of this integrated design: coldstart QP with ADMM (BaseADMM) spends 14.3 seconds for 14 agents trajectory optimization, hotstart reduces this to 121ms, and adding Riccati warmstart yields final time of 96ms. On our multi-agent collision avoidance benchmarks (2-14 agents, 20-step horizons), TurboADMM achieves up to 14.8× speedup over OSQP and up to 22.0× speedup over MOSEK, with the largest gains at higher agent counts. HPIPM succeeds for 2- agent problems but we observed convergence issues in 4+ agent scenarios despite feasible initialization. A C++ implementation is provided through this anonymous shareable GitFro
|
| |
| 14:20-14:40, Paper TuBT10.2 | |
| Conflict-Based Search for Multi-Agent Path Finding with Elevators (I) |
|
| He, Haitong | Harbin Institute of Technology |
| Wu, Xuemian | Shanghai Jiao Tong University |
| Zhao, Shizhe | Shanghai Jiao Tong University |
| Ren, Zhongqiang | Shanghai Jiao Tong University |
Keywords: Motion and Path Planning, Planning, Scheduling and Coordination
Abstract: This paper investigates Multi-Agent Path Finding with Elevators (MAPF-E), which seeks conflict-free paths for multiple agents whose start and goal locations may be on different floors, and the agents can use elevators to travel between floors. The existence of elevators complicates the interaction among the agents and introduces new planning challenges. On the one hand, elevators can cause many conflicts among agents due to their relatively long traversal time across floors, especially when many agents need to reach a different floor. On the other hand, the planner has to reason in a larger state space that includes elevator states in addition to agent locations. To address these challenges, this paper proposes CBS-E, an extension of Conflict-Based Search (CBS) to solve MAPF-E optimally. CBS-E introduces new concepts of elevator constraints to incorporate the state of the elevator into the conflict resolution process of CBS. We also extend Multi-Valued Decision Diagrams (MDDs) to CBS-E, referred to as MDD-E. MDD-E enables the intelligent selection of conflicts to resolve in each iteration, thereby improving the runtime efficiency. The results show that our CBS-E often doubles or triples the success rates of the baseline, and the proposed conflict reasoning can reduce the number of iterations of CBS-E by up to an order of magnitude.
|
| |
| 14:40-15:00, Paper TuBT10.3 | |
| Safe Autonomous Driving Via Data-Driven Robust Model Predictive Control (I) |
|
| Fang, Shiming | Binghamton University |
| Li, Xilin | Binghamton University |
| Wu, Changzhi | Chongqing Normal University |
| Yu, Kaiyan | Binghamton University |
Keywords: Robust/Adaptive Control, Autonomous Vehicle Navigation, Collision Avoidance
Abstract: Safe navigation in autonomous driving requires reliable obstacle avoidance despite modeling errors and sensing disturbances. Model predictive control (MPC) is widely used for this task, but simplified prediction models can lead to deviations between predicted trajectories and the actual vehicle motion. This paper presents a data-driven robust MPC framework for autonomous driving that improves safety under such uncertainties. Vehicle dynamics are described using a quasi-linear parameter-varying (LPV) bicycle model, where discrepancies between the model and the real system are captured as additive disturbances with unknown distributions. To handle these uncertainties, a distributionally robust chance-constrained formulation based on a Wasserstein ambiguity set is introduced, allowing safety constraints to be enforced probabilistically without assuming specific disturbance distributions. The resulting control problem is formulated as a quadratic program that can be solved efficiently for real-time implementation. Simulation studies and experiments on a scaled autonomous vehicle demonstrate that the proposed approach increases obstacle clearance and improves trajectory tracking consistency compared with nonlinear MPC and conventional LPV-MPC controllers.
|
| |
| 15:00-15:20, Paper TuBT10.4 | |
| When 99.8% Is Not Enough: Closed-Loop Evaluation Exposes the Metric Gap in Learned Microgrid Control |
|
| White, Meaghan | University of Queensland |
| Nadarajah, Mithulan | University of Queensland |
Keywords: Machine learning, Power and Energy Systems automation, Smart Grids
Abstract: Offline predictive metrics are common selection criteria for learned power system controllers, yet deployment is closed-loop, with each action changing the next state the controller must face. We evaluate this gap through 480 microgrid reconfiguration rollouts spanning 4 architectures, 4 training regimes and 30 operational scenarios. We find offline selection unreliable. Despite offline F1 scores clustered between 0.986 and 0.998, closed-loop cost ratios range from 0.17× to 1319× the genetic algorithm baseline, with 62.9% of rollouts exceeding the catastrophic threshold R ≥ 5. F1 is only weakly correlated with deployment cost (ρ = −0.33), while state mismatch rate has effectively no predictive value (ρ = 0.06). Per-step analysis further reveals single-step cost spikes above 3600× even in rollouts with acceptable aggregate cost. Targeted training perturbations show that robustness to endogenous state mismatch is critical, but architecture determines whether that robustness is induced, preserved or disrupted. These results provide a large-scale empirical characterization of the metric gap in learned microgrid control, showing not only that offline F1 can mislead, but that it can obscure orders-of-magnitude differences in deployment cost, transient behavior and post-divergence recovery. We therefore recommend selecting learned controllers by closed-loop utility over diverse scenarios, with tail-risk, transient and recovery metrics treated as deployment filters rather than afterthoughts.
|
| |
| 15:20-15:40, Paper TuBT10.5 | |
| Propeller-Assisted Robust 3D Hopping Robot with Hierarchical Force Allocation (I) |
|
| Zhang, Chuhan | Guangdong Technion Israel Institute of Technology |
| Zhang, Hongbo | The Chinese University of Hong Kong |
| Chen, Yanlin | The Chinese University of Hong Kong |
| Tang, Yunxi | The Chinese University of Hong Kong |
| Liu, Yunhui | Chinese University of Hong Kong |
| Liu, Mingyi | Guangdong Technion - Israel Institution of Technology |
| Chu, Xiangyu | The Chinese University of Hong Kong |
Keywords: Motion Control, Optimization and Optimal Control
Abstract: Monopedal hopping robots are conceptually simple but highly dynamic and inherently unstable. Achieving robust 3D hopping is still difficult because ground reaction forces are available only during the short stance phase, while the robot is underactuated in flight. A key unresolved issue is how to improve flight-phase control authority. Propeller assistance provides a promising solution, but it requires careful coordination of leg-generated contact forces and propeller thrusts across stance and flight. This paper presents a propeller-assisted 3D monopedal hopping robot with an active 3-RSR parallel leg and a trunk-mounted tri-rotor for auxiliary attitude regulation. To address the force coordination challenge, we propose a Hierarchical Force Allocation (HFA) framework based on a single rigid body (SRB) model. The leg generates the main stance contact wrench, while the tri-rotor provides auxiliary attitude regulation, compensating the residual attitude moment in stance and maintaining attitude during flight. Real-robot experiments in indoor and outdoor scenarios demonstrate sustained 3D hopping, including terrain transitions and impulsive push recovery, validating robustness under unmodeled contact and external disturbances.
|
| |
| 15:40-16:00, Paper TuBT10.6 | |
| CopilotTele: A Shared-Autonomy Teleoperation System for Bimanual Manufacturing Assembly with Risk-Guided Arbitration (I) |
|
| Chen, Chenrui | The Hong Kong University of Science and Technology (Guangzhou) |
| Zhang, Xudong | The Hong Kong University of Science and Technology (Guangzhou) |
| Wang, Xueting | Rochester Institute of Technology |
| Dengxiong, Xiwen | Rochester Institute of Technology |
| Xiao, Qinqin | ZRbot |
| Jing, Ke | ByteDance |
| Zhang, Yunbo | Hong Kong Institute of Science & Innovation, CAS, Hong Kong, China |
Keywords: Human-Centered Automation, Collaborative Robots in Manufacturing, Telerobotics and Teleoperation
Abstract: Industrial bimanual assembly remains difficult to automate because execution is multi-stage, coordination-intensive, and sensitive to local variation. In practice, pure teleoperation provides flexibility but is labor-intensive, whereas full autonomy is often too brittle to be trusted across the entire workflow. This paper presents CopilotTele, a shared-autonomy system for manufacturing-inspired bimanual assembly that enables autonomy-by-default execution with on-demand human intervention. CopilotTele learns short-horizon bimanual action priors offline from dual-operator VR demonstrations and deploys them online through an asynchronous runtime that integrates near-term action prediction, policy-side risk estimation, authority arbitration, and fail-safe behavior. During execution, the system autonomously performs near-term action chunks by default, while keeping the operator in the loop for proactive intervention or takeover when execution risk rises. We evaluate CopilotTele on an industrially inspired wire-harness assembly workflow consisting of six subtasks, with 50 trials per condition. Compared with manual single-operator teleoperation, the proposed system reduces average completion time by 13.7%, increases tasks per hour by 16.0%, and lowers retry count by 46.3%, while requiring only 0.88 intervention events and 1.09,s of intervention time per task on average. These results show that shared autonomy improves the efficiency and consistency of single-operator wire-harness assembly while recovering a substantial fraction of the coordination benefit of two-human collaboration at lower human cost.
|
| |
| TuCT1 |
201 |
| Data Analytics and Optimization for Smart Industry 3 |
Regular Session |
| |
| 16:15-16:35, Paper TuCT1.1 | |
| SSIL: Self-Supervised Imitation Learning for End-To-End Driving |
|
| Park, Jinbok | Sungkyunkwan University |
| Lee, Jinkyu | Handong Global University |
| Back, Muhyun | Handong Global University |
| Han, Hyun Min | Handong Global University |
| Ma, Tianwei | Texas a & M University-Corpus Christi |
| Won, Sang Min | Sungkyunkwan University |
| Hwang, Sung Soo | Handong Global University |
| Chun, Il Yong | Sungkyunkwan University |
Keywords: Autonomous Vehicle Navigation, Deep Learning in Robotics and Automation, Machine learning
Abstract: In autonomous driving, the end-to-end (E2E) driving approach that predicts vehicle control signals directly from sensor data is rapidly gaining attention. To learn a safe E2E driving system, one needs an extensive amount of driving data and human intervention. Vehicle control data is constructed by many hours of human driving, and it is challenging to construct large vehicle control datasets. Often, publicly available driving datasets are collected with limited driving scenes, and collecting vehicle control data is only available by vehicle manufacturers. To address these challenges, this paper proposes the first self-supervised learning framework, Self-Supervised Imitation Learning (SSIL), for E2E driving. The proposed SSIL framework can learn vision-based E2E driving networks without using driving command data or a pre-trained model. To construct pseudo steering angle data, proposed SSIL predicts a pseudo target from the vehicle’s poses at the current and previous time points that are estimated with light detection and ranging sensors. In addition, we propose a new cross-attention-based conditioning approach (CACA) for a vision encoder in E2E driving, where a high-level instruction serves as the conditioning signal for visual information. Our numerical experiments with three different benchmark datasets demonstrate that the proposed SSIL framework achieves very comparable E2E driving accuracy with the supervised learning counterpart. Furthermore, the proposed pseudo-label predictor outperformed an existing one using proportional integral derivative controller, and proposed CACA achieved superior performance over existing conditioning approaches.
|
| |
| 16:35-16:55, Paper TuCT1.2 | |
| UniST: A Pre-Trained Unified Spatio-Temporal Multi-Task Framework for Network Traffic Analysis |
|
| Zhang, Longfei | Nanjing University of Science and Technology |
| Xia, Pengcheng | Nanjing University of Science and Technology |
| Sun, Xiang | Nanjing University of Science and Technology |
| Ni, Yiyang | Jiangsu Second Normal University |
| Li, Jun | Southeast University |
Keywords: Big-Data and Data Mining, AI-Based Methods, Machine learning
Abstract: Spatio-temporal traffic data analysis is essential for 5G and intelligent network systems. However, existing methods often employ isolated models, neglecting shared spatio-temporal correlations and incurring prohibitive computational overhead. To address these issues, we propose UniST, a unified spatio-temporal multi-task learning framework for simultaneous prediction, imputation, and anomaly detection. The framework features a dual-pathway mechanism: a Spatio-Temporal Graph Embedding path explicitly models data with physical topology, while a Frequency Domain Enhancement path captures periodic variations in graph-free scenarios. These extracted features are processed by a shared GPT-2-based Transformer backbone to capture evolving temporal patterns. To mitigate gradient conflicts during joint multi-task optimization, we incorporate the Projecting Conflicting Gradients (PCGrad) algorithm to promote balanced convergence. Extensive experiments on four real-world datasets show that UniST achieves competitive and balanced performance across multiple tasks while reducing the overall training cost compared with deploying isolated task-specific models. Furthermore, few-shot experiments confirm its robust cross-domain transferability.
|
| |
| 16:55-17:15, Paper TuCT1.3 | |
| Robust Kinetic Parameter Extraction for Perovskite Oxygen Carriers Via Convex Optimization (Work-In-Progress) |
|
| Zhang, Hao | Northeastern Univ |
| Wang, Kun | Northeastern Univ |
| Zhang, Lingxi | Northeastern Univ |
| Wu, Haoyang | Northeastern Univ |
Keywords: Big-Data and Data Mining, Assembly
Abstract: Abstract—This study presents a robust single-curve kinetic analysis framework grounded in convex optimization theory to extract activation energy (Ea) from thermogravimetric data. By transforming nonlinear rate equations into a strictly convex loss function, our method guarantees unique global minima, eliminating local optima risks inherent in traditional fitting. Applied to CaFeO3 decomposition, the approach demonstrated high stability against sample mass variations and sensitivity to atmospheric changes. Comparative screening of synthesis methods revealed an 8.5% Ea increase for optimized protocols, confirming enhanced thermal stability. Preliminary results demonstrate the model’s robustness. Future work will focus on extending this framework to multi-cycle stability analysis and reactor-scale validation.
|
| |
| 17:15-17:35, Paper TuCT1.4 | |
| A UMAP-Embedded and MBO-Optimized K-Means Clustering Method for Aluminum Electrolysis Time-Series Data |
|
| Li, Aoyu | Northeastern University, Shenyang Aluminum and Magnesium Engineering and Research Institute Company Limited |
| Ma, Enjie | Shenyang Aluminum and Magnesium Engineering and Research Institute Company Limited |
Keywords: Big-Data and Data Mining, Machine learning, Factory Automation
Abstract: In the aluminum electrolysis production process, aluminum electrolytic cells accumulate a large amount of high-dimensional time-series data, which can reflect the operating status of the electrolytic cells. However, when using traditional clustering algorithms to analyze these time-series data, the high dimensions often lead to similar distances between different samples, resulting in poor discrimination of the clustering results. Therefore, this paper proposes a K-means clustering method that combines the Unified Manifold Approximation and Projection (UMAP) method with the Monarch Butterfly Optimization (MBO) algorithm. To avoid measuring the distances between high-dimensional time-series data of aluminum electrolytic cells based on the original distances, this paper first extracts multiple features of the original data, and then uses the UMAP method to project the high-dimensional feature vectors into a lower-dimensional feature space. Subsequently, the MBO algorithm is used to search for the initial clustering centers that minimize the within-cluster sum of squares in the feature space, which are then used as the input of the K-means algorithm, and finally obtain the sample clustering results. Through comparative experiments on real data from multiple electrolytic cells in an aluminum electrolysis plant, the proposed method has better clustering effects and stability, and the clustering results can assist the production managers of the aluminum electrolysis plant in making production decisions.
|
| |
| 17:35-17:55, Paper TuCT1.5 | |
| Spatial Field Prediction for Multicomponent Physics Simulation Using Structured Penalized Regression |
|
| Chen, Kangan | University of Wisconsin-Madison |
| Wang, Andi | University of Wisconsin-Madison |
Keywords: Big-Data and Data Mining, Machine learning, Modelling, Simulation and Validation of Cyber-physical Energy Systems
Abstract: High-fidelity physics simulations of multi-compartment engineering systems produce spatially distributed outputs that exhibit multiscale structure, geometric symmetry, and smoothness, yet existing surrogate modeling approaches address at most one or two of these properties in isolation. We propose a structured penalized regression framework for surrogate modeling of spatially distributed simulation outputs in multi-compartment systems. The model decomposes the coefficient field into a global component shared across all compartments and local components specific to each compartment type, with penalty terms representing spatial smoothness, positional symmetry, and within-compartment symmetry of the system geometry. The resulting penalized least squares problem admits a closed-form solution, with hyperparameters selected by cross-validation. On a prismatic high-temperature gas-cooled reactor fuel assembly with 19 compartments and 7 design parameters, the proposed method reduces test RMSE by 17% over pointwise OLS and 7% over functional linear modeling at n=80. A group lasso enables the selection of the effective inputs in the model, which significantly improves the model performance when the interactive effects are incorporated in the model and when the sample size are relatively small.
|
| |
| 17:55-18:15, Paper TuCT1.6 | |
| Optimization of Smart Building Performance Based on the Smart Readiness Indicator |
|
| Rana, Giuseppe Rocco | Polytechnic University of Bari |
| Roccotelli, Michele | Polytechnic University of Bari |
| Mangini, Agostino Marcello | Politecnico Di Bari |
| Fanti, Maria Pia | Polytechnic University of Bari |
Keywords: Building Automation, Smart Home and City, Optimization and Optimal Control
Abstract: This paper presents a rigorous mathematical model and an innovative methodology for optimizing smart building performance through the Smart Readiness Indicator (SRI), a key tool that can help address the energy inadequacy of Europe’s building stock, which requires intelligent retrofits to meet nZEB standards and emission reductions by 2050. The SRI (EU Directive 2018/844) assesses adaptability and energy efficiency, but current audits overlook budget constraints and regional differences. The aim of this work is to overcome the gap through a novel ILP (Integer Linear Programming) approach to find out optimal configurations of commercial devices in order to maximize the SRI within economic limits. The proposed model quantifies improvements and transforms the SRI into a planning tool for Decision Support Systems (DSS) enabling automated matching between owners and manufacturers for efficient retrofits aligned with EU sustainability goals. The article concludes with future integrations with the Building Information Model (BIM) for automated assessments toward intelligent and user-centric buildings.
|
| |
| TuCT2 |
202 |
| Frontier Technology of Industrial & Systems Engineering 3 |
Regular Session |
| |
| 16:15-16:35, Paper TuCT2.1 | |
| Statistical Analysis of Distributed Multi-Robot Routing Via Heartbeat Protocols and Intention Sharing |
|
| Woodman, Noelle | California Polytechnic State University |
| Diaz Alvarenga, Carlos | California Polytechnic State University, San Luis Obispo |
Keywords: Planning, Scheduling and Coordination, Agent-Based Systems, Robot Networks
Abstract: The use of multiple robotic agents for area coverage missions has recently received considerable attention from the robotics community. By employing a team of agents, these systems can increase robustness and efficiency. Furthermore, online methods allow the team to rapidly adapt to sudden changes in the environment or team composition because the need to idle the robot team and recompute a global solution is removed. This work demonstrates the use of distributed heartbeat protocols to ensure robust multi-agent area coverage. Through simulated benchmarks, we empirically show that our approach performs comparably to traditional distributed formulations in terms of efficiency and robustness. We benchmark our protocol in a range of scenarios and provide statistical analysis showing that the proposed method, despite being a much simpler implementation, performs comparably to commonly used solution paradigms for multi-robot routing when accounting for vehicle failure. The use of multiple robotic agents for area coverage missions has recently received considerable attention from the robotics community. By employing a team of agents, these systems can increase robustness and efficiency. Furthermore, online methods allow the team to rapidly adapt to sudden changes in the environment or team composition because the need to idle the robot team and recompute a global solution is removed. This work demonstrates the use of distributed heartbeat protocols to ensure robust multi-agent area coverage. Through simulated benchmarks, we empirically show that our approach performs comparably to traditional distributed formulations in terms of efficiency and robustness. We benchmark our protocol in a range of scenarios and provide statistical analysis.
|
| |
| 16:35-16:55, Paper TuCT2.2 | |
| A Reinforcement Learning-Enhanced Adaptive Genetic Algorithm for Aircraft Assembly Scheduling with Multi-Mode and Dynamic Resource Constraints |
|
| Chen, Chen | National Frontiers Science Center for Industrial Intelligence and Systems Optimization, Northeastern University |
| Dong, Yun | Key Laboratory of Data Analytics and Optimization for Smart Industry (Northeastern University) |
| Zhao, Ren | Liaoning Key Laboratory of Manufacturing System and Logistics Optimization |
| Sun, Defeng | Liaoning Engineering Laboratory of Data Analytics and Optimization for Smart Industry |
Keywords: Planning, Scheduling and Coordination, Compliant Assembly, Process Control
Abstract: To address a project scheduling problem with multi-mode execution and dynamic resource replenishment in aircraft assembly, this paper proposes a two-stage adaptive optimization framework. In the first stage, a hybrid genetic algorithm integrated with Q-learning is employed to select task mode and determine schedule. In the second stage, an adaptive large neighborhood search guided by Q-learning is applied for fine-tuning task sequence. Experimental results on benchmark datasets demonstrate that the proposed method significantly outperforms representative algorithms in terms of makespan. Real-world case studies further validate its effectiveness under multi-mode constraints and dynamic resource replenishment. This paper presents preliminary research progress, and a more complete study with extended experiments and in-depth analysis will be reported in future work.
|
| |
| 16:55-17:15, Paper TuCT2.3 | |
| Production–Inventory Planning in Equipment Manufacturing with Shared Molds and Sequence-Dependent Setups |
|
| Liu, Sibo | Northeastern University |
| Wang, Gongshu | Northeastern University |
Keywords: Planning, Scheduling and Coordination, Inventory Management, Manufacturing, Maintenance and Supply Chains
Abstract: Within the integrated manufacture-circulation industrial system, balancing operational efficiency with inventory costs in equipment manufacturing is complicated by shared resources and sequence-dependent operations. This paper studies a production–inventory planning problem that coordinates upstream raw-material replenishment, integer batch production, shared-mold allocation, and order delivery over a finite horizon. We formulate a mixed-integer programming (MIP) model that captures mold exclusivity, sequence-dependent changeovers, pooled minimum order quantity requirements, and the conversion from continuous raw-material weights to admissible production batches. These features ensure that procurement and inventory decisions remain consistent with executable shift-level production plans. To address this computational difficulty, we develop a column-generation-based kernel search (CG-KS) matheuristic. Column generation decomposes production decisions by line and coordinates them through a master problem, while the generated column pool and relaxation information guide kernel search. Computational experiments show that CG-KS is competitive on small instances and becomes more effective as instance size increases. On the large instance sets, CG-KS reduces total cost by 6.47% on average and by up to 10.76% compared with CPLEX under a 3600 s time limit.
|
| |
| 17:15-17:35, Paper TuCT2.4 | |
| Reinforcement Learning–Guided Integer Optimization for Production–Inventory–Distribution Planning in Manufacture-Circulation Industrial System |
|
| Cui, Jiawen | Northeastern University |
| Tang, Lixin | Northeastern University |
| Wang, Gongshu | Northeastern University |
| Yang, Yang | Northeastern University |
Keywords: Planning, Scheduling and Coordination, Logistics, Manufacturing, Maintenance and Supply Chains
Abstract: In the manufacture-circulation industrial system (MCIS), the integration of production, inventory, and distribution planning is critical for operational efficiency and cost minimization. This paper addresses the challenges of coordinating production and distribution decisions in a MCIS with multiple distribution centers and order-specific time windows. We propose a mixed-integer linear programming (MILP) model that jointly optimizes production lot-sizing, inventory levels, and distribution decisions while considering machine capacity, storage constraints, and delivery time windows. To solve the computationally challenging problem, we develop a matheuristic framework combining a relax-and-fix (RF) approach for generating feasible solutions and a reinforcement learning–guided fix-and-optimize (FO) algorithm for iterative solution improvement. Computational experiments on real-world and synthetic instances compare the proposed method with CPLEX and a conventional rolling-horizon style RFO baseline. This work contributes to the literature by addressing the integration of multi-stage production and distribution systems with customer order time windows, providing scalable solutions for complex manufacturing problems.
|
| |
| 17:35-17:55, Paper TuCT2.5 | |
| Robust Scheduling on Parallel Semi-Continuous Batching Machines with Incompatible Job Families |
|
| Sun, Yu | Northeastern University |
| Meng, Ying | Northeastern University |
| Tang, Lixin | Northeastern University |
Keywords: Robust Manufacturing, Intelligent and Flexible Manufacturing
Abstract: Motivated by walking beam heating systems, this paper studies the Parallel Semi-Continuous Batch Scheduling Problem with Incompatible Job Families (PSCBSP-IF) under processing time uncertainty. In this problem, the processing duration of each batch is jointly determined by the longest job in the batch, the batch size, and the machine capacity. The objective is to minimize the worst-case makespan by determining machine assignments and family-compatible batches before uncertainty is realized. To address this problem, we formulate a two-stage robust optimization model with a budgeted uncertainty set and develop two complementary solution approaches. The first approach is a strengthened inexact column-andconstraint generation (i-C&CG) framework, in which familylevel cuts are incorporated by exploiting problem-specific structural properties. The second approach is a Memetic framework that integrates assignment-only encoding, scenariopool-guided decoding, bottleneck-guided mutation, and tabubased intensification to obtain high-quality robust schedules within limited computation time. Computational experiments on generated steel-production instances show that the proposed Memetic framework consistently outperforms both the i-C&CG approach and its tabu-free variant across all tested uncertaintybudget levels, while maintaining stable performance under different budget settings.
|
| |
| 17:55-18:15, Paper TuCT2.6 | |
| Modeling and Statistical Cost Formulation under Metabolic Energy Constraints for Human-Powered Rover Navigation |
|
| Callata Suxo, Alvaro Daniel | The Mobi Company Llc |
| Acarapi Roca, Adrián Fernando | Independent Researcher |
| Velasquez Enriquez, Marcelo | Independent Researcher |
| Gaytan Bustamante, Angel Gael | Universidad Del Valle De Atemajac |
| Diaz Palacios, Fabio | Universidad Del Valle De Atemajac, Universidad Privada De Valle |
Keywords: Planning, Scheduling and Coordination, Probability and Statistical Methods, Human-Centered Automation
Abstract: This study presents a probabilistic optimization framework for stochastic route planning under metabolic energy constraints. The model is developed in the context of the NASA Human Exploration Rover Challenge (HERC). To address the limitations of deterministic planning, this approach applies machine learning algorithms to historical competition data. Specifically, it uses Hierarchical Bayesian Modeling, Gradient Boosting, and Learning to Rank to estimate probability distributions for obstacle success rates, traversal times, and metabolic energy costs. These statistical estimates are incorporated into a discrete probabilistic algorithm that evaluates potential routes based on expected score, outcome variance, and overall energy feasibility. The routing model was validated through a Monte Carlo simulation with 10,000 trials. The results demonstrate that accounting for the cumulative time delays of unpredictable failures improves route efficiency compared to standard planning heuristics. Furthermore, the methodology generalizes to other automation domains that operate under strict resource limits. By relating human fatigue to nonlinear battery discharge curves, this framework adapts routing and scheduling principles for autonomous vehicle fleets and drone delivery networks.
|
| |
| TuCT3 |
203 |
| Systems Technology of Industrial Automation and Control 3 |
Regular Session |
| |
| 16:15-16:35, Paper TuCT3.1 | |
| A Multi-Agent Framework for Robotic HMI Testing Using Multimodal Large Language Models |
|
| Wan, Dingyuan | Volkswagen AG |
| Rasheed, Umair | Volkswagen Group of America |
| Kovacs, David | Volkswagen AG |
| Vogt, Henning Sebastian | Volkswagen AG |
| Ansary, Jamal | VW |
| Konwalinka, Stefan | Volkswagen Nutzfahrzeuge |
| Pannek, Jürgen | TU Braunschweig |
Keywords: Agent-Based Systems, Cognitive Automation, Motion Control
Abstract: Traditional manual or template based validation methods for modern automotive Human–Machine-Interfaces (HMIs) struggle to keep pace with modern software development. This paper introduces an autonomous robotic HMI testing framework which is enabled by a multi-agent task planner and a robotic manipulator for physical interaction. The framework is designed to autonomously execute unseen multi-step HMI test cases. We utilized a planner to interpret test intents, while a grounding agent localizes the targeted UI elements, enabling adaptive behavior under UI modifications without any scripted templates. A cartesian stiffness controller is applied to a torque-controlled robot arm, ensuring precise interaction with a sensitive touchscreen. The system is validated on automotive HMI test cases integrated into a vehicle Hardware-in-the-Loop (HiL) testbench, demonstrating robust performance with a high success rate on unseen UI layouts. The results indicate that combining multi-agent reasoning with robotic physical interaction offers a scalable way towards modern HMI validation with reduced maintenance effort compared to static record-and-replay baselines.
|
| |
| 16:35-16:55, Paper TuCT3.2 | |
| Horizontal Beam Placement with Occupancy-Belief Space Planning |
|
| Li, Pusong | University of Illinois at Urbana-Champaign |
| Nagi, Rakesh | University of Illinois, Urbana-Champaign |
Keywords: Model Learning for Control, Deep Learning in Robotics and Automation, Compliant Assembly
Abstract: Object manipulation for construction assembly using a crane is a control problem with highly challenging dynamics, merging contact-rich manipulation, high dynamics uncertainty from the outdoors environment, coarse actuators, and an underactuated system. The inevitable contact from assembly presents both a potential to exacerbate small errors as well as an opportunity to limit uncertainty by restricting the possible configuration space. For a controller to take advantage of this beyond simple lift-up and lay-down operations, a system that is capable of considering how uncertainty evolves under contact dynamics is required. We approach this problem by introducing the novel concepts of textit{occupancy-belief space} (a space representing the probability the payload occupies each point in space at a given timestep) and textit{occupancy-belief dynamics} (the dynamics of the aforementioned space). We learn the occupancy-belief dynamics of the system under consideration using a visual foresight inspired model, then perform model predictive control using this learned model, with the objective of creating possible occupancy that corresponds to the payload in its intended target position. We evaluate this system in simulation and on a scale test model on the problem of I-beam assembly, specifically, aligning and inserting a horizontal box member between flanges of two opposing vertical I-beams.
|
| |
| 16:55-17:15, Paper TuCT3.3 | |
| Multi-Worker Assembly Line Rebalancing with Relevance-Guided Configuration Preservation |
|
| Vinetti, Martina | Chalmers University of Technology |
| Roselli, Sabino Francesco | Chalmers University of Technology |
| Fabian, Martin | Department of Electrical Engineering |
Keywords: Assembly, Human-Centered Automation, Optimization and Optimal Control
Abstract: In assembly line balancing, tasks are assigned to stations in order to satisfy a required cycle time. When production conditions change, the line must be rebalanced by modifying the current task allocation, typically aiming to move as few tasks as possible between stations. Similarity measures are commonly used to control such changes, but they generally evaluate configuration preservation by treating all tasks equally, which may not reflect their different practical importance. In this work, a emph{pruned Mean Similarity Factor} is proposed for assembly line rebalancing, evaluating similarity only over a subset of structurally relevant tasks identified through a relevance score. The proposed measure is integrated into a compact mixed-integer linear programming (MILP) formulation that considers practical aspects of manual assembly, specifically workload balance, ergonomic exposure, multi-worker stations, and positional constraints. Computational experiments on extended benchmark instances derived from the literature show that the proposed approach can obtain optimal rebalancing solutions within reasonable computational times, while maintaining high task colocation and balanced workload and ergonomic distributions. In particular, focusing the similarity evaluation on relevant tasks helps reduce the computational effort.
|
| |
| 17:15-17:35, Paper TuCT3.4 | |
| Lightweight Perception-Guided Docking and Suboptimal Control for Resource-Constrained Modular Micro-UAVs |
|
| Lu, Haolin | Wuhan University |
| Guo, Zhao | Wuhan University |
Keywords: Motion Control, Optimization and Optimal Control, Embedded Systems in Meso, Micro and Nano Scale
Abstract: Modular micro-UAV systems offer enhanced redundancy and operational flexibility, yet face critical challenges arising from severe on-board resource constraints. This paper proposes an integrated lightweight framework that jointly addresses autonomous docking and array trajectory tracking. For docking perception, a multi-zone Time-of-Flight (ToF) sensing scheme employing an O(n) region-growing clustering algorithm enables real-time target identification from 8times8 depth images. For array control, a suboptimal Model Predictive Control (MPC) method based on the Alternating Direction Method of Multipliers (ADMM) combined with Riccati recursion reduces computational latency from 7,ms to 0.05,ms. The unified system architecture enables seamless transition from individual docking to collective array flight. The iteration count is selected to preserve the closed-loop stability guarantees established for suboptimal ADMM-MPC, and extensive simulations validate the linear model assumptions and compare ToF against motion capture systems. Experiments on modified 70,g quadrotor modules demonstrate a docking success rate exceeding 80% within a 0.5,m initial separation and robust array shape maintenance within a 6% error bound.
|
| |
| 17:35-17:55, Paper TuCT3.5 | |
| Nonlinear Dynamic Modeling and Control of 2-DOF Tunable Prism Driven by Dielectric Elastomer (I) |
|
| Lv, Jianming | National University of Defense Technology |
| Huajie, Hong | National University of Defense Technology |
| Zihao, Gan | National University of Defense Technology |
| Zhuoqun, Hu | National University of Defense Technology |
| Zhang, Meng | National University of Defense Technology |
Keywords: Motion Control
Abstract: Dielectric elastomer-driven tunable prisms are novel liquid optical devices that enable two degree-of-freedom beam deflection. However, their periodic dynamic response exhibits hysteresis, creep, and rate-dependent viscoelastic behavior, which makes modeling the nonlinear dynamics and precise control challenging. This paper analyzes the dynamics and kinematics of a biconical tunable prism. Based on the principle of non-equilibrium thermodynamics, a nonlinear dynamic model of the tunable prism is established, taking into account the geometric configuration of the prism, the viscoelasticity of the DE, and the liquid resistance. We developed feedforward controller based on the model and combined with a PI feedback controller to form the hybrid controller. Experimental results show that the theoretical model can accurately predict the complex dynamic response of the system, with root mean square error of prediction lower than 5.59%. Single and dual axis deflection angle tracking root mean square error does not exceed 1.25%. These results demonstrate the effectiveness of the proposed dynamic modeling and tracking control method.
|
| |
| 17:55-18:15, Paper TuCT3.6 | |
| High-Precision Tracking Control of Multi-Axis Electro-Optical Systems Based on Accurate Identification of Dynamic Parameters (I) |
|
| Xu, Kehui | National University of Defense Technology |
| Xiao, Mubang | National University of Defense Technology |
| Xie, Xin | National University of Defense Technology |
| Wen, Zhijie | National University of Defense Technology |
| Fan, Shixun | National University of Defense Technology |
| Fan, Dapeng | National University of Defense Technology |
Keywords: Motion Control, Industrial and Service Robotics
Abstract: High-precision tracking control is essential for multi-axis electro-optical systems (EOSs) to defense low-slow-small (LSS) moving targets. The inherent rotational inertia and nonlinear joint friction degrade the accuracy in both dynamic and low-speed tracking scenarios, necessitating model-based feedforward controllers that require precise parameter identification. To solve this issue, this study proposes a motor-coupled dynamic parameters set (MCDPS) identification framework that incorporates physical feasibility constraints, specifically designed for EOSs that lack joint torque sensors. An asymmetric nonlinear friction model is proposed to characterize the nonlinearity and direction-dependence of joint friction. Then, an Iterative Composite Least Squares (ICLS) algorithm is employed to identify the inertial and nonlinear friction parameters through iterative linear regression. A backpropagation neural network (BPNN) using joint position and load data as inputs is also proposed to further reduce the modeling error of joint nonlinearities. Experiments are performed using a multi-axis EOS to track typical joint trajectory and LSS simulated target. Results demonstrate that the proposed ICLS algorithm combined with the position- and load-related BPNN outperforms traditional identification methods. Additionally, the model-based feedforward controller is compared with proportional integral derivative (PID), extended kalman filter (EKF) and extended state observer (ESO) with sliding-mode control (SMC) approaches, showing a 21.29% improvement in tracking accuracy.
|
| |
| TuCT4 |
204 |
| Enabling Technology of Industrial Intelligence 3 |
Regular Session |
| |
| 16:15-16:35, Paper TuCT4.1 | |
| Biomimetic Visual Perception Network for Industrial Image Anomaly Detection (I) |
|
| Tang, Siang | Huazhong University of Science and Technology |
| Yang, Hua | Huazhong University of Science and Technology |
| Chen, Jiankui | Huazhong University of Science and Technology |
| Yin, Zhouping | Professor, School of Mechanical ScienceandEngineering, Huazhong University of Science&Technology, China |
Keywords: Computer Vision for Manufacturing, Semiconductor Manufacturing, Intelligent and Flexible Manufacturing
Abstract: Anomaly detection in the industrial manufacturing process is important for controlling product quality. In real-world scenarios, anomalies manifest in a number of diverse and unforeseen ways and can be categorized into two main types: structural anomalies, such as scratches and breakages, and logical anomalies, such as missing components and dislocations. Current mainstream methods tend to focus on knowledge generalization but are inadequate for detailed and logical detection tasks. To address this issue, this study proposes Biomimetic Visual Perception Network (BVPN), a biomimetic design-based novel paradigm that simulates biological vision and human logical discrimination with detailed distillation dependencies and semantic component information. BVPN introduces three convolution-based networks to extract information from different regions, imitating biological central, para-central, and peripheral vision, and facilitates the integration of disparate information through the process of knowledge distillation. On the basis, BVPN develops two modules based on the student-teacher network: Focus Module, for local detailed detection, and Peripheral Module, for global context detection. Moreover, BVPN proposes the Logical Module, a clustering-based segmentation branch, assessing logical anomalies by the number and color information of segmentation components. Through a fused structure highly aligning with the perception mode of human system, BVPN achieves excellent performance in the detection of structural and logical anomalies. Extensive experiments with mainstream anomaly detection datasets and real-world inkjet printing dataset demonstrate that BVPN outperforms state-of-the-art competitors in accuracy.
|
| |
| 16:35-16:55, Paper TuCT4.2 | |
| GEM-Lane:Geometry-Aware Efficient Mamba for Real-Time Lane Detection |
|
| Sun, Yangchun | Xi'an University of Posts and Telecommunications |
Keywords: Computer Vision for Transportation, Intelligent Transportation Systems
Abstract: 。车道识别在自动驾驶中起着关键作用, 然而,这需要在高速之间做出一个挑战性的权衡 推断和精确定位。准确捕捉 细长几何结构与远距离建模 依赖依然困难,尤其是在严重情况下 阻塞。为了解决这些挑战,本文 提出了一种几何感知的高效曼巴线 检测(GEM-Lane),基于混合型的实时框架 状态空间模型(SSM)。GEM-Lane 集成选择性 通过层级细化扫描以增强效果 拓扑感知。具体来说,是一种混合特征 聚合策略在计算效率方面取得了平衡 稳健的特征表示。这里是残差VMamba 模块在高层次集成以建模全局 语义依赖,而轻量级 EVMamba 模
|
| |
| 16:55-17:15, Paper TuCT4.3 | |
| Ultra-Low Power Industrial Perception: Sub-Degree Gauge Reading Via CIM-Optimized Polar Networks |
|
| Zhang, Tong | China Mobile Research Institute |
| Pan, Weiping | China Mobile Research Institute |
| Xu, Qingqing | China Mobile Research Institute |
| Li, Xiaotao | China Mobile Research Institute |
| Wang, Yaqi | China Mobile Research Institute |
| Niu, Yawen | China Mobile Research Institute |
Keywords: Computer Vision in Automation, Factory Automation, Cyber-physical Production Systems and Industry 4.0
Abstract: In industrial Internet of Things (IIoT) applications, automated analog gauge reading faces a practical conflict between high-precision recognition and stringent edge power constraints. Standard microcontrollers, restricted by the von Neumann bottleneck, are inefficient for executing vision models at milliwatt power levels. To address this problem, we propose a hardware-algorithm co-designed reading system tailored for the WTM2101 Computing-in-Memory (CIM) chip. We introduce WTM-Polar-Net, which uses physics-aware polar coordinate encoding to simplify rotational feature learning. This design enables a stacked 3×3 small-kernel architecture with dense decoding and no skip connections, keeping intermediate activations within the CIM SRAM budget while reducing over-smoothing of slender pointer features. In addition, a focus-weighted MSE loss and Model Exponential Moving Average (EMA) are used to improve heatmap localization stability under low-resolution and noisy conditions. With Quantization-Aware Training (QAT), the model is adapted for INT8 deployment on the CIM platform. Evaluated on the public RoboGauge-10k dataset from Roboflow Universe, our method achieves a Mean Absolute Error of 0.96º and a 93.5% accuracy within a 2º tolerance. Hardware measurements show 7.2 mW power consumption and 148 ms latency on WTM2101, corresponding to 373× lower energy per inference than the STM32F407 MCU under the same quantized model setting.
|
| |
| 17:15-17:35, Paper TuCT4.4 | |
| Enhancing the Real-Time Capability of Manufacturing Digital Twins through FPGA with a Model-Based Design Approach |
|
| Gu, Jingming | University of Auckland |
| Xu, Xun | University of Auckland |
| Biglari-Abhari, Morteza | University of Auckland |
| Polzer, Jan | University of Auckland |
Keywords: Cyber-physical Production Systems and Industry 4.0, Intelligent and Flexible Manufacturing, Factory Automation
Abstract: Digital Twin (DT) plays an increasingly important role in smart manufacturing by enabling real-time monitoring, analysis, and intelligent decision-making. However, existing DT implementations struggle to meet the strict latency and determinism requirements in time-sensitive industrial applications. To address this limitation, this paper proposes a Field-Programmable Gate Array (FPGA)-enabled real-time DT architecture that integrates hardware-level perception, processing, and decision execution at the edge. To facilitate the development and deployment of FPGA-enabled modules within this framework, an FPGA-oriented Model-Based Design (MBD) workflow is introduced, enabling systematic design, verification, and hardware implementation through FPGA-in-the-Loop (FIL) testing. A representative case study based on an FIR filter for sensor signal processing is presented to demonstrate the proposed workflow. Experimental results show that the FPGA implementation achieves ultra-low and strictly deterministic latency while preserving functional correctness, validating the suitability of FPGA-based computation for real-time DT applications in smart manufacturing systems.
|
| |
| 17:35-17:55, Paper TuCT4.5 | |
| Measurement-Calibrated Multi-Camera Fusion for Vision-Based Indoor Localization |
|
| Toro Diz, Mateo | Technische Hochschule Rosenheim |
| Hoss, Jonathan | Rosenheim University of Applied Sciences |
| Klarmann, Noah | Rosenheim University of Applied Sciences |
Keywords: Data fusion, Computer Vision for Transportation, Surveillance Systems
Abstract: Indoor vision-based localization systems are affected by detection noise, occlusions, and limited camera coverage, leading to uncertainty at multiple stages of the pipeline. While multi-camera data fusion is widely used to mitigate these issues, it is typically treated as a black-box component and evaluated solely end-to-end, obscuring its mechanistic contributions. To address this gap, this work investigates whether explicitly characterizing single-camera localization errors can be leveraged to calibrate and optimize multi-camera data fusion. We introduce a novel, measurement-calibrated fusion approach that integrates component-wise error quantification— specifically isolating homography calibration, human detection, and motion tracking. A component-wise evaluation is conducted to quantify error contributions from homography calibration, human detection, and motion tracking. Experimental results show that data fusion improves localization accuracy compared to single-camera baselines. While measurement-calibrated fusion provides only limited improvement in absolute accuracy over standard fusion, it substantially reduces trajectory variance and improves motion smoothness, which are critical for applications requiring stable and continuous motion estimates. These results highlight the value of explicit error characterization when designing data fusion strategies for vision-based indoor positioning systems.
|
| |
| 17:55-18:15, Paper TuCT4.6 | |
| A Calibration-Free Two-Stage Visual Servoing Framework for Industrial Bin-Picking Via Online Imitation Learning |
|
| Ying, Zixi | Zhejiang University |
| Kong, Shuhang | University of Nottingham Ningbo China |
| Kong, Xiaowu | Zhejiang University |
| Liu, Xuanyang | W-Innovation Technology (hangzhou) Co.Ltd |
Keywords: Deep Learning in Robotics and Automation, Computer Vision in Automation, Intelligent and Flexible Manufacturing
Abstract: Industrial bin-picking remains a challenging task in automated manufacturing. Existing methods often depend on accurate depth sensing and calibration, limiting their robustness and applicability. Visual servoing (VS) offers a closed-loop alternative but struggles with large initial offsets and textureless parts. In this paper, we propose a two-stage pipeline that combines object detection and learning-based visual servoing using only two monocular RGB cameras. The first stage isolates a target part using YOLOv8. The second stage employs an online imitation learning policy that maps dual wrist-mount images directly to 6-DoF end-effector action, eliminating the need for hand-eye calibration. To ensure deterministic supervision and reduce human labor, we introduce an automatic expert trajectory generator and an improved Greedy-DAgger online training algorithm. Real-world experiments show that our method achieves sub-millimeter translation and sub-degree rotation accuracy under various disturbances, with a 94.23% part clearance rate in a tight-clearance (0.8mm) placement task, demonstrating its robustness and industrial potential.
|
| |
| TuCT5 |
205 |
| Emerging Technology of Automation System 3 |
Regular Session |
| |
| 16:15-16:35, Paper TuCT5.1 | |
| Implicit Point-To-Voxel LiDAR-IMU SLAM (I) |
|
| Dong, Yan | Huazhong University of Science and Technology |
| Xu, Enci | Huazhong University of Science and Technology |
| Yang, Junyu | Huazhong University of Science and Technology |
| Qiu, Shaoqiang | Huazhong University of Science and Technology |
| Wang, Jiong | Huazhong University of Science and Technology |
| Han, Bin | Huazhong University of Science and Technology |
Keywords: Autonomous Vehicle Navigation, Sensor Fusion, Deep Learning in Robotics and Automation
Abstract: Existing LiDAR SLAM methods typically rely on parameterizable explicit local models -- planes, quadric surfaces, or Gaussian distributions -- to formulate geometric constraints for state estimation. In unstructured environments, however, structures like tree trunks, weeds and uneven terrains are difficult to be precisely represented by simple parametric models, leading to fewer reliable constraints and introducing modeling errors. To address this issue, we propose a novel implicit point-to-voxel observation model that directly predicts residuals and their uncertainty from local voxel geometries without explicit fitting. The resulting measurements are incorporated into an iterated error-state Kalman filter (IESKF) for pose update. To support efficient online retrieval and update of local geometric context, we further introduce an implicit voxel map representation that partitions the map into voxels and maintains an online-updated implicit representation for each voxel. Experiments show that the proposed method provides more stable effective constraints and improves localization accuracy in unstructured scenes, while maintaining real-time performance on a CPU platform.
|
| |
| 16:35-16:55, Paper TuCT5.2 | |
| Analysis of Calibration Poses and Algorithms for Kinematic Calibration of Collaborative Robots Using Laser Tracker Measurements |
|
| Stoeckl, Florian | DHBW Karlsruhe |
| Müller, Silvan | Baden-Württemberg Cooperative State University |
| Rettig, Oliver | DHBW Karlsruhe |
| Strand, Marcus | Baden-Wuerttemberg Cooperative State University Karlsruhe |
| Gardill, Markus | Brandenburg University of Technology Cottbus-Senftenberg |
Keywords: Calibration and Identification, Collaborative Robots in Manufacturing, Factory Automation
Abstract: Kinematic calibration of collaborative robotic arms is essential for achieving high absolute positioning accuracy in industrial applications. As Ma et al. showed, the quality of the calibration critically depends on the number and spatial distribution of the measurement poses. The choice of processing algorithm also influences the results. This work-in-progress paper investigates how different calibration poses and their algorithmic interpretation influence the accuracy achieved for a Universal Robots UR5e collaborative arm, measured with a FARO Vantage E laser tracker. So far we have analysed the influence of transformation matrices on the calibration results and present these findings in this paper. To calculate the transformation matrix, we compare the Kabsch algorithm with circle-fitting-based calculations, using both automatic and manual measurements. The pose distribution and amount are analysed for calibration based on a rigid analysis of the robot’s positioning accuracy in its workspace. Preliminary experimental results indicate that automatically creating the transformation matrix based on distributed poses and processing it with the Kabsch algorithm yields maximum absolute errors of 0.181 mm, 0.257 mm, and 0.146 mm in X, Y, and Z, respectively, whereas a circle around the base analysed with a circle-fitting algorithm yields maximum absolute errors of 0.296 mm, 0.332 mm, and 0.800 mm in X, Y, and Z, respectively. This is compared to manual measurements, which achieve maximum absolute errors of 0.174 mm, 0.286 mm, and 0.149 mm in X, Y, and Z, respectively. A large measurement of 10 000 poses is still outstanding for the recommendation of suitable calibration poses.
|
| |
| 16:55-17:15, Paper TuCT5.3 | |
| Camera-Projector-Robot Calibration for Human-Robot Collaboration |
|
| Angleraud, Alexandre | Tampere University |
| Houbre, Quentin | Tampere University |
| Siltala, Niko | Tampere University |
| Pieters, Roel S. | Tampere University |
Keywords: Calibration and Identification, Collaborative Robots in Manufacturing, Human-Centered Automation
Abstract: Effective and safe human-robot collaboration relies on the clear communication of information relevant to the shared task. Contextual information, such as work areas, object placement locations and safety zones can be projected on the work space, observable by the human operator, in an unobtrusive way. To facilitate this, a projector-camera system needs to be installed alongside a robotic arm as perception and communication tool. As the camera, projector and robot all operate in their own work space, calibration is needed to align the workspaces and compensate for distortion effects. This paper describes the calibration processes and all required tools for such system, including calibration between projector and camera, robot end-effector and camera, and camera intrinsics. As outcome the calibration should find the relationship between projected points and the world coordinate frame. The approach is evaluated with two workcells; a collaborative robot workcell and an industrial robot workcell, each with different configurations for the projector-camera system. The results demonstrate the procedures for all three calibration methods in detail, showing that an accurate calibration can be obtained for both workcells. The software is available at https://github.com/cogrob-tuni/projector-interface.
|
| |
| 17:15-17:35, Paper TuCT5.4 | |
| Robot Exploration under Limited and Occluded Field-Of-View: A Direction-Aware Deep Reinforcement Learning Approach Via Node-Sector Modeling |
|
| Ye, Yiru | National Taiwan University (NTU) |
| Lian, Feng-Li | National Taiwan University |
| Chien, Shaoyu | Delta Electrics Inc |
| Huang, Pei-Chun | National Taiwan University |
|
|
| |
| 17:35-17:55, Paper TuCT5.5 | |
| Speech-To-Trajectory: An LLM-Enabled Human-Robot Collaborative Framework for Generating New Assembly Tasks from Verbal Guidance |
|
| Zhao, Jiabao | The Pennsylvania State University |
| Li, Hongliang | The Pennsylvania State University |
| Lim, Jonghan | The Pennsylvania State University |
| Kovalenko, Ilya | The Pennsylvania State University |
Keywords: Collaborative Robots in Manufacturing, Assembly, Industrial and Service Robotics
Abstract: Human-Robot Collaboration has enhanced manufacturing efficiency and productivity by integrating the precision and repeatability of robots with the adaptability of human operators. The growing demand for customized products frequently requires human operators to manually reprogram robot task plans. This process demands specialized skills such as coding and robotic programming, making it costly and inaccessible to non-expert human operators. In this study, we propose a Large Language Model-enabled Human-Robot Collaboration framework that facilitates human-robot interaction through natural language. Using the proposed framework, a human operator can teach the robot to assemble new products via verbal guidance, while the robot learns and stores the manipulation sequences for future reuse. In addition, a motion trajectory simplifier is integrated into the framework to refine assembly motion for reuse, further improving manufacturing efficiency. The proposed framework is evaluated on the NIST Assembly Task Board, demonstrating a promising direction for Human-Robot Collaboration in manufacturing.
|
| |
| 17:55-18:15, Paper TuCT5.6 | |
| Contact-Aware 3D Hand Trajectory Reconstruction from Monocular RGB Video for Robot Manipulation |
|
| Huang, Jinyi | The University of Auckland |
| Zheng, Hao | New York University Abu Dhabi |
| Wan, Wenqing | Xi'an Jiaotong University |
| Zhu, Xu | Dalian University of Technology |
| Polzer, Jan | The University of Auckland |
| Xu, Xun | University of Auckland |
Keywords: Collaborative Robots in Manufacturing, Computer Vision in Automation, Learning and Adaptive Systems
Abstract: Enabling collaborative robots to learn manipulation trajectories from human demonstrations is essential for flexible automation in human-centric manufacturing. While video-based learning offers a scalable approach, its effectiveness is often constrained by monocular depth ambiguity and temporal tracking instability in hand trajectory reconstruction, as well as the geometric inconsistency between hand trajectories and robot task-space requirements. To address these challenges, this paper proposes a contact-aware 3D hand trajectory reconstruction framework that generates temporally coherent manipulation trajectories from monocular RGB videos for robot execution. It detects hand-object contact phases and identifies synchronized motion segments within them, exploiting these as geometric anchors to propagate depth corrections and suppress trajectory jitter. A geometry-aware mapping module then translates the reconstructed trajectories into executable robot end-effector waypoint sequences through task-space coordinate transformation. Experiments conducted on the HA-ViD and Assembly101 datasets across multiple viewpoints demonstrate improved trajectory fidelity to object motion, depth consistency, and motion smoothness. These results confirm that contact-aware geometric constraints can effectively bridge the gap between noisy monocular observations and reliable robot execution, offering a practical pathway for robot motion acquisition from human demonstration videos.
|
| |
| TuCT6 |
503 |
LLM-Based Application for Manufacture-Circulation Industrial System (MCIS)
3 |
Regular Session |
| |
| 16:15-16:35, Paper TuCT6.1 | |
| Towards Applying Automation to Retrofitting Industrial Variable Frequency Drives |
|
| Shojaeifard, Leyla | University of Helsinki |
| Kilamo, Terhi | Tampere University |
| Systä, Kari | Tampere University |
| Männistö, Tomi | University of Helsinki |
Keywords: Power and Energy Systems automation, Manufacturing, Maintenance and Supply Chains, Product Design, Development and Prototyping
Abstract: Variable Frequency Drives (VFDs) can produce cost savings by improving energy efficiency and reducing motor wear and tear, thereby lowering maintenance costs. With proper maintenance and modernization, the drives’ lifespans can extend to decades, but they will eventually need to be replaced. This case study explores the challenges of replacing VFDs in existing industrial installations when backward compatibility is not possible. We interviewed 22 employees of a large, internationally operating VFD manufacturer about retrofitting and its challenges. Our findings include a four-phase process for retrofitting and considerations on how the application domain of the VFD use case can affect the retrofitting process.
|
| |
| 16:35-16:55, Paper TuCT6.2 | |
| Voltage-Prioritized Dynamic Current Allocation for Transient Synchronization of Hybrid GFM-GFL Systems |
|
| Lou, Xiyue | Economic and Technical Research Institute of State Grid Jilin Electric Power Company |
| Jiang, Minglei | Economic and Technical Research Institute of State Grid Jilin Electric Power Company |
| Xin, She | Economic and Technical Research Institute of State Grid Jilin Electric Power Company |
| Feng, Fan | Economic and Technical Research Institute of State Grid Jilin Electric Power Company |
| Zheng, Huicong | Economic and Technical Research Institute of State Grid Jilin Electric Power Company |
| Wang, Kaiping | Economic and Technical Research Institute of State Grid Jilin Electric Power Company |
| Shi, Shengyao | Economic and Technical Research Institute of State Grid Jilin Electric Power Company |
Keywords: Power and Energy Systems automation, Smart Grids, Robust/Adaptive Control
Abstract: Hybrid microgrids comprising Grid-Forming (GFM) and Grid-Following (GFL) converters face severe transient synchronization challenges under weak-grid or voltage-stressed conditions, where fault-induced voltage depression and converter current saturation reduce the effective Point of Common Coupling (PCC) voltage stiffness. During severe faults, conventional Circular Current Limiting (CCL) proportionally scales down both active and reactive current components, which may weaken reactive voltage support during the critical post-fault recovery interval and aggravate power-frequency oscillations. To address this issue, this paper proposes an adaptive voltage-prioritized dynamic current allocation strategy. By monitoring the PCC voltage magnitude, the proposed allocator introduces a bounded q-axis voltage-support reservation during deep voltage sags and dynamically refills the residual converter current capacity between the d- and q-axes. Different from fixed reactive-priority current limiting, the proposed method does not maximize reactive current at all times; instead, the dynamic refilling mechanism balances voltage support and active-current headroom within the same current limit. An Equal-Area-Criterion-inspired analysis is used to interpret how PCC voltage recovery affects the post-fault synchronization margin. Time-domain simulations, including ablation and reservation-coefficient sensitivity tests, verify that the proposed method improves post-fault frequency recovery and suppresses large power-frequency oscillations without noticeably increasing converter current stress.
|
| |
| 16:55-17:15, Paper TuCT6.3 | |
| AED: Automatic Discovery of Effective and Diverse Vulnerabilities for Autonomous Driving Policy with Large Language Models |
|
| Qiu, Le | Tsinghua University |
| Xu, Zelai | Tsinghua University |
| Tan, Qixin | Tsinghua University |
| Tang, Wenhao | Tsinghua University |
| Yu, Chao | Tsinghua University |
| Wang, Yu | Tsinghua University |
Keywords: Reinforcement, Intelligent Transportation Systems, AI-Based Methods
Abstract: Assessing the safety of autonomous driving policy is of great importance, and reinforcement learning (RL) has emerged as a powerful method for discovering critical vulnerabilities in driving policies. However, existing RL-based approaches often struggle to identify vulnerabilities that are both effective—meaning the autonomous vehicle is genuinely responsible for the accidents—and diverse—meaning they span various failure types. To address these challenges, we propose AED, a framework that uses large language models (LLMs) to Automatically discover Effective and Diverse vulnerabilities in autonomous driving policies. We first utilize an LLM to automatically design reward functions for RL training. Then we let the LLM consider a diverse set of accident types and train adversarial policies for different accident types in parallel. Finally, we use preference-based learning to filter ineffective accidents and enhance the effectiveness of each vulnerability. Experiments across multiple simulated traffic scenarios and tested policies show that AED uncovers a broader range of vulnerabilities and achieves higher attack success rates compared with expert-designed rewards, thereby reducing the need for manual reward engineering and improving the diversity and effectiveness of vulnerability discovery.
|
| |
| 17:15-17:35, Paper TuCT6.4 | |
| Robotic Agentic Platform for Intelligent Electric Vehicle Disassembly |
|
| Allen, Zachary | University of Colorado Boulder |
| Conway, Max | University of Colorado Boulder |
| Antieau, Lyle | University of Colorado Boulder |
| Augustin Ponraj, Allen Devaraj | University of Colorado Boulder |
| Correll, Nikolaus | University of Colorado at Boulder |
Keywords: Remanufacturing, Plug-in Electric Vehicles, AI-Based Methods
Abstract: Electric vehicles create an urgent need for scalable battery recycling, yet disassembly of large EV battery packs remains largely manual due to high voltages and significant design variability. We present a human–robot collaborative research platform for EV battery disassembly, designed to investigate perception-driven manipulation, flexible automation, and AI-assisted robot programming in realistic recycling scenarios. The system integrates a gantry-mounted industrial manipulator, RGB-D perception, and an automated nut-running tool for fastener removal on a full-scale EV battery pack. An open-vocabulary object detection pipeline achieves 0.918 mAP50, enabling reliable identification of screws, nuts, busbars, and other components. We experimentally evaluate (n=204) three fastener removal strategies, achieving one-shot success rates of 83% using taught poses, 57% with one-shot vision execution, and 82% with visual servoing, with average disassembly times of 22.9–36.3 minutes for the battery's top cover screws. To support flexible interaction, we introduce agentic AI specifications for robotic disassembly tasks, allowing LLM agents to translate high-level instructions into robot actions through structured tool interfaces and ROS services. SmolAgents using GPT-4o-mini and Qwen 3.5 9B, the latter running fully offline on edge hardware. Tool-based interfaces achieve 100% task completion, while automatic ROS service discovery shows 43.3% failure rates, highlighting the need for structured robot APIs for reliable LLM-driven control. This platform enables systematic investigation of human–robot collaboration, agentic robot programming, and increasingly autonomous disassembly workflows, providing a practical foundation for research toward scalable robotic battery recycling.
|
| |
| 17:35-17:55, Paper TuCT6.5 | |
| Automated Cleaning System for Photovoltaic Panels in Desert Environments |
|
| Hefny, Anan | Rochester Institute of Technology |
| Rongo, Zachary | Rochester Institute of Technology |
| Gittelsohn, Daniel | Rochester Institute of Technology |
| Bork, Noah | Rochester Institute of Technology |
| Jackson, Alec | Rochester Institute of Technology |
| Rodgers, Tyler | Rochester Institute of Technology |
| Yan, Bing | Rochester Institute of Technology |
Keywords: Renewable Energy Sources, Sustainability and Green Automation, Energy and Environment-aware Automation
Abstract: Dust accumulation significantly reduces the power output of photovoltaic (PV) panels. Commonly used water sprayers consume large amounts of water annually for cleaning PV panels which poses environmental and economic concerns, especially in desert environments where water is scarce. In this work, an autonomous PV panel cleaning system with a novel threshold-based control strategy is developed to maximize energy savings through optimized cleaning operations. The threshold is selected based on an energy savings analysis to ensure economic feasibility. The effect of dust density on PV power is investigated experimentally, and a method for quantifying the dust accumulated is developed. The device is successful at autonomous operation without manual intervention and achieves a cleaning rate of 97%, with no degradation after consecutive cleaning cycles. The proposed system can be adapted to larger arrays and integrated with Internet of Things (IoT) based monitoring systems, offering a practical solution for utility-scale PV installations in arid regions.
|
| |
| 17:55-18:15, Paper TuCT6.6 | |
| Optimization-Based Exploration of Steelmaking Decarbonization Routes |
|
| Haikarainen, Carl | Åbo Akademi University |
| Saxen, Henrik | Abo Akademi University |
Keywords: Sustainability and Green Automation, Intelligent and Flexible Manufacturing, Sustainable Production and Service Automation
Abstract: The steelmaking industry is under heavy decarbonization pressure, but the industrial sector is capital intensive and fossil-free alternative processes are neither fully developed nor economically feasible. Direct reduction of iron ore using hydrogen in shaft furnaces is a promising decarbonization alternative, but adoption of the technology is expected to take several decades, during which new and old processes will operate side by side. This is a challenge for efficient material and energy integration in the steel plant. This paper presents an approach where the transition toward sustainable operation is formulated as an optimization problem, with accumulated costs over a project period being minimized, and with discrete break points where new unit processes can be introduced and old units can be decommissioned. The resulting mixed-integer quadratically constrained problem can be solved under different boundary conditions to study the economic feasibility of transition paths. The approach is illustrated by example cases featuring a steel plant with initially two blast furnaces and a basic oxygen furnace where the carbon intensity is decreased by a gradual adoption of direct reduction technology combined with electric arc furnaces. The arising solutions are presented and discussed at some length, and some suggestions for future work are also given.
|
| |
| TuCT7 |
504 |
Intelligent Solutions for PCB & Semiconductor Smart Manufacturing and
Manufacturing Excellence |
Special Session |
| Organizer: Chien, Chen-Fu | National Tsing Hua University |
| Organizer: Ehm, Hans | Infineon Technologies AG |
| Organizer: Kuo, Hsuan-An | National Cheng Kung University |
| Organizer: Wang, Hung-Kai | National Tsing Hua University |
| Organizer: Fu, Wenhan | University of Shanghai for Science and Technology |
| Organizer: Huang, Yu-Chieh | Zhen Ding Tech. Group (ZDT) and Leading Tech. (LIST) |
| |
| 16:15-16:35, Paper TuCT7.1 | |
| Two-Phase Rolling Copper Price Forecasting for Min–Max Regret Procurement Decision and an Empirical Study for PCB Smart Production (I) |
|
| Chen, Ze-Xi | National Tsinghua University |
| Chien, Chen-Fu | National Tsing Hua University |
Keywords: Manufacturing, Maintenance and Supply Chains, Big-Data and Data Mining, AI-Based Methods
Abstract: Abstract— Printed circuit board (PCB) industry is a core contributor to artificial intelligence (AI) infrastructure, and digital transformation has become increasingly critical for maintaining competitive capability. In general, copper-related raw materials account for around 30–40% of the cost structure in the PCB manufacturing. However, the price of copper fluctuates due to demand–supply mismatches as well as international economic and political factors. This study develops a two–phase data–driven decision framework to predict weekly copper prices and support procurement decision-making. In the first phase, a Bi–LSTM–CNN–CBAM model was employed to predict weekly copper prices and to augment the initial output with residuals using a lightweight multilayer perceptron (Light MLP). The second phase translated forecast residuals into 81 decision scenarios. A Min–Max Regret (MMR) criterion considers each possible price outcome generated from the forecasting scenarios to derive procurement strategies that minimize the maximum regret relative to ex–post optimal decisions. The data-driven decision framework was validated via an empirical analysis of a world-leading PCB company. This study demonstrates that the two-phase framework provides a robust procurement decision to reduce supply chain risk under price volatility, thereby empowering digital transformation.
|
| |
| 16:35-16:55, Paper TuCT7.2 | |
| On-Line Hierarchical RF Signal Anomaly Inspection and Localization on Production Lines (I) |
|
| Kuo, Chen-Yi | National Tsing Hua University |
| Huang, Chang-Kai | Zhen Ding Tech. Group Technology (ZDT) |
| Lai, Jingyuan | Avary Holding |
| Hsieh, Hsin-Lung | Zhen Ding Tech. Group Technology (ZDT) |
| Chien, Chen-Fu | National Tsing Hua University |
Keywords: Probability and Statistical Methods, Failure Detection and Recovery, Intelligent and Flexible Manufacturing
Abstract: Identifying diverse abnormal events in electronics fabrication requires efficient monitoring of high-dimensional spectral profiles across intricate antenna networks. To alleviate the intensive manual diagnostic burden and complex verification complexity in on-line environments, this study proposes a hierarchical framework for on-line RF signal anomaly inspection and localization. The methodology integrates an ensemble of linear and non-linear learners, incorporating PCA, ICA, and their kernel-based extensions, to construct a robust global screening model. Instead of exhaustive manual analysis, a multiscale partitioning strategy is employed to provide normalized anomaly scoring for localized spectral refinement. An empirical study conducted at a leading PCB manufacturing company validates that the proposed hierarchical logic effectively filters nominal units while pinpointing suspicious frequency regions base on a majority voting mechanism. The results demonstrate that the framework enhances diagnostic resolution and operational throughput, offering a scalable decision-support solution for maintaining communication reliability in high-volume manufacturing environments.
|
| |
| 16:55-17:15, Paper TuCT7.3 | |
| Predictive AGV Trigger Time to Minimize Material Waiting Time in PCB Manufacturing (I) |
|
| Ku, Chien-Chun | Zhen Ding Tech. Group |
| Hsieh, Hsin-Lung | Zhen Ding Tech. Group Technology (ZDT) |
| Huang, Yu-Chieh | Zhen Ding Tech. Group (ZDT) and Leading Tech. (LIST) |
| Chien, Chen-Fu | National Tsing Hua University |
Keywords: Manufacturing, Maintenance and Supply Chains, Robust Manufacturing, Agent-Based Systems
Abstract: The rapid transition toward Industry 4.0 demands a paradigm shift in Printed Circuit Board (PCB) manufacturing, where the operational efficiency of Automated Guided Vehicle (AGV) systems dictates factory throughput. Current industrial practices often struggle with dynamic material flows and fixed fleet sizes, leading to intense resource contention. Traditional dispatching methods heavily depend on spatial proximity-based heuristics, such as the Nearest-Vehicle (NV) rule. These reactive approaches evaluate only instantaneous system states, inherently lacking the look-ahead capability required to prevent premature resource allocation and minimize severe delays for critical lots. To address these operational bottlenecks, this research proposes a proactive predictive scheduling framework that integrates standard dispatching algorithms with Real-time Equipment Data Analytics from the Equipment Automation Program (EAP). By dynamically parsing equipment state variables and event logs, the framework accurately forecasts process completion times. This allows the system to synchronize AGV movements with imminent machine completions rather than reacting passively to immediate requests. Empirical evaluations demonstrate that the proposed EAP-integrated approach significantly reduces the average material waiting time and total makespan, optimizing finite AGV utilization in highly dynamic manufacturing environments.
|
| |
| 17:15-17:35, Paper TuCT7.4 | |
| Impedance Prediction Based on Analytical Formula and TPE-MLP for Signal Integrity to Enhance PCB Reliability Assurance (I) |
|
| Fu, Wenhan | University of Shanghai for Science and Technology |
| Wang, Qiuyang | University of Shanghai for Science and Technology |
| Chien, Chen-Fu | National Tsing Hua University |
Keywords: Semiconductor Manufacturing, Intelligent and Flexible Manufacturing, Cyber-physical Production Systems and Industry 4.0
Abstract: With the rapid advancement of high-frequency, high-speed, and high-density electronic systems, signal integrity has become increasingly critical in printed circuit boards (PCB). Characteristic impedance plays a key role in maintaining signal transmission quality, making accurate impedance prediction essential during PCB design and reliability assurance. Conventional approaches either rely on empirical analytical formulas with limited accuracy or computationally expensive electromagnetic simulations. To address these limitations, this paper presents a hybrid impedance prediction framework that integrates analytical modeling with machine learning, where a multilayer perceptron (MLP) is developed for impedance prediction and its hyperparameters are automatically optimized using the Tree-structured Parzen Estimator (TPE). The experimental dataset is generated via Latin hypercube sampling based on the modified Hammerstad–Jensen formula and industrial PCB process specifications. The results demonstrate that the proposed method achieves high prediction accuracy, with an R² of 0.9996 and an MSE of 0.5213, outperforming several baseline models. The proposed framework provides an efficient and reliable approach for impedance prediction, supporting improved signal integrity analysis and PCB design optimization.
|
| |
| 17:35-17:55, Paper TuCT7.5 | |
| A Large Language Model-Augmented Semantic Framework for Knowledge Management: Facilitating Pricing Strategies in the More-Than-Moore Supply Chain (I) |
|
| Sun, Pei-Ching | Zhen Ding Technology Holding Limited |
| Ehm, Hans | Infineon Technologies AG |
| Kuo, Hsuan-An | National Cheng Kung University |
| Bonik, Marta | Infineon Technologies AG |
| Chien, Chen-Fu | National Tsing Hua University |
Keywords: Manufacturing, Maintenance and Supply Chains, AI-Based Methods, Planning, Scheduling and Coordination
Abstract: The More-than-Moore supply chain is characterized by heterogeneous technologies, shortened product life cycles, and capacity-constrained operations, making pricing decisions increasingly dependent on lead-time differentiation, customer classification, technology maturity, and capacity allocation. Existing pricing approaches primarily focus on revenue optimization but lack a formal mechanism for representing and validating the semantic relationships among these interdependent decision factors. To address this gap, this study proposes a UNISON-based semantic framework augmented by a large language model (LLM) for knowledge management facilitating lead-time sensitive pricing strategies. The proposed framework formalizes pricing-related entities and influence relationships through ontology-based modeling, supports logical consistency verification using SPARQL-based semantic querying, and incorporates LLM-assisted reasoning for policy plausibility assessment and scenario validation. A scenario-based validation study inspired by semiconductor practices in Germany and Taiwan is conducted to evaluate the proposed framework. The results show that lead-time-based pricing can be consistently represented and interpreted within a broader pricing ontology while supporting explainable reasoning and semantic consistency across pricing, contract governance, and capacity management. This study contributes a structured approach for integrating pricing governance, capacity management, and knowledge management in advanced semiconductor supply chains.
|
| |
| 17:55-18:15, Paper TuCT7.6 | |
| A Large Language Model Approach for Intelligent Reasoning in Shipbuilding Processes with Integrated Data Privacy Protection (I) |
|
| Guo, Mengyuan | Guangdong University of Technology |
| Chen, Chong | Guangdong University of Technology |
| Wang, Tao | Guangdong University of Technology |
| Cheng, Lianglun | Guangdong University of Technology |
Keywords: Intelligent and Flexible Manufacturing, Hybrid Logical/Dynamical Planning and Verification, Cognitive Automation
Abstract: Large language models (LLMs) are driving shipbuilding process knowledge services from retrieval-based question answering toward regulation-driven reasoning and decision-making. However, process specifications, operation records, and quality inspection texts usually contain personally identifiable information (PII), such as persons, organizations, locations, and dates. These entities are often embedded in evidence chains through cross-sentence coreference and multi-field associations, making explicit intermediate outputs prone to privacy residuals and leakage during generation. To address this issue, this work studies privacy-aware structured reasoning for shipbuilding process decision workflows under explicit privacy-control requirements. The proposed approach builds a stage-wise reasoning chain with a Translator, a Planner, a Solver, and a Verifier, while introducing an input-side Anonymizer and a Translator-side Privacy Guard for unified privacy processing and chain-level auditing. Specifically, the input side performs sample-level consistent anonymization and PII detection; the Translator output is further subjected to residual auditing and local repair; and key intermediate products as well as final outputs are uniformly audited using PII residual rate. This work introduces privacy auditing and residual control into multiple stages of structured reasoning workflows to improve privacy controllability during multi-step reasoning. Experiments on ProntoQA, ProofWriter, FOLIO, and the Ship dataset show that the reasoning workflow preserves the accuracy and usability of multi-step constraint reasoning while satisfying engineering privacy-control requirements.
|
| |
| TuCT8 |
501 |
| Virtual Session 3 |
Regular Session |
| |
| 16:15-16:35, Paper TuCT8.1 | |
| Simulation-Based Planning of Motion Sequences for Automated Procedure Optimization in Multi-Robot Assembly Cells |
|
| Schneider, Loris | Karlsruhe Institute of Technology |
| Ungen, Marc | Robert Bosch GmbH |
| Huber, Elias | Robert Bosch GmbH |
| Klein, Jan-Felix | KTH Royal Institute of Technology |
Keywords: Planning, Scheduling and Coordination, Motion and Path Planning, Intelligent and Flexible Manufacturing
Abstract: Reconfigurable multi-robot cells offer a promising approach to meet fluctuating assembly demands. However, the recurrent planning of their configurations introduces new challenges, particularly in generating optimized, coordinated multi-robot motion sequences that minimize the assembly duration. This work presents a simulation-based method for generating such optimized sequences. The approach separates assembly steps into task-related core operations and connecting traverse operations. While core operations are constrained and predetermined, traverse operations offer significant potential for optimization. Therefore, the scheduling of core operations is formulated as an optimization problem, while feasible traverse operations are incorporated through a decomposition-based motion planning strategy. Several solution techniques are explored, including a sampling heuristic, tree-based search and gradient-free optimization. For motion planning, a decomposition method is proposed that identifies specific subproblems in the schedule, which can be solved independently with modified centralized path planning algorithms. The proposed method generates collision-free multi-robot assembly procedures that outperform a baseline relying on decentralized, robot-individual motion planning. Its effectiveness is demonstrated through simulation experiments.
|
| |
| 16:35-16:55, Paper TuCT8.2 | |
| Explainable Geospatial AI for Satellite Ground Station Siting Using LiDAR-Derived Terrain Intelligence |
|
| Sarkar, Shohini | University of Maryland, College Park |
| Mahendran, Smithi | University of Maryland, College Park |
| Chudasama, Rishi | University of Maryland, College Park |
| Mannam, Varun | University of Maryland, College Park |
| Luthra, Arav | University of Maryland, College Park |
| Rekhi, Yuvraj | University of Maryland, College Park |
| Nadig, Vivek | University of Maryland, College Park |
| Goenka, Arsh | University of Maryland, College Park |
Keywords: Big-Data and Data Mining, Data fusion, Probability and Statistical Methods
Abstract: Representative clutter height (RCH) is a key parameter in radio propagation and interference analysis because it captures the dominant height of local obstructions that drive terminal clutter loss. Current practice often relies on fixed clutter heights assigned to land-use classes in Recommendation ITU-R P.452-18, but this misses within class variation and can lead to conservative exclusion zones and poor site ranking for low Earth orbit ground-station siting and spectrum coordination. We present an interpretable, globally deployable machine learning framework for predicting RCH from open geospatial data. The model is trained using LiDAR-derived labels from the U.S. Geological Survey 3D Elevation Program and inference-time features from global land-cover, terrain, demographic, thermal, and optical remote sensing products. We define RCH using a robust 75th-percentile clutter height statistic, evaluate multiple regressors, and select LightGBM for its accuracy, efficiency, and compatibility with feature attribution analysis. The final model achieves a mean absolute error of 1.79m and an R^2=0.765, reducing absolute error by more than 60% relative to the ITU baseline. Beyond aggregate fit, we evaluate domain facing criteria relevant to RF planning, including meter scale error, tolerance-band accuracy, over and under estimation tails, agreement with ITU clutter height regimes, and SHAP-based physical plausibility. SHAP identifies tree canopy cover, land-cover semantics, and spectral reflectance as the most influential predictors. Studies on segmentation derived features, non-forest ablations, and land-cover matched international validation show that open geospatial data can improve clutter modeling at scale without sacrificing interpretability or deployability.
|
| |
| 16:55-17:15, Paper TuCT8.3 | |
| Deep Reinforcement Learning for Flexible Job Shop Scheduling with Random Job Arrivals |
|
| Tang, Yu | ETH Zurich |
| Zakwan, Muhammad | ETH Zurich |
| Balta, Efe | Inspire AG |
| Lygeros, John | ETH Zurich |
| Rupenyan, Alisa | Zurich University of Applied Sciences |
Keywords: Cyber-physical Production Systems and Industry 4.0, Intelligent and Flexible Manufacturing, Planning, Scheduling and Coordination
Abstract: The Flexible Job Shop Scheduling Problem (FJSP) is the optimal allocation of a set of jobs to machines. Two primary challenges persist in FJSP: the unpredictable arrival of future jobs and the combinatorial complexity of the problem, rendering it intractable for conventional mixed-integer linear programming solvers. This paper proposes an event-based Deep Reinforcement Learning (DRL) approach to solve FJSP with random job arrivals. Specifically, we employ the Proximal Policy Optimization algorithm and use lightweight Multi-Layer Perceptrons to train the DRL agent for minimizing the total completion time of all jobs. We design the state representation to be directly accessible from the environment, and limit the learning agent to selecting from among a set of well-established dispatching rules. Simulations show that our DRL approach outperforms any of the individual dispatching rules on datasets with varying heterogeneity and job arrival rates. We benchmark our DRL against an arrival-triggered mixed integer linear programming solution and show that our method achieves good performance, especially when the datasets are heterogeneous.
|
| |
| 17:15-17:35, Paper TuCT8.4 | |
| AerialGuard: Edge-Deployed Zero-Shot Anomaly Detection for UAVs |
|
| Rathore, Yash Pratap Singh | Indian Institute of Technology Mandi |
| Singh, Shruti | Indian Institute of Technology Mandi |
| Vatsi, Ritali | Indian Institute of Technology Mandi |
| Shukla, Amit | Indian Institute of Technology Mandi |
| Goyal, Pawan | IIT Kharagpur |
Keywords: Deep Learning in Robotics and Automation, Computer Vision in Automation, Surveillance Systems
Abstract: Autonomous UAV patrol requires real-time perception capable of identifying novel visual anomalies without prior labeled examples. Closed set detectors fail on unseen objects, while CLIP-based anomaly methods designed for ground-level imagery have not been validated for the aerial domain. We present AerialGuard, an end-to-end zero-shot visual anomaly detection system deployed entirely on a real autonomous UAV. AerialGuard combines a domain-aligned visual encoder finetuned on aerial imagery via contrastive learning, a Mahalanobis one-class scorer trained solely on normal examples, and a deterministic behavior router that maps anomaly confidence directly to closed-loop flight actions. A key finding is that contrastive finetuning creates a polarity inversion in embedding space, normals drift farther from the centroid than anomalies, causing Euclidean scoring to fail, which Mahalanobis whitening corrects. The scorer achieves 0.9968 AUROC on HERIDAL, a geographically disjoint alpine benchmark, with no anomaly training examples, a single threshold fitted on normal VisDrone crops transfers without modification across all evaluation conditions. The full pipeline runs at 6.4 FPS on a Jetson Orin Nano during autonomous flight. A five-flight validation at three altitudes confirms zero false state transitions and correct anomaly triggering, demonstrating that zero-shot visual scoring can drive reliable closed-loop UAV behavior under edge resource constraints.
|
| |
| 17:35-17:55, Paper TuCT8.5 | |
| Intent-State Transformer for Stable Vessel Intent Recognition in Autonomous Maritime Systems |
|
| Meepaganithage, Ayesh | University of Nevada Reno |
| Nicolescu, Mircea | University of Nevada, Reno |
| Nicolescu, Monica | University of Nevada, Reno |
Keywords: Deep Learning in Robotics and Automation, Autonomous Agents, Machine learning
Abstract: Reliable intent recognition is essential for autonomous maritime systems operating in safety-critical environments. Although recent deep learning models achieve high accuracy in sliding-window vessel behavior classification, predictions often fluctuate across adjacent time steps, leading to unstable intent estimates. To address this limitation, we propose the Intent-State Transformer with Sequential Decision Head (IST-SD), a structured sequence model that explicitly represents the evolution of behavioral intent over time. The architecture combines a Transformer encoder for contextual sequence representation with a latent intent state that is updated sequentially through a distance-conditioned gating mechanism. This design enables the model to accumulate behavioral evidence across the observation window while maintaining stable intent predictions. We evaluate the proposed approach across multiple deep learning baselines and feature-selection strategies using both full and balanced maritime encounter datasets. Experimental results show that IST-SD consistently outperforms conventional RNN and Transformer models across multiple distance regimes, produces more temporally stable predictions across successive sliding windows, and achieves the highest overall classification performance. The findings demonstrate that explicitly modeling the evolution of intent can significantly improve the robustness and reliability of maritime behavior prediction systems.
|
| |
| 17:55-18:15, Paper TuCT8.6 | |
| TwinQC: Digital Twin-Driven Contrastive Learning for Zero-Shot Surface Defect Detection in High-Mix Electronics Assembly |
|
| Elsayed, Saher | University of Pennsylvania |
| Ali, Mohamed | Texas A&M University |
| Abubaker, Samer | State University of New York at Binghamton |
| Abdelgawad, Mohamed | Kutztown University of Pennsylvania |
| Alhajj, Rayane | Virginia Tech |
Keywords: Machine learning, Computer Vision in Automation, AI-Based Methods
Abstract: High-mix electronics assembly lines introduce new printed circuit board (PCB) variants every 2-4 weeks, yet surface-mount defect detection systems require hundreds of labeled anomaly images per variant to reach production-grade accuracy, a labeling bottleneck that costs 180-320K annually per line in engineering labor and delayed yield ramp. We present TwinQC, a digital twin-driven contrastive learning framework that achieves zero-shot defect detection on new PCB variants with no labeled defect images by training exclusively on photorealistic renders from an automatically parameterized digital twin. Three co-designed contributions enable this: (1) a Twin Render Engine (TRE) that generates unlimited variant-specific normal and defect images from CAD/BOM data via domain-randomized physically based rendering in under 3 min per variant; (2) a Hierarchical Contrastive Defect Encoder (HCDE) trained on twin renders that learns a compact 256-D anomaly-sensitive embedding via multi-scale patch contrast, achieving strong real-domain transfer; and (3) an Online Domain Alignment (ODA) module that adapts the encoder to each production line's camera and lighting profile using only 20 unlabeled real images, with no defect labels required, closing the sim-to-real gap in under 8 min. Evaluated on the PCB HMix2025 benchmark, including 18 PCB variants and 4,312 annotated real defect images across 9 defect classes, and in a live SMT line pilot at a Tier-1 contract manufacturer, TwinQC achieves 96.4% AUROC on zero-shot new variants, a 12.1 percentage-point gain over the next-best label-free baseline, with 35 ms inference latency on an NVIDIA Jetson AGX Orin, satisfying the 50 ms takt-time constraint of the assembly line. Dataset, twin templates, and code are released publicly.
|
| |
| TuCT9 |
502 |
| Emerging Data Science in Manufacturing 2 |
Special Session |
| Organizer: Lee, Chia-Yen | National Taiwan University |
| Organizer: Hsu, Chia-Yu | National Tsing Hua University |
| Organizer: Lugaresi, Giovanni | KU Leuven |
| Organizer: Fan, Shu-Kai S. | National Taipei University of Technology |
| Organizer: Jang, Young Jae | Korea Advanced Institute of Science and Technology |
| Organizer: Hung, Yu-Hsin | National Yang Ming Chiao Tung University |
| Organizer: Lin, Chin-Yi | University of Texas at El Paso |
| Organizer: Tsai, Tsung-Han | National Taipei University of Business |
| Organizer: Ma, Kang-Ting | National Dong Hwa University |
| |
| 16:15-16:35, Paper TuCT9.1 | |
| Congestion-Aware Automated Roadmap Design for Automated Material Handling Systems (I) |
|
| Kwon, San | Korea Advanced Institute of Science and Technology |
| Jang, Young Jae | Korea Advanced Institute of Science and Technology |
Keywords: Factory Automation, Optimization and Optimal Control, Autonomous Vehicle Navigation
Abstract: Automated Material Handling Systems (AMHS) using Autonomous Mobile Robots (AMRs) are expanding rapidly in manufacturing, and as the fleet grows the AMR roadmap, the node and edge graph that governs where and how robots move on the factory floor, becomes essential for linking machines and workstations and maximizing production throughput. In industry these roadmaps are still hand-drawn by engineers from intuition rather than any systematic rule, producing designs prone to congestion at heavily utilized points such as ports and intersections. This study proposes an end-to-end pipeline that automatically generates roadmaps aware of these congestion points. From a CAD drawing, the pipeline extracts the floor topology and solves MILP to generate a roadmap. It then analyzes the roadmap's congestion using queueing theory and iteratively refines the graph until the predicted congestion stabilizes. The refined roadmap is validated by a factory emulator from DAIM Research Corporation, closing the gap to deployment.
|
| |
| 16:35-16:55, Paper TuCT9.2 | |
| Automated Vision-Based Congestion Detection System for Overhead Hoist Transports (I) |
|
| Sun, Chienshih | Korea Advanced Institute of Science and Technology |
| Jang, Young Jae | Korea Advanced Institute of Science and Technology |
Keywords: Semiconductor Manufacturing, Computer Vision for Manufacturing, Factory Automation
Abstract: This paper presents an automated vision-based framework for congestion detection in overhead hoist transport (OHT) systems used in semiconductor fabs. The proposed system processes top-view video to perform track segmentation, OHT detection, and congestion analysis without additional sensing infrastructure. Track topology is extracted through an OpenCV-based pipeline, and OHTs are detected using a YOLOv8l model trained by semiconductor recording data, achieving a mAP50 of 0.831 on an unseen layout. Key performance indicator (KPI) validation shows that sections flagged by the M/D/1-based criterion exhibit higher stop frequency and lower OHT speed, confirming the congestion definition. End-to-end experiments across three traffic densities show latency ranging from 0.81 s to 1.88 s.
|
| |
| 16:55-17:15, Paper TuCT9.3 | |
| An Empirical Study on Capacitor Defect Classification with Noisy Labels Using Semi-Supervised Learning (I) |
|
| Hsu, Chia-Yu | National Tsing Hua University |
| Teng, Yu-Han | National Taiwan University of Science and Technology |
Keywords: AI-Based Methods, Computer Vision for Manufacturing, Deep Learning in Robotics and Automation
Abstract: This study proposes a robust framework for defect classification in multi-layer ceramic capacitors (MLCCs) under noisy-label conditions by integrating a two-stage noisy-sample selection scheme with a noise-resilient learning model. In practical industrial inspection, training labels are often corrupted by annotator fatigue or insufficient expertise, which can severely impair model optimization and degrade classification accuracy. To address this challenge, a class-wise Gaussian mixture model is first employed to identify candidate clean samples, followed by confidence-based matching to further refine sample selection. Through this two-stage procedure, the training set is partitioned into clean and noisy subsets. The clean subset is used as labeled data, whereas the noisy subset is treated as unlabeled data after label removal for subsequent robust training. Based on this data refinement strategy, the proposed model is trained on high-quality labeled samples and further optimized by iteratively generating pseudo-labels for the unlabeled samples. Extensive experiments demonstrate that the proposed framework achieves reliable MLCC defect classification even under substantial label corruption. Under both symmetric and asymmetric noise settings, it consistently outperforms conventional supervised learning approaches across noise rates ranging from 10% to 60%. These findings confirm the effectiveness and robustness of the proposed method and underscore its potential for real-world industrial defect inspection
|
| |
| 17:15-17:35, Paper TuCT9.4 | |
| GMM-Based Air Compressor Operating Mode Clustering in Industrial Environments (I) |
|
| Wen, Han | University of Connecticut |
| Tu, Jiachen | University of Connecticut |
| Zhang, Liang | University of Connecticut |
Keywords: Machine learning
Abstract: In energy-efficiency assessments, current sensors and data loggers are routinely employed to monitor air-compressor performance over extended time horizons, generating datasets that capture a wide range of operating conditions. In this study, current time-series measurements obtained from an industrial manufacturing facility are analyzed to infer the operating modes of air compressors. To this end, a data-driven framework is proposed to automatically identify operating states by combining sliding-window time-series segmentation with Gaussian Mixture Model (GMM) clustering. The root-mean-square (RMS) value of the current signal is computed over overlapping windows to quantify the load level, and a temporal backfilling procedure is introduced to enhance the temporal consistency of the resulting state labels. The numerical results demonstrate that the proposed methodology robustly discriminates among all states.
|
| |
| 17:35-17:55, Paper TuCT9.5 | |
| An Embodied Perception Approach for Tool Remaining Useful Life in Multi-Agent Collaborative Manufacturing (I) |
|
| Li, Dingkun | Nanjing University of Aeronautics and Astronautics |
| Wang, Liping | Nanjing University of Aeronautics and Astronautics |
| Liu, Changchun | Nanjing University of Aeronautics and Astronautics, College of Mechanical and Electrical Engineering |
| Tang, Dunbing | Nanjing University of Aeronautics and Astronautics, College of Mechanical and Electrical Engineering |
| Zhu, Haihua | Nanjing University of Aeronautics and Astronautics |
Keywords: Agent-Based Systems, Diagnosis and Prognostics, Intelligent and Flexible Manufacturing
Abstract: In dynamic manufacturing scenarios, the task allocation and autonomous decision-making of embodied multi-agent systems critically depend on accurate perception of equipment health states. However, tool wear and breakage in CNC machine tools can severely affect machining quality, production continuity, and scheduling reliability. To address this issue, this study proposes a dual-stream perception network for tool remaining useful life prediction by jointly modeling machining process parameters and time-series monitoring signals. In the proposed framework, a ResNeXt-based dynamic branch is developed to extract degradation-sensitive temporal features from monitoring signals, while the static parameters branch utilizes a fully connected neural network to parse the physical attributes of machining parameters. Furthermore, a multi-head self-attention mechanism is introduced to achieve cross-modal feature fusion, thereby extracting spatiotemporal features sensitive to tool degradation states and realizing the prediction of remaining useful life. Experimental results on multi-condition milling data demonstrate that the proposed method achieves reliable RUL prediction, with a root mean square error as low as 0.036 and a coefficient of determination up to 0.961, indicating superior prediction accuracy, robustness, and cross-condition generalization capability. Furthermore, the predicted RUL values are stored in a database as state-awareness information accessible to equipment agents, laying a foundation for future research on perception-driven multi-agent decision-making.
|
| |
| 17:55-18:15, Paper TuCT9.6 | |
| Digital Twin-Enabled Embodied Intelligence with Multi-Agent Systems for Real-Time Perception-Decision-Execution in Manufacturing (I) |
|
| Du, ZhiWen | Yangzhou University |
| Geng, Junsai | Yangzhou University |
| Nie, Qingwei | Yangzhou University |
| Liu, Changchun | Nanjing University of Aeronautics and Astronautics |
Keywords: Agent-Based Systems
Abstract: Data acquisition on the manufacturing floor is no longer the primary bottleneck; instead, translating data analytics results into actionable physical operations has become the critical obstacle to autonomous manufacturing. Traditional intelligent manufacturing methods over-rely on predefined rules and offline modeling, often causing delayed system responses and significant deficiencies in flexibility and adaptability when facing unanticipated production disturbances, non-standard customized orders, or human-robot collaboration. This limitation stems from the inherent disconnect between the cognitive and execution layers. To bridge this gap, this paper proposes an embodied intelligence paradigm for constructing manufacturing systems with real-time closed-loop "perception-decision-execution". For implementation, digital twin technology provides a cost-effective virtual environment for trial-and-error. High-fidelity bidirectional mapping allows complex algorithm training and production line reconfiguration verification to be conducted in virtual space, eliminating downtime risks and hardware wear from physical commissioning. Meanwhile, a multi-agent system adopts distributed coordination, enabling each manufacturing unit to autonomously adjust behavior based on dynamic tasks and reduce reliance on centralized scheduling. Further analysis shows that transitioning to highly autonomous manufacturing systems depends on resolving four interrelated technical challenges: distributed agile decision-making in dynamic environments, high-frequency synchronous mapping between virtual and physical data, robust safety control for physical human-machine interaction, and the design of task-oriented layered embodied architectures. Overcoming these challenges will enable the manufacturing i
|
| |
| TuCT10 |
Convention Hall A |
| Best Paper Award Session 3 |
|
| |
| 16:35-16:55, Paper TuCT10.2 | |
| Dynamic Inpatient Admission Scheduling Optimization Based on Reinforcement Learning Methods |
|
| Zhao, Can | Zhejiang University |
| Zhang, Zheng | Zhejiang University |
| Ma, Yinghua | Zhejiang University |
| Ding, Weihao | University of Birmingham |
Keywords: Scheduling in Healthcare, Reinforcement
Abstract: Admission scheduling is widely applied in healthcare systems to balance admission demand and resource capacity. Existing studies largely focus on immediate admission or rejection decisions, with limited consideration of delayed admission. In this study, we investigate dynamic inpatient admission scheduling with delayed admission, where delayed patients are placed on a waiting list with guaranteed admission deadlines. We develop a two-stage decision framework that first responds to new requests through direct admission, delay, or rejection, and then dynamically admits patients from the waiting list. The problem is formulated as an infinite-horizon Markov decision process (MDP). However, tracking each patient’s admission deadline significantly increases the state dimensionality. To address this challenge, we propose a congestion score-based Actor–Critic (CSAC) algorithm and design an interpretable response policy. We further adopt an adaptive admission rule to allocate patients from the waiting list. Numerical experiments demonstrate the value of delayed admission and the effectiveness of the proposed CSAC algorithm, and provide insights for hospital operations in diverse environments.
|
| |
| 16:55-17:15, Paper TuCT10.3 | |
| Prior-Knowledge-Guided Reinforcement Learning for Musculoskeletal Locomotion |
|
| Tang, Jiayu | The Hong Kong University of Science and Technology (Guangzhou) |
| Huang, Zhenmin | The Hong Kong University of Science and Technology |
| Qin, Dezhi | The Hong Kong University of Science and Technology |
| Zhou, Jinni | Hong Kong University of Science and Technology (Guangzhou) |
| Ma, Jun | The Hong Kong University of Science and Technology |
Keywords: Rehabilitation, Biomimetics, Motion Control
Abstract: Reinforcement learning (RL) has achieved notable success in locomotion of robotic systems, yet remains challenged by high-dimensional and overactuated musculoskeletal systems. In this work, we propose a synergy-guided RL framework that incorporates biologically inspired priors extracted from pre-trained locomotion policies. Specifically, we derive low-dimensional muscle activation representations and functional muscle group structures from coordinated activation patterns. These priors are embedded into policy learning, enabling the control of structured muscle synergies instead of individual muscles. By reducing effective action dimensionality while preserving intrinsic coordination patterns, the proposed approach improves sample efficiency and training stability. Extensive experiments demonstrate that our method achieves fast convergence, high rewards, and robust locomotion under standard training conditions, while also generalizing effectively to unseen environmental settings, including stair, hilly, and rough terrain.
|
| |
| 17:15-17:35, Paper TuCT10.4 | |
| Signed Voxel Distance Field–Constrained Clinically Admissible Safe Milling Trajectory Planning for Autonomous Robotic Cochlear Implantation |
|
| Dong, Yufei | Nankai University |
| Sun, Xiaoxue | Nankai University |
| Song, Pei-Cheng | Nankai University |
| Wang, Hongpeng | Nankai University |
| Yu, Hongjian | Harbin Institute of Technology |
Keywords: Medical Robots and Systems, Motion and Path Planning, Robotics and Automation in Life Sciences
Abstract: Robotic cochlear implantation requires precise and safety-critical trajectory planning to avoid damage to nearby neural structures such as the facial nerve (FN) and chorda tympani (CT). Existing methods mainly rely on straight-line geometric trajectories and discrete distance checks, providing limited safety assessment in the complex temporal bone anatomy. This paper proposes a voxel-based Signed Distance Field (SDF)--constrained clinically admissible safe milling trajectory planning method for autonomous robotic cochlear implantation. An initial straight-line geometric access channel is constructed from the entry point to the cochlear target, and a localized voxelized milling region is established around it. By fusing multiple critical neural structures, a composite SDF is generated to enforce voxel-wise safety constraints and remove unsafe voxels that do not satisfy the predefined safety threshold, resulting in a continuously connected and robot-executable safe milling region along the initial path. Based on this region, a milling trajectory satisfying all safety constraints is generated. Full-path distance profiling quantitatively evaluates distances between trajectory points and nearby neural structures. Experiments on clinical CBCT datasets demonstrate that the method consistently maintains safety distances while producing continuous, anatomically consistent, and robot-executable milling paths.
|
| |
| 17:35-17:55, Paper TuCT10.5 | |
| A Clinically Informed Cross-Attention Deep Learning Framework for Automated Diagnosis of Obstructive Sleep Apnea |
|
| Li, Lingyi | Arizona State University |
| Si, Bing | Arizona State University |
Keywords: Automation in Life Science: Biotechnology, Pharmaceutical and Health Care, AI and Machine Learning in Healthcare, Machine learning
Abstract: Obstructive sleep apnea (OSA) is a highly prevalent breathing-related sleep disorder associated with a variety of cardiovascular comorbidities. Polysomnography (PSG) is the clinical gold standard for OSA diagnosis, with disease severity quantified by the apnea hypopnea index (AHI). However, deriving AHI requires manual annotation of overnight PSG recordings by certificated technicians, making the OSA diagnostic process costly and time-consuming. While it is highly desirable to automate this process with a reliable machine learning pipeline, conventional models often rely on handcrafted feature engineering and may therefore not sufficiently capture the complex dynamics of physiological signals. Deep learning models can learn nonlinear representations directly from raw PSG signals, but many existing works study PSG independently and give less consideration to clinical variables that are known to influence OSA risk, e.g., age, sex. More importantly, most prior research has focused on a single dataset, raising concerns about the robustness and generalizability of these models. To tackle these limitations, this project proposes a clinical knowledge-guided cross-attention deep learning framework for automated PSG annotation and AHI prediction from overnight PSG recordings, combined with clinical variables. The proposed framework employs a deep encoder to learn epoch-level features from PSG data and classify epochs based on whether an adverse respiratory event occurs, providing epoch-level PSG annotations that mimic manual evaluation in current practice. Then, the epoch-level embeddings will be integrated with clinical variables using a multi-head attention mechanism for patient-level AHI prediction. The framework is validated across four large real-world datasets.
|
| |