Abstract:
To address the problems of a fixed desliming particle size, independent control of the “pre-deslim-ing + dense-medium separation” processes, and the difficulty in achieving optimal economic benefits under fluctuating coal quality during thermal coal preparation, a multi-agent deep reinforcement learning-based collaborative optimization decision-making method is proposed. The raw coal desliming and dense-medium separation processes are modeled as two cooperative agents, and a team reward function that considers both economic benefits and product quality is designed. An attention mechanism is introduced to optimize the policy network, and a digital twin-based simulation environment is developed. Based on the Centralized Training with Decentralized Execution framework, collaborative training and simulation testing of the agents are conducted. To verify the superiority of this strategy, comparative analyses were performed against a fixed-parameter strategy and an independent single-agent control strategy. The results show that under typical coal quality fluctuation conditions, this strategy increases cumulative economic benefits by more than 5.6%, achieves a clean coal ash content qualification rate of up to 99.8%, and maintains an average fluctuation in separation density of 0.004 g/cm
3. Under extreme conditions involving abrupt changes in coal quality, the strategy still maintains positive economic growth. Even under severe conditions where the fine-particle content exceeds 50%, economic benefits can still be improved by 1.5%. The proposed method overcomes the limitations of local optimization at individual stages of the coal preparation system and enables cross-process collaborative control of raw coal desliming and dense-medium separation, providing theoretical and methodological support for the transformation of coal preparation intelligence from “perception–control” toward “cognition–decision-making.”