| 摘 要: 针对工程量清单计价评审中依赖经验判定、单一阈值难以区分异常类型及纯机器学习模型缺乏可解释性等行业痛点,构建了涵盖量价偏差、文本语义、项目环境及无监督异常的高维多层特征体系,采用六种异质树模型进行Stacking集成并引入孤立森林提升对未知分布异常的敏感度,设计12条评审规则实现量价分解归因与新增项细分,通过优先级融合策略整合规则结论与模型预测。实验表明,固定阈值15%方法准确率仅90.84%、漏检率17.91%、归因正确率50.45%,而融合方法5折交叉验证准确率达98.09%,宏平均F1值0.9432,漏检率0.37%;在跨项目验证中,LOPO平均准确率97.72%,变异系数6.82%,实现可解释的细粒度评审,满足评审规范对证据链可追溯性的要求。 |
| 关键词: 集成学习 规则引擎 工程量清单 异常检测 |
|
中图分类号:
文献标识码:
|
|
| Anomaly Detection Method for Bill of Quantities Based on Ensemble Learning and Rule Engine |
|
ZHANG You
|
Shijiazhuang Construction Project Review and Evaluation Center
|
| Abstract: To address the industry pain points in bill-of-quantity (BOQ) pricing review—such as reliance on empirical judgment, difficulty in distinguishing anomaly types with a single threshold, and lack of interpretability in pure machine learning models—this study constructs a high-dimensional multi?layer feature system covering quantity–price deviations, textual semantics, project context, and unsupervised anomalies. Six heterogeneous tree?based models are integrated via Stacking, with Isolation Forest incorporated to enhance sensitivity to anomalies with unknown distributions. Twelve review rules are designed to enable quantity–price decomposition attribution and subdivision of newly added items, and a priority fusion strategy is adopted to integrate rule?based conclusions with model predictions. Experimental results show that the fixed 15% threshold method achieves only 90.84% accuracy, 17.91% miss rate, and 50.45% attribution accuracy, whereas the proposed fusion method attains 98.09% accuracy in 5?fold cross?validation, with a macro?averaged F1?score of 0.9432 and a miss rate of 0.37%. In cross?project validation, the LOPO average accuracy reaches 97.72% with a coefficient of variation of 6.82%, demonstrating interpretable fine?grained review that satisfies the requirement for evidence?chain traceability in review specifications. |
| Keywords: Ensemble learning Rule engine Bill of quantities Anomaly detection |