Language-model assistants are frequently evaluated at the moment they generate SQL or transformation code, although production data engineering succeeds or fails across a much longer artifact lifecycle. Requirements are interpreted, sources and dependencies are selected, code and configuration are created, tests are executed, jobs are submitted, outputs are observed, and changes are revised or rolled back. This evidence synthesis argues that the artifact, rather than the chat session or model response, should be the primary control object for LLM-assisted data engineering. SiriusBI demonstrates how modular coordination, clarification, and adaptive SQL-generation strategies can produce analytic artifacts from business intent. SiriusDeliver moves further into operational delivery with hierarchical skill orchestration, pre- and post-execution artifact control, and trace-driven skill evolution. Their combination reveals a practical control plane built from versioned states, explicit invariants, provenance, staged authorization, and executable quality checks. Research on data provenance, automated data-quality verification, production readiness, technical debt, human-AI interaction, dataset documentation, and tool-using agents supplies complementary mechanisms. The synthesis defines a lifecycle state model, analyzes control gates and failure containment, and identifies governance requirements for learning from traces without institutionalizing past mistakes.
- Jiang, J., Xie, H., Yang, J., Shen, S., Wang, Z., Zheng, Y., ... & Jiang, J. (2024). Siriusbi: A comprehensive llm-powered solution for data analytics in business intelligence. *arXiv preprint arXiv:2411.06102*.
- Xie, H., Zhou, X., Yang, J., Shen, S., Wang, Z., Zheng, Y., ... & Jiang, J. (2026). SiriusDeliver: Automating Data Warehouse Delivery at Tencent. *arXiv preprint arXiv:2608.09185*.
- Amershi, S., Weld, D., Vorvoreanu, M., Fourney, A., Nushi, B., Collisson, P., et al. (2019). Guidelines for human-AI interaction. In *Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems* (pp. 1-13). https://doi.org/10.1145/3290605.3300233 DOI
- Breck, E., Cai, S., Nielsen, E., Salib, M., & Sculley, D. (2017). The ML test score: A rubric for ML production readiness and technical debt reduction. In *2017 IEEE International Conference on Big Data* (pp. 1123-1132). https://doi.org/10.1109/BigData.2017.8258038 DOI
- Buneman, P., Khanna, S., & Tan, W. C. (2001). Why and where: A characterization of data provenance. In *Database Theory - ICDT 2001* (pp. 316-330). Springer. https://doi.org/10.1007/3-540-44503-X_20 DOI
- Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Daume III, H., & Crawford, K. (2021). Datasheets for datasets. *Communications of the ACM, 64*(12), 86-92. https://doi.org/10.1145/3458723 DOI
- Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., et al. (2019). Model cards for model reporting. In *Proceedings of the Conference on Fairness, Accountability, and Transparency* (pp. 220-229). https://doi.org/10.1145/3287560.3287596 DOI
- Sambasivan, N., Kapania, S., Highfill, H., Akrong, D., Paritosh, P., & Aroyo, L. M. (2021). Everyone wants to do the model work, not the data work: Data cascades in high-stakes AI. In *Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems* (pp. 1-15). https://doi.org/10.1145/3411764.3445518 DOI
- Schelter, S., Lange, D., Schmidt, P., Celikel, M., Biessmann, F., & Grafberger, A. (2018). Automating large-scale data quality verification. *Proceedings of the VLDB Endowment, 11*(12), 1781-1794. https://doi.org/10.14778/3229863.3229867 DOI
- Schick, T., Dwivedi-Yu, J., Dessi, R., Raileanu, R., Lomeli, M., Hambro, E., et al. (2023). Toolformer: Language models can teach themselves to use tools. *Advances in Neural Information Processing Systems, 36*. https://doi.org/10.52202/075280-2997 DOI
- Sculley, D., Holt, G., Golovin, D., Davydov, E., Phillips, T., Ebner, D., et al. (2015). Hidden technical debt in machine learning systems. *Advances in Neural Information Processing Systems, 28*, 2503-2511.
- Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2023). ReAct: Synergizing reasoning and acting in language models. In *International Conference on Learning Representations*. https://openreview.net/forum?id=WE_vluYUL-X
- Journal
- Journal of Artificial Intelligence and Interdisciplinary Research
- Volume
- 1 (2026)
- Article number
- aji20260005
- License
- CC BY 4.0
