Data-platform agents operate in environments where schemas, APIs, organizational policies, and local practices change continuously. Static prompts and fixed tool descriptions therefore decay, while unconstrained learning from successful interactions risks reproducing obsolete shortcuts and unsafe habits. This research agenda examines delivery traces as a source of situated knowledge for continually improving analytics agents. SiriusBI supplies an upstream view in which multi-round clarification and data-conditioned generation adapt business questions to domain context. SiriusDeliver supplies a downstream view in which hierarchical skills, artifact lifecycle control, and trace-driven skill evolution support warehouse delivery at production scale. The agenda connects these systems with research on retrieval-augmented generation, reasoning and acting, learned tool use, production technical debt, readiness testing, data-quality verification, data cascades, documentation, human-AI interaction, and limits of intrinsic self-correction. It proposes a trace-to-skill pipeline with selection, abstraction, specification, testing, approval, deployment, monitoring, and retirement stages. It also defines an evaluation program for learning benefit, safety, drift, and organizational distribution. The central claim is that traces should be treated as governed evidence, not as demonstrations that automatically deserve imitation. Continual improvement becomes credible only when learned procedures remain scoped, testable, reversible, and attributable.
- Jiang, J., Xie, H., Yang, J., Shen, S., Wang, Z., Zheng, Y., ... & Jiang, J. (2024). Siriusbi: A comprehensive llm-powered solution for data analytics in business intelligence. *arXiv preprint arXiv:2411.06102*.
- Xie, H., Zhou, X., Yang, J., Shen, S., Wang, Z., Zheng, Y., ... & Jiang, J. (2026). SiriusDeliver: Automating Data Warehouse Delivery at Tencent. *arXiv preprint arXiv:2608.09185*.
- Amershi, S., Weld, D., Vorvoreanu, M., Fourney, A., Nushi, B., Collisson, P., et al. (2019). Guidelines for human-AI interaction. In *Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems* (pp. 1-13). https://doi.org/10.1145/3290605.3300233 DOI
- Breck, E., Cai, S., Nielsen, E., Salib, M., & Sculley, D. (2017). The ML test score: A rubric for ML production readiness and technical debt reduction. In *2017 IEEE International Conference on Big Data* (pp. 1123-1132). https://doi.org/10.1109/BigData.2017.8258038 DOI
- Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Daume III, H., & Crawford, K. (2021). Datasheets for datasets. *Communications of the ACM, 64*(12), 86-92. https://doi.org/10.1145/3458723 DOI
- Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2024). Large language models cannot self-correct reasoning yet. In *International Conference on Learning Representations*. https://openreview.net/forum?id=IkmD3fKBPQ
- Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., et al. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. *Advances in Neural Information Processing Systems, 33*.
- Sambasivan, N., Kapania, S., Highfill, H., Akrong, D., Paritosh, P., & Aroyo, L. M. (2021). Everyone wants to do the model work, not the data work: Data cascades in high-stakes AI. In *Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems* (pp. 1-15). https://doi.org/10.1145/3411764.3445518 DOI
- Schelter, S., Lange, D., Schmidt, P., Celikel, M., Biessmann, F., & Grafberger, A. (2018). Automating large-scale data quality verification. *Proceedings of the VLDB Endowment, 11*(12), 1781-1794. https://doi.org/10.14778/3229863.3229867 DOI
- Schick, T., Dwivedi-Yu, J., Dessi, R., Raileanu, R., Lomeli, M., Hambro, E., et al. (2023). Toolformer: Language models can teach themselves to use tools. *Advances in Neural Information Processing Systems, 36*. https://doi.org/10.52202/075280-2997 DOI
- Sculley, D., Holt, G., Golovin, D., Davydov, E., Phillips, T., Ebner, D., et al. (2015). Hidden technical debt in machine learning systems. *Advances in Neural Information Processing Systems, 28*, 2503-2511.
- Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2023). ReAct: Synergizing reasoning and acting in language models. In *International Conference on Learning Representations*. https://openreview.net/forum?id=WE_vluYUL-X
- Journal
- Journal of Artificial Intelligence and Interdisciplinary Research
- Volume
- 1 (2026)
- Article number
- aji20260006
- License
- CC BY 4.0
