Kumbh 2015 Telecom Dataset for Reproducible Mass Gathering Research
Main Article Content
Abstract
This paper describes telecom data from a panel of 194,520 telecom records across eight telecom operators and 1,621 operator-site pairs, collected over five nonconsecutive dates for Kumbh 2015. An executable Python pipeline preserves 169,932 counts where observed, restores previous estimates for 715 counts estimated from the preceding three observed hours, and sets 23,873 unsupported gaps to missing. Temporal analysis shows that two operators have constant counts within a site-day and that 99.15% of adjacent observed pairs for a third operator are also unchanged. An expanding-date benchmark evaluates persistence, median hourly targets across past periods, and pooled log-ridge regression over 125,169 hourly target observations. Persistence produces the lowest pooled mean absolute error for every test date. The contribution is a documented event dataset, transparent preprocessing, and reference benchmarks for reuse.
Article Details

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
IJCERT Policy:
The published work presented in this paper is licensed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license. This means that the content of this paper can be shared, copied, and redistributed in any medium or format, as long as the original author is properly attributed. Additionally, any derivative works based on this paper must also be licensed under the same terms. This licensing agreement allows for broad dissemination and use of the work while maintaining the author's rights and recognition.
By submitting this paper to IJCERT, the author(s) agree to these licensing terms and confirm that the work is original and does not infringe on any third-party copyright or intellectual property rights.
References
D. Helbing, A. Johansson, and H. Z. Al-Abideen, “Dynamics of crowd disasters: An empirical study,” Physical Review E, vol. 75, no. 4, Art. no. 046109, 2007. DOI: 10.1103/PhysRevE.75.046109.
P. Deville et al., “Dynamic population mapping using mobile phone data,” Proceedings of the National Academy of Sciences of the USA, vol. 111, no. 45, pp. 15888–15893, 2014. DOI: 10.1073/pnas.1408439111.
F. Ricciato, P. Widhalm, F. Pantisano, and M. Craglia, “Beyond the single-operator, CDR-only paradigm: An interoperable framework for mobile phone network data analyses and population density estimation,” Pervasive and Mobile Computing, vol. 35, pp. 65–82, 2017. DOI: 10.1016/j.pmcj.2016.04.009.
M. D. Wilkinson et al., “The FAIR guiding principles for scientific data management and stewardship,” Scientific Data, vol. 3, Art. no. 160018, 2016. DOI: 10.1038/sdata.2016.18.
T. Gebru et al., “Datasheets for datasets,” Communications of the ACM, vol. 64, no. 12, pp. 86–92, 2021. DOI: 10.1145/3458723.
V. D. Blondel et al., “Data for development: The D4D challenge on mobile phone data,” arXiv preprint arXiv:1210.0137, 2012. DOI: 10.48550/arXiv.1210.0137.
G. Barlacchi et al., “A multi-source dataset of urban life in the city of Milan and the Province of Trentino,” Scientific Data, vol. 2, Art. no. 150055, 2015. DOI: 10.1038/sdata.2015.55.
Y.-A. de Montjoye, Z. Smoreda, R. Trinquart, C. Ziemlicki, and V. D. Blondel, “D4D-Senegal: The second mobile phone data for development challenge,” arXiv preprint arXiv:1407.4885, 2014. DOI: 10.48550/arXiv.1407.4885.
O. E. Martínez-Durive, S. Mishra, C. Ziemlicki, S. Rubrichi, Z. Smoreda, and M. Fiore, “The NetMob23 dataset: A high-resolution multi-region service-level mobile data traffic cartography,” arXiv preprint arXiv:2305.06933, 2023. DOI: 10.48550/arXiv.2305.06933.
J. Lloret, J. Tomas, A. Canovas, and L. Parra, “An integrated IoT architecture for smart metering,” IEEE Communications Magazine, vol. 54, no. 12, pp. 50–57, 2016. DOI: 10.1109/MCOM.2016.1600647CM.
G. F. Scaranti, L. F. Carvalho, S. Barbon Junior, J. Lloret, and M. L. Proença Jr., “Unsupervised online anomaly detection in software defined network environments,” Expert Systems with Applications, vol. 191, Art. no. 116225, 2022. DOI: 10.1016/j.eswa.2021.116225.
A. Cini, I. Marisca, and C. Alippi, “Filling the Gaps: Multivariate time series imputation by graph neural networks,” in Proceedings of the International Conference on Learning Representations (ICLR), 2022.
H. Wickham, “Tidy data,” Journal of Statistical Software, vol. 59, no. 10, pp. 1–23, 2014. DOI: 10.18637/jss.v059.i10.
NIST/SEMATECH, “Box plot,” in e-Handbook of Statistical Methods. Accessed: Sep. 16, 2026.
L. J. Tashman, “Out-of-sample tests of forecasting accuracy: An analysis and review,” International Journal of Forecasting, vol. 16, no. 4, pp. 437–450, 2000. DOI: 10.1016/S0169-2070(00)00065-0.
A. E. Hoerl and R. W. Kennard, “Ridge regression: Biased estimation for nonorthogonal problems,” Technometrics, vol. 12, no. 1, pp. 55–67, 1970. DOI: 10.1080/00401706.1970.10488634.
R. J. Hyndman and A. B. Koehler, “Another look at measures of forecast accuracy,” International Journal of Forecasting, vol. 22, no. 4, pp. 679–688, 2006. DOI: 10.1016/j.ijforecast.2006.03.001.
G. K. Sandve, A. Nekrutenko, J. Taylor, and E. Hovig, “Ten simple rules for reproducible computational research,” PLoS Computational Biology, vol. 9, no. 10, Art. no. e1003285, 2013. DOI: 10.1371/journal.pcbi.1003285.
S. Kapoor and A. Narayanan, “Leakage and the reproducibility crisis in machine-learning-based science,” Patterns, vol. 4, no. 9, Art. no. 100804, 2023. DOI: 10.1016/j.patter.2023.100804.
Y.-A. de Montjoye, C. A. Hidalgo, M. Verleysen, and V. D. Blondel, “Unique in the crowd: The privacy bounds of human mobility,” Scientific Reports, vol. 3, Art. no. 1376, 2013. DOI: 10.1038/srep01376.