Kumbh 2015 Telecom Dataset for Reproducible Mass Gathering Research

Main Article Content

Lavanya Addepalli
Sandip Kashinath Shinde
Sachin Raghunath Pachorkar
Girish Pandit Pagare
Jaime Lloret

Abstract

This paper describes telecom data from a panel of 194,520 telecom records across eight telecom operators and 1,621 operator-site pairs, collected over five nonconsecutive dates for Kumbh 2015. An executable Python pipeline preserves 169,932 counts where observed, restores previous estimates for 715 counts estimated from the preceding three observed hours, and sets 23,873 unsupported gaps to missing. Temporal analysis shows that two operators have constant counts within a site-day and that 99.15% of adjacent observed pairs for a third operator are also unchanged. An expanding-date benchmark evaluates persistence, median hourly targets across past periods, and pooled log-ridge regression over 125,169 hourly target observations. Persistence produces the lowest pooled mean absolute error for every test date. The contribution is a documented event dataset, transparent preprocessing, and reference benchmarks for reuse.

Article Details

How to Cite
[1]
Lavanya Addepalli, Sandip Kashinath Shinde, Sachin Raghunath Pachorkar, Girish Pandit Pagare, and Jaime Lloret, “Kumbh 2015 Telecom Dataset for Reproducible Mass Gathering Research”, Int. J. Comput. Eng. Res. Trends, vol. 13, no. 3, pp. 26–34, Sep. 2026.
Section
Research Articles

References

D. Helbing, A. Johansson, and H. Z. Al-Abideen, “Dynamics of crowd disasters: An empirical study,” Physical Review E, vol. 75, no. 4, Art. no. 046109, 2007. DOI: 10.1103/PhysRevE.75.046109.

P. Deville et al., “Dynamic population mapping using mobile phone data,” Proceedings of the National Academy of Sciences of the USA, vol. 111, no. 45, pp. 15888–15893, 2014. DOI: 10.1073/pnas.1408439111.

F. Ricciato, P. Widhalm, F. Pantisano, and M. Craglia, “Beyond the single-operator, CDR-only paradigm: An interoperable framework for mobile phone network data analyses and population density estimation,” Pervasive and Mobile Computing, vol. 35, pp. 65–82, 2017. DOI: 10.1016/j.pmcj.2016.04.009.

M. D. Wilkinson et al., “The FAIR guiding principles for scientific data management and stewardship,” Scientific Data, vol. 3, Art. no. 160018, 2016. DOI: 10.1038/sdata.2016.18.

T. Gebru et al., “Datasheets for datasets,” Communications of the ACM, vol. 64, no. 12, pp. 86–92, 2021. DOI: 10.1145/3458723.

V. D. Blondel et al., “Data for development: The D4D challenge on mobile phone data,” arXiv preprint arXiv:1210.0137, 2012. DOI: 10.48550/arXiv.1210.0137.

G. Barlacchi et al., “A multi-source dataset of urban life in the city of Milan and the Province of Trentino,” Scientific Data, vol. 2, Art. no. 150055, 2015. DOI: 10.1038/sdata.2015.55.

Y.-A. de Montjoye, Z. Smoreda, R. Trinquart, C. Ziemlicki, and V. D. Blondel, “D4D-Senegal: The second mobile phone data for development challenge,” arXiv preprint arXiv:1407.4885, 2014. DOI: 10.48550/arXiv.1407.4885.

O. E. Martínez-Durive, S. Mishra, C. Ziemlicki, S. Rubrichi, Z. Smoreda, and M. Fiore, “The NetMob23 dataset: A high-resolution multi-region service-level mobile data traffic cartography,” arXiv preprint arXiv:2305.06933, 2023. DOI: 10.48550/arXiv.2305.06933.

J. Lloret, J. Tomas, A. Canovas, and L. Parra, “An integrated IoT architecture for smart metering,” IEEE Communications Magazine, vol. 54, no. 12, pp. 50–57, 2016. DOI: 10.1109/MCOM.2016.1600647CM.

G. F. Scaranti, L. F. Carvalho, S. Barbon Junior, J. Lloret, and M. L. Proença Jr., “Unsupervised online anomaly detection in software defined network environments,” Expert Systems with Applications, vol. 191, Art. no. 116225, 2022. DOI: 10.1016/j.eswa.2021.116225.

A. Cini, I. Marisca, and C. Alippi, “Filling the Gaps: Multivariate time series imputation by graph neural networks,” in Proceedings of the International Conference on Learning Representations (ICLR), 2022.

H. Wickham, “Tidy data,” Journal of Statistical Software, vol. 59, no. 10, pp. 1–23, 2014. DOI: 10.18637/jss.v059.i10.

NIST/SEMATECH, “Box plot,” in e-Handbook of Statistical Methods. Accessed: Sep. 16, 2026.

L. J. Tashman, “Out-of-sample tests of forecasting accuracy: An analysis and review,” International Journal of Forecasting, vol. 16, no. 4, pp. 437–450, 2000. DOI: 10.1016/S0169-2070(00)00065-0.

A. E. Hoerl and R. W. Kennard, “Ridge regression: Biased estimation for nonorthogonal problems,” Technometrics, vol. 12, no. 1, pp. 55–67, 1970. DOI: 10.1080/00401706.1970.10488634.

R. J. Hyndman and A. B. Koehler, “Another look at measures of forecast accuracy,” International Journal of Forecasting, vol. 22, no. 4, pp. 679–688, 2006. DOI: 10.1016/j.ijforecast.2006.03.001.

G. K. Sandve, A. Nekrutenko, J. Taylor, and E. Hovig, “Ten simple rules for reproducible computational research,” PLoS Computational Biology, vol. 9, no. 10, Art. no. e1003285, 2013. DOI: 10.1371/journal.pcbi.1003285.

S. Kapoor and A. Narayanan, “Leakage and the reproducibility crisis in machine-learning-based science,” Patterns, vol. 4, no. 9, Art. no. 100804, 2023. DOI: 10.1016/j.patter.2023.100804.

Y.-A. de Montjoye, C. A. Hidalgo, M. Verleysen, and V. D. Blondel, “Unique in the crowd: The privacy bounds of human mobility,” Scientific Reports, vol. 3, Art. no. 1376, 2013. DOI: 10.1038/srep01376.