Number of Issues

5

Article View

30,158

PDF Download

33,520

View Per Article

478.7

PDF Download Per Article

532.06

Number of Submissions

108

Rejected Submissions

7

Reject Rate

6

Accepted Submissions

61

Acceptance Rate

56

Time to Accept (Days)

172

Number of Indexing Databases

10

Number of Reviewers

97

The Journal of Data Science and Modeling is an open-access, double-blind, peer-reviewed journal published by Allameh Tabataba’i University the leading university in Humanities and Social Sciences in Iran. Data Science and Modeling has been established to provide an intellectual platform for national and international researchers working on issues related to Data Science and Modeling.

To allow for easy and worldwide access to the most updated research findings, the journal is set to be an open-access journal. All submitted papers should report original and unpublished experimental or theoretical research results until they will be reviewed. Papers submitted to the journal should meet those criteria and must not be under consideration for publishing elsewhere. The journal is published in both a print and an online version.

About Journal of Data Science and Modeling:

  • Country of Publication: Iran
  • Publisher: Allameh Tabataba’i University Press
  • Format: Print and Online
  • Start Publishing: 2022
  • Print-ISSN: 2676-5926
  • E-ISSN: 2980-9010
  • Available from: ATU Press, Google Scholar, Magiran, Civilica, ...
  • Impact Factor (ISC): --
  • Frequency: Semiannual
  • Publication Dates: 2022
  • Language: English
  • Scope: Computational Statistics, Statistical modeling,  Data science.
  • Article Processing Charges: No
  • Type of Journal: Academic / Scholarly
  • Open Access: Yes
  • Indexed & Abstracted: Google Scholar, Civilica
  • Policy: Peer Review
  • Initial Review Time: up to ten days
  • Review Time: Three Weeks, approximately
  • Contact E-mail: jcsm@atu.ac.ir
  • Alternate E-mail: askandari@atu.ac.ir
Neural Network

Applying the Modified Sinc Neural Network for Weather Forecasting

Pages 1-28

https://doi.org/10.22054/jdsm.2025.84908.1065

Ghasem Ahmadi

Abstract Accurate weather prediction plays a vital role in many sectors, such as agriculture, disaster preparedness, transportation systems, and urban planning. Traditional meteorological models face challenges in capturing complex atmospheric dynamics, leading to increased reliance on artificial neural networks (ANNs) for improved forecasting accuracy. ANNs have been widely applied in meteorology due to their ability to model nonlinear relationships and temporal dependencies. Based on the Sinc numerical methods, the modified Sinc neural network (MSNN) has been introduced recently. This model uses the advantages of the Sinc function, such as smoothness and fluctuation, and at the same time improves the ability to model nonlinear dependencies and temporal dynamics in environmental data. This work utilizes the MSNN for time series forecasting where its parameters are adjusted with a discrete-time online Lyapunov-based learning algorithm. Then, it is applied to enhance the weather forecasting. This model is evaluated on datasets containing various meteorological variables. The data used in this article is related to the city of Khorramabad in Iran. The results show that despite its simple structure, MSNN has a high efficiency in weather forecasting.

Machine Learning

PCA by Shrinkage Estimation: A Comprehensive Mathematical and Statistical Analysis

Pages 29-47

https://doi.org/10.22054/jdsm.2025.86705.1074

Parviz Nasiri, Heydar Mokhtari Farivar

Abstract Principal Component Analysis (PCA) is a cornerstone technique for dimensionality
reduction and data analysis. However, classic PCA can exhibit instability in
high-dimensional settings where the number of variables significantly exceeds the
number of observations. Shrinkage-based PCA addresses this limitation by incorporating
regularization into the covariance matrix estimation process, leading to
more stable and interpretable results. This paper provides a robust mathematical
and statistical foundation for shrinkage-based PCA, compares its performance with
classic PCA, and demonstrates its advantages through theoretical analysis, numerical
simulations, and real-world data experiments. It is important to note that using the idea of a contraction estimator increases the efficiency of the estimator. mean time in this paper, it is shown that the covariance matrix estimator resulting from the contraction estimator is very efficient.
It is also worth mentioning that to increase the efficiency of the contraction estimator, the recently discussed interval contraction estimator can be used.
keywords: principal component analysis, Shrinkage-based, Estimation, Covariance Structures, Simulation.

Mathematical Computing

Unsteady numerical simulation and data analysis of hysteresis associated with water uptake in capillary tube

Pages 49-78

https://doi.org/10.22054/jdsm.2025.88414.1078

Farnood Freidooni, Ali Rajabpour, Sina Nasiri

Abstract Capillary action and water uptake are tremendously fundamental and practical phenomena which used in a wide range of applications from industries and medical to agricultural fields. This work aims to provide a detailed numerical investigation and statistical data sampling of capillary action and water uptake, considering hysteresis associated with density, surface tension, contact angle, gravity force, tube diameter, and inclination angle effects. The main selected domain is 1mm and 5mm in diameter and height. The solver type is chosen as a pressure-based solver, and time-dependent data sampling is utilized. The flow field is selected as incompressible, constant properties, Newtonian homogeneous fluid. The finite volume method on a co-located grid system is used. The code uses algebraic multigrid schemes to accelerate the solution. The message passing interface parallelized code is used. The bisection algorithms are used for partitioning. The pressure and velocity fields were coupled using the PISO algorithm. The results show that increasing capillary tube diameter or surface tension enhances uptake velocity by 98–100% and reduces filling time by 49–50%, respectively, though inertial/dissipative effects caused minor deviations (1–12%) in surface tension cases. Flow velocity scaled linearly with contact angle doubling filling time, while gravitational acceleration induced only marginal delays with negligible meniscus impact, supporting its omission in engineering models. Transient meniscus asymmetry occurred in inclined tubes (45°) due to contact angle disparity between halves, yet filling duration remained identical to vertical and horizontal orientations despite geometric differences in meniscus evolution.

Neural Network

‎Recurrent Neural Networks for Loan Default Prediction: A Dual Deterministic and Uncertainty-Aware Framework

Pages 79-93

https://doi.org/10.22054/jdsm.2026.89154.1082

Navid Ashraf, Shokouh Shahbeyk, Hossein Teimoori Faal

Abstract Abstract of paper: This study explores the application of Recurrent Neural Networks (RNNs) for predicting loan defaults, with a particular emphasis on incorporating uncertainty estimation into the predictive framework. Conventional RNN models demonstrate high accuracy, but they fail to provide quantitative measures of prediction uncertainty. To address this limitation, a dual modeling approach is proposed: a standard RNN model for achieving high predictive accuracy and an uncertainty-aware RNN model incorporating Bayesian inference. The uncertainty-aware model enables enhanced risk assessment through confidence level estimation and improved capture of complex temporal dependencies in financial data. Experimental results indicate that both proposed models outperform traditional methods, with the uncertainty-aware variant offering superior risk evaluation capabilities through its probabilistic outputs. These findings contribute to advancing credit risk assessment methodologies and offer practical value for financial institutions seeking more robust default prediction systems.
Keyword: Recurrent Neural Networks (RNNs), Loan Default Prediction, Uncertainty Quantification, Credit Risk Assessment‎.

Statistical Computing

Diagnostic measures based on restricted ridge estimator in linear mixed measurement error models

Pages 95-122

https://doi.org/10.22054/jdsm.2026.86329.1070

Fatemeh Ghapani

Abstract This article focuses on diagnostic measures for identifying high-leverage points, and influential observations in linear mixed measurement error (LMME) models. It achieves by imposing the stochastic restrictions on the parameters and incorporating the ridge estimator to tackle the issue of multicollinearity. To this end, generalized leverage matrices are defined using the restricted ridge estimator (RRE) to identify high-leverage observations. Additionally, analogs of Cook’s distance and likelihood distance are proposed to determine influential observations through a case deletion approach. Simulation studies and real-life applications support the theoretical results. To the best of our knowledge, there has been no significant attention in the literature regarding diagnostics for leverage and influence measures concerning the outcomes of the RRE in LMME models. Hence, this paper evaluates the influence of observations by using leverage and influence measures to identify influential observations on the RRE’s of fixed effects and the prediction of random effects in LMME models

Computational Statistics

Zero-Inflated Two-Parameter Distribution for Modeling Overdispersed Count Data

Pages 123-142

https://doi.org/10.22054/jdsm.2026.85641.1069

zahra Karimiezmareh, Behdad Mostafaiy

Abstract In this paper, we propose a new two-parameter discrete distribution based on central Bell expansion, which is zero-inflated and designed to effectively model overdispersed count data. We study several structural properties of the proposed distribution and demonstrate that it is infinitely divisible, which adds theoretical strength and potential for wider applicability. The paper also discusses parameter estimation techniques for the distribution, focusing on two common approaches: the method of moments and the maximum likelihood estimation method. Both methods are developed and explained in detail. To evaluate the accuracy and reliability of these estimators, a simulation study is conducted across different sample sizes, allowing us to assess their performance under various conditions. To illustrate the practical importance and usefulness of the new distribution, we apply it to two real data sets and show how well it fits the observed data, reinforcing its value as a flexible tool for analyzing count data.

Statistical Computing

The unit linear exponential distribution: properties, quantile regression model and applications

Pages 143-174

https://doi.org/10.22054/jdsm.2026.88441.1077

Lazhar BENKHELIFA

Abstract The paper suggests a novel model defined on the unit interval, termed the unit linear exponential distribution, constructed via an inversion of the exponential function. The proposed model contains the unit exponential and the unit Rayleigh distributions as special submodels. Fundamental properties of the introduced distribution are discussed which are stochastic ordering property, quantile function, incomplete moments, moments, probability weighted moment, order statistics, stochastic orderings, stress strength reliability, and Tsallis and Renyi entropies. The distribution has two unknown parameters, which are estimated utilizing the following methods: maximum likelihood, maximum product spacing, least and weighted least squares, Cramér-von Mises, and Anderson-Darling. The behavior of these estimators is assessed through a simulation study. Furthermore, the paper develops a novel quantile regression model based on suggested distribution, which is shown to be a good alternative to existing models like the Kumaraswamy, beta, and unit Chen quantile regression models. We estimate the parameters of the regression model utilizing maximum likelihood. Two well-known real data applications are given to prove the modeling capability of the newly suggested distribution and quantile regression model.

Machine Learning

CheatingRank: A Multi-Stage Approach for Detecting Cheating in Online Assessments

Pages 175-199

https://doi.org/10.22054/jdsm.2026.86564.1071

Seyede Fatemeh Noorani, Mohammad Reza Mohammadi, Maryam Karimi, Nasrin Taherkhani

Abstract In this study, a multi-stage approach based on online exam data analysis was proposed to identify and rank students' Cheating Ranks. Initially, student response sheets were clustered using the $K-means++$ algorithm with dynamic determination of K, forming groups with similar characteristics. Subsequently, in later stages, each student's Cheating Rank was determined based on various behavioral and performance parameters.

Results demonstrated that the Cheating Rank derived from online exam data effectively differentiates between students suspicious of cheating and normal students, with statistically significant differences between the two groups. These findings underscore the validity and efficacy of the proposed method in detecting cheating in online exams.

Additionally, the impact of threshold selection for group differentiation highlighted the importance of appropriate parameter tuning in enhancing detection accuracy and influencing model sensitivity and reliability. The use of in-person exam scores as a reference criterion strengthened the results' credibility and enabled more objective model evaluation.

Given the limitations in sample size and data scope, future research should focus on larger, more diverse, and multidimensional datasets to improve both diagnostic accuracy and model generalizability. Furthermore, integrating this approach with advanced machine learning techniques and behavioral analytics could significantly enhance online exam integrity monitoring systems.

Overall, this research represents a significant step toward developing cost-effective, reliable, and efficient methods to reduce cheating in online educational environments, fostering greater trust among instructors and students in assessment processes.

Bayesian Computation Statistics

On weight and variance uncertainty in neural networks for regression tasks

Pages 201-222

https://doi.org/10.22054/jdsm.2026.88820.1080

Morteza Amini, Moein Monemi, Mahmoud Taheri, Mohammad Arashi

Abstract We investigate the problem of weight uncertainty originally proposed by [Blundell et al. (2015). Weight uncertainty in neural networks. In International conference on machine learning, 1613-1622, PMLR.] in the context of neural networks designed for regression tasks, and we extend their framework by incorporating variance uncertainty into the model. Our analysis demonstrates that explicitly modeling uncertainty in the variance parameter can significantly enhance the predictive performance of Bayesian neural networks. By considering a full posterior distribution over the variance, the model achieves improved generalization compared to approaches that treat variance as fixed or deterministic. We evaluate the generalization capability of our proposed approach through a function approximation example and further validate it on the riboflavin genetic dataset. Our exploration encompasses both fully connected dense networks and dropout neural networks, employing Gaussian and spike-and-slab priors respectively for the network weights, providing a comprehensive assessment of how variance uncertainty affects model performance across different architectural choices.

Machine Learning

Supporting Critical Services with Intelligent Placement in the Computing Continuum

Pages 223-253

https://doi.org/10.22054/jdsm.2026.86787.1073

Reza Sookhtsaraei, Mehdi Sakhaei-nia, Fereshteh Azadi Parand

Abstract Critical services increasingly rely on distributed computing infrastructures that can deliver fast and reliable responses under dynamic conditions. Traditional cloud-centric deployments often struggle to meet these demands due to latency and reliability constraints. This paper addresses these challenges by exploring intelligent service placement across the cloud–edge computing continuum. We introduce an adaptive placement framework that continuously adjusts deployment decisions in response to changing system states and service requirements. To improve robustness, the framework integrates learning transfer mechanisms and a criticality-aware resilience strategy that prioritizes service continuity during failures. The proposed approach enhances responsiveness and reliability, leading to a higher proportion of services meeting strict timing constraints. Comprehensive experimental evaluations confirm that the framework provides more effective support for critical and delay-sensitive services compared to existing placement strategies in continuum-based environments.

Bayesian Computation Statistics

Bayesian Analysis of the Weighted Marshall-Olkin Bivariate Exponential Model

Pages 255-280

https://doi.org/10.22054/jdsm.2026.85970.1068

Ali Sakhaei, Iman Makhdoom

Abstract The Weighted Marshall-Olkin Bivariate Exponential (WMOBE) distribution was first proposed by
Jamalizadeh and Kundu (2013), who examined its different characteristics and properties. Bayesian
estimation of the model parameters is carried out using both the squared error loss (SEL) function,
which is symmetric, and the linear-exponential (LINEX) loss function, which is asymmetric. These
estimators are derived under both informative and non-informative gamma priors. Given the complexity
of the four-parameters model, explicit analytical solutions for the Bayesian estimators are not attainable,
making it necessary to employ the Gibbs sampling procedure. Markov Chain Monte Carlo (MCMC)
methods are widely utilized to compute and implement these estimates. Furthermore, the convergence
behavior of the Markov chain toward a stationary distribution is carefully analyzed. Credible intervals,
particularly the highest posterior density (HPD) intervals for the unknown parameters, are also
constructed. To assess and compare the effectiveness of these estimation approaches, Monte Carlo simulations are performed. Finally, the methodology is applied to a real-world dataset for illustrative purposes.

Bayesian Computation Statistics

Health Monitoring of Industrial Equipment Based on a Single Output Parameter Using a Bayesian Two-Sample Test in Hilbert Space

Pages 281-300

https://doi.org/10.22054/jdsm.2026.86633.1072

Mohammad Mehdi Abdollahi, M. Bameni Moghadam

Abstract This study investigates the application of Bayesian two-sample testing in Hilbert space to monitor the health status of industrial equipment. The proposed method evaluates distributional differences between operational data samples to detect faults or anomalies. Unlike traditional multivariate techniques, our approach pro-vides higher sensitivity to subtle distributional shifts and supports visual insights into posterior distributions. Real-world experiments on industrial sensor outputs demonstrate the method’s effectiveness in early fault detection, reducing mainte-nance costs and downtime. This makes Bayesian two-sample testing in Hilbert space a powerful tool for predictive maintenance strategies.
This study investigates the application of Bayesian two-sample testing in Hilbert space to monitor the health status of industrial equipment. The proposed method evaluates distributional differences between operational data samples to detect faults or anomalies. Unlike traditional multivariate techniques, our approach pro-vides higher sensitivity to subtle distributional shifts and supports visual insights into posterior distributions. Real-world experiments on industrial sensor outputs demonstrate the method’s effectiveness in early fault detection, reducing mainte-nance costs and downtime. This makes Bayesian two-sample testing in Hilbert space a powerful tool for predictive maintenance strategies.

Computational Statistics

A goodness-of-fit test for progressively first-failure-censored data from a proportional hazard rate model

Articles in Press, Accepted Manuscript, Available Online from 23 July 2026

https://doi.org/10.22054/jdsm.2026.88330.1076

Mohammad Vali Ahmadi, Hamid Reza Moheghi

Abstract Progressive first-failure censoring schemes are potentially useful in practical applications where budget constraints exist or rapid testing is required. Moreover, several common sampling schemes such as first-failure censoring, progressive Type-II censoring, Type-II censoring, and complete sampling can be viewed as special cases of the progressive first-failure censoring scheme.
In this article, we propose a goodness-of-fit test statistic to assess whether a progressively first-failure-censored sample originates from a distribution belonging to the proportional hazard rate model. This model encompasses several well-known lifetime distributions, including the exponential, Rayleigh, Lomax, Burr Type-XII, Weibull, Gompertz, and Pareto distributions, among others.
We derive the null distribution of the proposed test statistic. Through Monte Carlo simulations, we evaluate the power of the proposed test under a range of different alternative distributions, assuming the null distribution to be exponential, Rayleigh, or Pareto. Finally, we present several numerical real-world examples to demonstrate the practical applicability of the proposed goodness-of-fit test. We also summarize the main findings and provide concluding remarks.

Bayesian Computation Statistics

Bayesian Analysis of Missing Data and Its Application in Categorical Data of TIMSS Test

Articles in Press, Accepted Manuscript, Available Online from 29 July 2026

https://doi.org/10.22054/jdsm.2026.92095.1087

Zahra Ghaffari Irdmousa

Abstract Abstract
In this article, unlike recently developed methods that can only distinguish associations between pairs of variables, Bayesian Multilevel Latent Class (BMLC) model for the multiple imputation of nested categorical data is considered which is flexible enough to automatically deal with complex interactions in the joint distribution of the variables to be estimated. After presenting the model, I run the model on a real data set, that is, mathematics test of TIMSS 2015 which has missing data in itself, and at the end the completed data for both levels (Level 1, students and Level 2, schools) are obtained. To assess the performance of the BMLC model, we compare it with list wise deletion (LD) with the help of R software. Conclusion: Results show that the BMLC model has reliability and is able to cover the Bayesian estimator with lower risk in this research.
Keywords: test of TIMSS, Bayesian analysis, missing data, categorical data, Bayesian Multilevel Latent Class model.

Bayesian Computation Statistics

Lognormal structured additive regression model for spatio-temporal data and its application to breast cancer data in Iran

Articles in Press, Accepted Manuscript, Available Online from 03 August 2026

https://doi.org/10.22054/jdsm.2026.88625.1079

Soudabe Sajjadipanah, Hasti Hashemi, Seyyed mahmoud Mirjalili

Abstract Medical data are typically recorded across geographic regions and over successive time points, giving rise to spatio-temporal structures that present important statistical modelling challenges. In this study, we propose a spatio-temporal regression model based on the lognormal distribution within the structured additive regression framework. The lognormal assumption accommodates the positive skewness and heteroscedasticity of the incidence rate data. Bayesian inference for the model parameters was conducted using the integrated nested Laplace approximation (INLA), which delivers substantial computational gains over Markov chain Monte Carlo methods while maintaining comparable accuracy. We systematically compared four types of spatio-temporal interactions (Types I–IV) to identify the most parsimonious and best-fitting dependence structure. To evaluate performance, we applied the proposed model and several competitors to breast cancer incidence data from all provinces of Iran over the period 2010–2019. The results demonstrated the clear superiority of the lognormal structured additive regression (LNSTAR) model, with the Type II (temporal-only) interaction providing the best fit. Beyond its methodological contribution, this analysis provides actionable risk stratification to inform targeted screening policies in Iran.

Keywords Cloud