Volume & Issue: Volume 3, Issue 2, June 2025 
Computational Statistics

A goodness-of-fit test for progressively first-failure-censored data from a proportional hazard rate model

Pages 1-26

https://doi.org/10.22054/jdsm.2026.88330.1076

Mohammad Vali Ahmadi, Hamid Reza Moheghi

Abstract Progressive first-failure censoring schemes are potentially useful in practical applications where budget constraints exist or rapid testing is required. Moreover, several common sampling schemes such as first-failure censoring, progressive Type-II censoring, Type-II censoring, and complete sampling can be viewed as special cases of the progressive first-failure censoring scheme.
In this article, we propose a goodness-of-fit test statistic to assess whether a progressively first-failure-censored sample originates from a distribution belonging to the proportional hazard rate model. This model encompasses several well-known lifetime distributions, including the exponential, Rayleigh, Lomax, Burr Type-XII, Weibull, Gompertz, and Pareto distributions, among others.
We derive the null distribution of the proposed test statistic. Through Monte Carlo simulations, we evaluate the power of the proposed test under a range of different alternative distributions, assuming the null distribution to be exponential, Rayleigh, or Pareto. Finally, we present several numerical real-world examples to demonstrate the practical applicability of the proposed goodness-of-fit test. We also summarize the main findings and provide concluding remarks.

Bayesian Computation Statistics

Bayesian Analysis of Missing Data and Its Application in Categorical Data of TIMSS Test

Pages 27-43

https://doi.org/10.22054/jdsm.2026.92095.1087

Zahra Ghaffari Irdmousa

Abstract Abstract
In this article, unlike recently developed methods that can only distinguish associations between pairs of variables, Bayesian Multilevel Latent Class (BMLC) model for the multiple imputation of nested categorical data is considered which is flexible enough to automatically deal with complex interactions in the joint distribution of the variables to be estimated. After presenting the model, I run the model on a real data set, that is, mathematics test of TIMSS 2015 which has missing data in itself, and at the end the completed data for both levels (Level 1, students and Level 2, schools) are obtained. To assess the performance of the BMLC model, we compare it with list wise deletion (LD) with the help of R software. Conclusion: Results show that the BMLC model has reliability and is able to cover the Bayesian estimator with lower risk in this research.
Keywords: test of TIMSS, Bayesian analysis, missing data, categorical data, Bayesian Multilevel Latent Class model.

Bayesian Computation Statistics

Lognormal structured additive regression model for spatio-temporal data and its application to breast cancer data in Iran

Pages 45-68

https://doi.org/10.22054/jdsm.2026.88625.1079

Soudabe Sajjadipanah, Hasti Hashemi, Seyyed mahmoud Mirjalili

Abstract Medical data are typically recorded across geographic regions and over successive time points, giving rise to spatio-temporal structures that present important statistical modelling challenges. In this study, we propose a spatio-temporal regression model based on the lognormal distribution within the structured additive regression framework. The lognormal assumption accommodates the positive skewness and heteroscedasticity of the incidence rate data. Bayesian inference for the model parameters was conducted using the integrated nested Laplace approximation (INLA), which delivers substantial computational gains over Markov chain Monte Carlo methods while maintaining comparable accuracy. We systematically compared four types of spatio-temporal interactions (Types I–IV) to identify the most parsimonious and best-fitting dependence structure. To evaluate performance, we applied the proposed model and several competitors to breast cancer incidence data from all provinces of Iran over the period 2010–2019. The results demonstrated the clear superiority of the lognormal structured additive regression (LNSTAR) model, with the Type II (temporal-only) interaction providing the best fit. Beyond its methodological contribution, this analysis provides actionable risk stratification to inform targeted screening policies in Iran.

Pattern Recognition

Fuzzy-AutoRepair: Expert-Validated Repair-Pattern-Conditioned Re-Ranking and Search-Cost Analysis for C/C++ Program Repair

Pages 69-85

https://doi.org/10.22054/jdsm.2026.94090.1099

Amir Ghanati, Saeed Parsa, Habib Izadkhah

Abstract Automated program repair (APR) for C/C++ programs is hindered by uncertain fault-localisation signals, large candidate spaces, costly test-based validation, and limited interpretability in learned or prompt-based repair decisions. This paper introduces Fuzzy-AutoRepair, an uncertainty-aware fuzzy learning-to-rank framework that prioritises explicit repair-pattern families prior to candidate generation. Unlike open-ended code generation, the proposed method represents each suspicious code region using 40 Boolean fault-local and contextual features and constrains candidate generation to 13 auditable AST-level repair-pattern families. The central novelty is a fuzzy inference layer that converts overlapping syntactic, data-dependence, control-context, API-call, enumeration, loop-header, and structural-risk evidence into linguistic repair suitability scores. These scores are fused with pairwise learning-to-rank relevance so that the exploration order is determined jointly by data-driven prediction and expert-readable fuzzy rationale, while type, scope, syntactic, and structural preconditions remain hard validity constraints. On 2,244 held-out pattern-selection samples and a 69-defect C/C++ benchmark, defect-level execution logs show that the balanced fuzzy-LTR configuration gives the strongest observed ranking and search-cost behaviour among the evaluated ordering variants. Compared with the LTR-only AutoRepair baseline, it improves Recall@1 from 0.45 to 0.50 (+11.1%), Recall@3 from 0.74 to 0.80 (+8.1%), and NDCG@5 from 0.78 to 0.84 (+7.7%). It also reduces generated candidates from 914 to 620 per defect (-32.2%), compilation attempts from 337 to 238 (-29.4%), and mean runtime from 49.38 to 41.70 minutes (-15.6%). The logs additionally show secondary observed differences in plausible repairs (33/69 versus 36/69) and correct-first repairs (14/69 versus 17/69).

Pattern Recognition

Non parametric Modeling of Fuzzy Time Dependent Data Based on Support Vector Machine

Pages 87-114

https://doi.org/10.22054/jdsm.2026.92040.1086

Mohammadghasem Akbari, Reza Zarei

Abstract This paper proposes a novel nonparametric support vector machine (SVM)-based approach for modeling and predicting fuzzy time-dependent data. Unlike traditional parametric methods such as ARIMA models and linear regression, which rely on rigid assumptions, the proposed framework leverages the flexibility of SVMs to capture complex, nonlinear temporal patterns in observations represented as fuzzy numbers. The nonparametric approach is chosen for its flexibility, robustness to model misspecification, and ability to accommodate irregular patterns and outliers common in fuzzy time series data. A practical example from software development quality monitoring motivates the need for this approach, where conventional SVM methods cannot directly handle fuzzy observations or uncertainty captured by spreads. To evaluate performance, the model is applied to both simulated fuzzy time series and a real-world dataset. A comprehensive set of evaluation criteria is employed: a similarity-based measure (MSM) for overall fuzzy similarity, root mean squared error (RMSE) for center accuracy, and average spread error (ASE) for spread accuracy. Empirical results demonstrate that the proposed method consistently outperforms existing methods across all three metrics. Furthermore, a comparison with a linear SVM model confirms the necessity of the nonlinear kernel for capturing complex temporal patterns, while the BDS test confirms the presence of nonlinear structure in the real dataset. The findings also indicate that the fitted models are robust and remain stable even in the presence of outliers. Overall, the results suggest that this SVM-based nonparametric framework offers a powerful, accurate, and robust tool for modeling fuzzy time-dependent data in various practical applications.

Neural Network

Predicting Temperament and Cardiac Strength from Persian Medicine Pulsology Using Artificial Neural Networks: An Algorithmic and Sensitivity Analysis

Pages 115-136

https://doi.org/10.22054/jdsm.2026.91444.1084

mohammad dehghandar, Mahdi Alizadeh Vaghasloo, Ghasem Ahmadi

Abstract Persian Medicine (PM) pulsology is a central diagnostic modality for assessing
an individual’s temperamental state and cardiac strength; however, its practical
application is inherently subjective and dependent on practitioner expertise. To
enhance objectivity and reproducibility in PM diagnostics, this study proposes
an artificial neural network (ANN) framework for the quantitative estimation of
four core temperamental qualities—warmness, coldness, wetness, and dryness—
together with cardiac strength, based solely on measurable pulse characteristics
Clinical data were collected from 69 individuals, comprising 11 pulse-derived features
and gender as inputs, with six corresponding diagnostic outputs. A multilayer
perceptron (MLP) architecture was trained and optimized using two learning
algorithms: Levenberg–Marquardt (LM) and Scaled Conjugate Gradient (SCG).
Comparative evaluation demonstrated the superiority of the SCG-trained network,
with an optimal configuration of 12 hidden neurons. This model achieved a test
accuracy of 96.43% and a low mean squared error (MSE) of 0.0139. Notably, cardiac
strength was predicted with 100% accuracy, while temperamental qualities
were classified with accuracies ranging from 71.43% to 92.86%. To enhance interpretability,
a sensitivity analysis was conducted, revealing pulse strength and pulse
frequency as the most influential predictors across multiple diagnostic outputs.
The proposed ANN-based system provides a stable and objective computational
surrogate for traditional PM pulsology. It offers practical utility for practitioner
training, supports diagnostic standardization, and establishes a methodological
foundation for future integrative and intelligent diagnostic platforms in Persian
Medicine.

Neural Network

A parametric activation function based on Wendland RBF

Pages 137-168

https://doi.org/10.22054/jdsm.2026.89529.1083

Majid Darehmiraki

Abstract This paper introduces a novel parametric activation function based on Wendland radial basis functions (RBFs) for deep neural networks. Wendland RBFs, known for their compact support, smoothness, and positive definiteness in approximation theory, are adapted to address limitations of traditional activation functions like ReLU, sigmoid, and tanh. The proposed enhanced Wendland activation combines a standard Wendland component with linear and exponential terms, offering tunable locality, improved gradient propagation, and enhanced stability during training. Theoretical analysis rigorously examines derivative behavior, smoothness, gradient flow, and saturation properties, demonstrating advantages over ReLU, GELU, and Swish including strictly positive gradients, $C^2$-continuity, near-unity gradient decay, and well-conditioned Jacobians. Empirical experiments on synthetic tasks and benchmark datasets confirm competitive performance, with comprehensive diagnostic analyses validating the theoretical claims. Results show that the Wendland-based activation achieves superior accuracy in certain scenarios, particularly in regression tasks, while maintaining computational efficiency. The study bridges classical RBF theory with modern deep learning, suggesting that Wendland activations can mitigate overfitting and improve generalization through localized, smooth transformations.

Machine Learning

BRAF: A Behavioral Response Analysis Framework for Student Patterns in the Era of Digital Assessment

Pages 169-199

https://doi.org/10.22054/jdsm.2026.93565.1096

Seyede Fatemeh Noorani, Fatemeh orooji, Amir Hushang TajFar

Abstract In the era of digital assessment, learners' behavioral data offer new opportunities for analyzing response patterns. This study presents the Behavioral Response Analysis Framework (BRAF), integrating behavioral data and machine learning to analyze responses in a randomized sequential online test. Data from 282 students completing a six-item test were analyzed using $t$-tests, Pearson correlation, Isolation Forest, one-way ANOVA, and Random Forest.
Isolation Forest identified 29 students (10.3\%) with anomalous response-position patterns, whose mean score was significantly lower than that of other students (3.76 vs. 4.70; $p < 0.001$). No significant difference was found between median-based PositionBias groups. Overall, 45.7\% of students changed their answers at least once, with answer-changing associated with a 0.61-point lower mean score ($p < 0.001$). Random Forest assigned greater importance to TotalChanges than PositionBias (53.3\% vs. 46.7\%), although predictive performance was limited.
The findings suggest that multivariate analysis may improve screening for unusual response-position patterns. Answer-changing was associated with lower performance, but causality cannot be established. BRAF is therefore best viewed as a screening framework for targeted human review rather than a diagnostic tool.

Mathematical Computing

Multi-Stage Risk Modeling Based on GARCH EVT Copula with Bootstrap Uncertainty Analysis

Pages 201-225

https://doi.org/10.22054/jdsm.2026.91882.1085

Azar Ghyasi

Abstract This paper develops a comprehensive framework for portfolio risk measurement integrating GARCH models for conditional volatility, Extreme Value Theory (EVT) for tail risk, and copula functions for dependence. The framework is validated using 1,000 simulated daily observations for two assets, with GARCH(1,1) models for volatility and a $t$-copula for dependence. EVT thresholds are selected using mean excess plots and stability analysis \citep{Coles:et:al:2001}.
Results indicate heavy tails, with kurtosis up to 8.74, volatility persistence exceeding 0.98, GPD shape parameters of 0.23 and 0.29, and a $t$-copula tail dependence coefficient of 0.374. Monte Carlo simulation yields one-day-ahead VaR and CVaR estimates of $-0.74\%$ and $-1.14\%$, respectively. Compared with historical simulation and alternative models, the proposed framework produces VaR estimates 46.67\% lower, while bootstrap confidence intervals demonstrate substantial uncertainty in extreme risk estimates. Kupiec backtesting does not reject the VaR forecasts ($p$-value = 0.5632), although the Christoffersen independence test indicates potential for improvement through more flexible volatility specifications.

Computational Statistics

A Mahalanobis-Distance-Based Copula Model Averaging Approach for Dependence Modelling with an Application to the Tehran Province Earthquake Catalogue

Pages 227-251

https://doi.org/10.22054/jdsm.2026.94057.1098

Sedigheh Shams, Mahnaz Rahmani

Abstract Copula models provide a flexible framework for modelling complex dependence structures between continuous variables. However, several competing copula families may provide similarly satisfactory fits, making inference based on a single selected model potentially unstable. To address this issue, a Mahalanobis-distance-based copula model-averaging approach is adopted to account for model uncertainty and obtain more stable estimates of conditional exceedance probabilities. Its performance is evaluated through simulation studies across a range of Kendall's $\tau$ values and illustrated using an earthquake catalogue from Tehran Province, Iran, to examine dependence between earthquake magnitude and inter-event time. The simulation results demonstrate stable estimation across different dependence structures. In the real-data application, appropriate marginal models are fitted before applying the copula framework, and the estimated conditional exceedance probabilities vary only slightly across elapsed-time thresholds. These findings suggest that elapsed time alone provides limited predictive information about the magnitude of subsequent earthquakes and highlight the value of incorporating model uncertainty into copula-based dependence modelling.

Computational Statistics

Short-Term Highway Traffic Volume Prediction: A Case Study of the I-94 Interstate in the Saint Paul–Minneapolis Metropolitan Area Using Hybrid Machine Learning Models and Multidimensional Temporal Features

Pages 253-286

https://doi.org/10.22054/jdsm.2026.93135.1092

Pooyan Fallah, Vadood Keramati

Abstract This study investigates short-term highway traffic-volume prediction as a key component of Intelligent Transportation Systems (ITS), supporting traffic management, infrastructure planning, and reductions in delays and emissions. The Metro Interstate Traffic Volume dataset (48,204 records; 9 variables) was preprocessed through data-quality checks and removal of unrealistic observations. To capture temporal and weather-related effects, systematic feature engineering incorporated day/night status, working days, week position, hourly bins, seasons, lagged traffic volume, and environmental variables such as temperature, precipitation, and weather conditions. Several models were evaluated, including linear and negative binomial regression, nonparametric methods, regression trees, and random forests, with hyperparameters optimized using cross-validation-based randomized and grid search. The selected model achieved strong test performance ($R^2=0.9753$, SMAPE $=9.26\%$), substantially outperforming the raw-feature linear baseline ($R^2\approx0.0484$). Traffic volume at time $t$ was predicted using contemporaneously available environmental variables and lagged traffic observations. The main contribution is a lag-aware, multidimensional feature-engineering and ensemble-learning framework for data-driven intelligent traffic management.

Machine Learning

Ridge Logistic Regression for Binary Water Quality Classification under Severe Multicollinearity: A Monte Carlo and Real River Study

Pages 287-302

https://doi.org/10.22054/jdsm.2026.94308.1101

Razieh Jafaraghaie

Abstract Classification of water quality into acceptable and unacceptable categories is essential for public health and environmental protection. However, standard maximum likelihood logistic regression (MLE) often becomes unstable in real-world hydrochemical datasets due to severe multicollinearity. This study evaluates the performance of ridge-penalized logistic regression compared to standard MLE through a Monte Carlo simulation and a real-world application using river water data from Buenos Aires, Argentina.

While MLE demonstrated marginally better in-sample predictive metrics in controlled simulations with moderate collinearity, it suffered from severe non-convergence and coefficient divergence in the real dataset, where predictors exhibited near-perfect correlation. In contrast, ridge regularization provided highly stable, interpretable coefficient estimates and robust performance by effectively managing the bias-variance tradeoff.
These findings emphasize that ridge regularization is not merely a theoretical improvement, but a practical necessity for binary water quality classification when hydrochemical predictors are highly correlated. I recommend routine use of cross-validated ridge penalization in environmental modeling, while explicitly acknowledging the limitations of in-sample performance metrics in small, imbalanced datasets.