V.M. Santi, K.A. Notodiputro, B. Sartono
Variable selection is an important topic in linear regression analysis. In practice, a large number of predictors usually are introduced at the initial stage of modeling to attenuate possible modeling biases. stepwise deletion and subset selection are usually used which can be computationally expensive and ignore stochastic errors in the variable selection process. In addition, the best subset selection of variables suffers from several disadvantages, the most severe of which is its a lack of stability. In this article, penalized likelihood approaches are proposed to handle these kinds of problems. The proposed methods select variables and estimate coefficients simultaneously. Some of penalty functions are used to produce sparse solutions. Based on the RMSE and Generalized Information Criterion (GIC) criteria, it was found that the factors affecting Indonesian mathematics scores, where LASSO produces 11 important variables for the model while SCAD has 6 variables which mean that the LASSO model is more complex than SCAD. The MCP produces a simpler model with 5 important variables but has excessive biassed. The results also showed that the SCAD penalty function had the best performance compared to LASSO, Ridge and MCP. Ridge penalty has a worst performance based on all criteria. © 2019 IOP Publishing Ltd.
Statistics Study Program, Faculty of Mathematics and Natural Sciences, State University of Jakarta, Indonesia; Department of Statistics, Faculty of Mathematics and Natural Sciences, Bogor Agricultural University, Bogor, Indonesia
Research at a Glance
Register to unlockTopics & SDG Alignment
Register to unlockCollaboration
Register to unlockAuthor Profile (Selected)
Register to unlockReferences Overview
Register to unlockJournal & Source
Register to unlockMetadata & Integrity
Register to unlock