School of Statistics

Theses and dissertations submitted to the School of Statistics

Items in this Collection

Scoping reviews are a form of evidence synthesis that map the extent, range, and nature of research activity on a broad topic, providing a foundation for identifying knowledge gaps and informing future research and policy. Abstract screening, typically conducted by two or more independent reviewers, is one of the most resource-intensive phases of scoping reviews, motivating interest in automation through large language models (LLMs). This study evaluated LLMs as zero-shot binary classifiers for screening abstracts for a scoping review. The review consisted of 1,169 abstracts on LGBTIQ inclusion in the Philippines screened by three human experts. A prompt-based simulation design was implemented, with 2,300 records sampled with replacement per scenario across ten configurations: a benchmark, a model comparison across five models from an early snapshot of the GPT-5 family (GPT-5, GPT-5.1, GPT-5.2, GPT-5-mini, and GPT-5-nano), a reasoning effort comparison evaluating GPT-5.2 across three effort levels (low, medium, and high), and a prompt design variant evaluating whether removing the uncertain output category and enforcing a strictly binary classification constraint alters screening performance. Performance was evaluated using accuracy, precision, recall, and F1-score. Findings inform whether LLM-assisted screening achieves performance suitable for integration into evidence synthesis workflows across model architectures, reasoning mode configurations, and prompt output granularity.


Reviews are an avenue to better understand how air travel is perceived by travelers. They contain personal evaluations and experiences, good or bad, and creates a frame of input. This study explored the use of airline reviews from Skytrax's air travel review website to assess the public's perception towards air travel and identify talking points available in the compilation of reviews. A combination of text and image data was for unimodal, bimodal early fusion, and bimodal late fusion to predict review sentiments, while Latent Dirichlet Allocation and BERTopic were used to generate topic clusters. In summary, recently published reviews on airlines with image attachments highlighted never experiences in air travel. Dissatisfaction was mainly attributed to airport and flight timing and baggage handling and fees. Despite the decline in positive sentiment, premium cabin classes and certain airline regions fared comparatively better, while seating, amenities, and Asia-Pacific routes report positively experiences.


Digital banks have emerged as pivotal drivers of digital finance, addressing issues of financial exclusion through fully digital experiences and incentives like higher interest rates. However, the lack of physical branches poses challenges, especially in customer service, as users often turn to online platforms to voice concerns. Despite the establishment of a digital banking framework five years ago, studies focusing on customer satisfaction in digital banks remain limited. This research addresses this gap by employing static and dynamic topic modeling (DTM) techniques to analyze user reviews from digital banking applications, uncovering latent topics and identifying key feedback and concerns. The study created baseline static topic models using FASTopic and BERTopic. Results showed that BERTopic identified hidden topics in the data that are uniquely specific (e.g., Notification & In-App Ads, Games & Gambling Controls, and Dark Mode Requests). FASTopic, on the other hand, derived broad but balanced topics (e.g., Rewards, Perks & Everyday Utility, Biometric & Password Authentication, App Bugs & Technical Glitches, Transaction Errors, and Smooth Navigation). Although high-level insights were derived from BERTopic, the output of FASTopic was determined to better support the downstream objectives of the study, such as analyzing and characterizing the evolution of topics through time. The study leveraged the output of FASTopic to further isolate and analyze persistent issues of digital banks. The study uncovered trends and patterns for each entity covering both negative and positive reviews. Top concerns or issues were further tracked using the DTM and topic activity over time function. The study translated these trends and patterns into actionable insights per entity, explaining its importance on the overall customer journey. The study extended the analysis to include first quarter 2026 data in order to provide relevant recommendations. With the findings specified in this study, digital banks are encouraged to check internally if there are manifestations related to the concerns and issues raised. Institutions are also encouraged to recreate and enrich the study (i.e., with more data source and more sophisticated methods) to pro-actively monitor persisting concerns of their customers and, ultimately, implement timely corrective measures. Serving as a proof-of-concept in the industry, this study provides regulators and compliance teams with preliminary insights that may help identify potential regulatory gaps and evolving consumer needs. The findings may further inform service improvement initiatives and support efforts to mitigate market conduct and reputational risks within the Philippine digital banking sector.


Alternative proteins are critical to sustainable food systems, yet their adoption is moderated by complex sensory, psychological, and socioeconomic factors. This study utilized Bayesian Structural Equation Modeling (BSEM) to investigate the relationships between PROP sensitivity, preference, neophobia, and sociodemographic variables and attitude and perception toward alternative proteins among Filipino consumers (n = 125). Sensory evaluation was conducted via the generalized Labeled Magnitude Scale (gLMS), while psychological constructs were modeled as reflective latent variables. Results showed a a predominately medium-taster population (68.80%), with high stated preferences for alternative proteins based on health and sustainability, though conventional products maintained a sensory advantage in taste. The structural model revealed that Food Technology Neophobia was the most robust predictor of consumer response, positively correlating with negative attitudes (³ = 0.388) and negatively correlating with product perception (³ =
-0.285). Furthermore, higher income significantly mitigated negative attitudes (³ = - 0.311), while prior consumption experience was a requisite for favorable perception (³ = -0.205$). Notably, biological bitterness sensitivity (PROP) and general food neophobia (FN) demonstrated negligible effects on the endogenous constructs. These findings suggest that for Filipino consumers, the "neophobic" barrier to alternative proteins is specifically technological and socioeconomic rather than biological.
Consequently, transition strategies should prioritize demystifying food processing technologies and enhancing economic accessibility over sensory-only interventions.


The Fay-Herriot (FH) model is widely used in small area estimation for producing reliable domain-level estimates by combining direct survey data with auxiliary infor-mation through area-specific random effects. A well-known limitation of the EBLUP under this framework, however, is that it requires a direct survey estimate to compute the shrinkage factor. When certain areas have zero sample sizes, the EBLUP collapses to the synthetic estimator, effectively setting the area-specific random effect to zero, and loses the ability to reflect between-area heterogeneity. This study asks: when the EBLUP can no longer recover a random effect for a zero-sample area, can the available cluster or spatial structure in the data be used to construct a reasonable proxy?
This study proposes Random Effects Augmentation (REA), a two-stage approach that improves zero-sample estimation while retaining the classical FH structure. In the first stage, the standard FH model is fitted to sampled areas to obtain FH-shrunken random effects. In the second stage, proxy random effects for zero-sample areas are constructed by averaging these shrunken effects within a local reference group under a local exchangeability assumption. Two variants are developed: REA-ML, which uses covariate-defined groups, and REA-SP, which uses geographic proximity or adminis-trative boundaries.
The methods are evaluated in a simulation study with 100 areas across 13 scenarios varying structure type, structure strength, and zero-sample proportion. Results show that REA offers little gain when cluster or spatial structure is weak, provides meaningful im-provement over the synthetic estimator when structure is moderate, and is outperformed by dedicated multilevel or spatial models when structure is strong. The methods are also applied to Philippine rice production data for 81 provinces under 20%, 30%, and 40% zero-sample settings. Overall, REA provides a practical middle-ground approach for improving zero-sample prediction within the Fay-Herriot framework while clarifying when more specialized models should be preferred.