Abstract
Artificial Intelligence (AI)-powered features have rapidly proliferated across mobile apps in various domains, including productivity, education, entertainment, and creativity. However, how users perceive, evaluate, and critique these AI features remains largely unexplored, primarily due to the overwhelming volume of user feedback. In this work, we present the first large-scale, data-driven study of user feedback on AIpowered mobile apps, leveraging a curated dataset of 292 AIdriven apps across 14 categories with nearly one million AI-specific reviews from Google Play. We develop and validate a multi-stage analysis pipeline that begins with a human-labeled benchmark and systematically evaluates large language models (LLMs) and prompting strategies. Each stage, including review classification, aspect-sentiment extraction, and clustering, is validated for accuracy and consistency. Our pipeline enables scalable, highprecision analysis of user feedback, extracting over one million aspect-sentiment pairs clustered into 18 positive and 15 negative user topics. Our analysis reveals that users consistently focus on a narrow set of themes: positive comments emphasize productivity, reliability, and personalized assistance, while negative feedback highlights technical failures (e.g., scanning and recognition), pricing concerns, and limitations in language support. Our pipeline surfaces both satisfaction with one feature and frustration with another within the same review. These fine-grained, co-occurring sentiments are often missed by traditional approaches that treat positive and negative feedback in isolation or rely on coarsegrained analysis. To this end, our approach provides a more faithful reflection of the real-world user experiences with AIpowered apps. Category-aware analysis further uncovers both universal drivers of satisfaction and domain-specific frustrations. We expect our findings to advance understanding of user-centered development of AI-powered features and provide actionable guidance to software engineers and app developers who seek to align AI features of apps with user expectations.
| Original language | English |
|---|---|
| Title of host publication | Proceedings - 2025 IEEE International Conference on Big Data, BigData 2025 |
| Editors | Cheng-Zhong Xu, Leong Hou U, Xueqi Cheng, Jing Gao, Giuseppe Polese, Hong Mei, Paul Boniol, Michiaki Tatsubori, Chen Zhao, Dawei Zhou, Xiaohua Hu |
| Pages | 500-509 |
| Number of pages | 10 |
| Edition | 2025 |
| ISBN (Electronic) | 9798331594473 |
| DOIs | |
| State | Published - 2025 |
| Event | 2025 IEEE International Conference on Big Data, BigData 2025 - Macau, China Duration: Dec 8 2025 → Dec 11 2025 |
Conference
| Conference | 2025 IEEE International Conference on Big Data, BigData 2025 |
|---|---|
| Country/Territory | China |
| City | Macau |
| Period | 12/8/25 → 12/11/25 |
Bibliographical note
Publisher Copyright:© 2025 IEEE.
Funding
This material is based upon the work supported in part by the National Science Foundation (NSF) under Grants No. 2449694 and 2449695. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the authors and do not necessarily reflect the views of the NSF.
| Funders | Funder number |
|---|---|
| National Science Foundation Arctic Social Science Program | 2449695, 2449694 |
Keywords
- AI
- Mobile Applications
- User reviews
ASJC Scopus subject areas
- Artificial Intelligence
- Computer Networks and Communications
- Computer Science Applications
- Information Systems
- Information Systems and Management
- Safety, Risk, Reliability and Quality
Fingerprint
Dive into the research topics of 'Large-Scale Analysis of User Feedback on AI-Powered Mobile Apps'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver