Oregon State University

Predictive Analytics Platform for Student Success

published

Description

Universities collect extensive enrollment, academic, demographic, and engagement data, but advisors and administrators often lack timely, integrated insights into students who may be at academic risk.

In collaboration with the Oregon State University College of Engineering, students will design and develop a student success intelligence platform using historical institutional data. The platform will support prediction of student outcomes, analysis of risk factors, and data-driven decision-making. Students will work through an end-to-end workflow including data engineering, exploratory analysis, predictive modeling, evaluation, visualization, and deployment. Depending on team interests, they may also explore large language models, explainable AI, agentic AI, and interactive decision-support dashboards.

Problem statement

Advisors, faculty, and administrators need to identify students who may require support before academic difficulties become barriers to success. Although universities collect academic, demographic, enrollment, and engagement data each semester, these data are often fragmented across systems and analyzed manually or reactively.

This project will address that gap for the Oregon State University College of Engineering by developing a platform that turns historical institutional data into predictive and interpretable insights. The system will identify students at academic risk, examine factors associated with outcomes, and present information that can support timely intervention and operational decisions.

Objectives

By the end of the capstone, the team will deliver:

  • Prediction engine: A trained and evaluated machine learning model for predicting a student-success outcome, such as retention, persistence, or graduation risk, with documented accuracy, precision, recall, and ROC-AUC results.
  • Data pipeline: An automated, documented workflow for data cleaning, feature engineering, preprocessing, and schema definition.
  • Interactive dashboard: A dashboard showing risk predictions, success indicators, feature importance, and summary statistics or trends.
  • Explainability: Model interpretations and visual explanations using SHAP or feature-importance analysis.
  • Documentation: System architecture, modeling methodology, installation or deployment instructions, and user guidance.
  • Demonstration: An end-to-end demonstration using institutional data, accompanied by a presentation of the design, implementation, results, and future enhancements.

If core deliverables are complete, the team may explore LLM-generated summaries, advisor notifications, what-if analysis, cloud deployment with authentication, APIs, or additional forecasting.

Minimum qualifications

  • Senior standing in Computer Science, Computer Engineering, Data Science, or a related field.
  • Proficiency in at least one programming language, preferably Python.
  • Working knowledge of functions, classes, and data structures.
  • Basic statistics and data-analysis skills.
  • Ability to learn machine-learning concepts and collaborate on a multidisciplinary team.

Preferred qualifications

  • Experience with Python data science libraries (e.g., Pandas, NumPy, Scikit-learn).
  • Coursework or projects involving machine learning, artificial intelligence, or data mining.
  • Experience working with SQL and relational databases.
  • Familiarity with Git/GitHub for version control.
  • Experience developing web applications or dashboards (e.g., Streamlit, Flask, Power BI, Tableau).
  • Familiarity with cloud platforms (AWS, Azure, or GCP) or APIs.
  • Strong communication skills and an interest in solving real-world problems using AI.

Licensing / IP / NDA

This project requires an NDA or IP agreement.

Students must sign a Non-Disclosure Agreement (NDA) because the project processes sensitive data.