Sunday, October 4, 2026
HomeData ScienceData Science Project Management – Agile, CRISP-DM and Delivering ML Projects

Data Science Project Management – Agile, CRISP-DM and Delivering ML Projects

Table of Content

📋 KEY INSIGHTS

  • The most common reason data science projects fail is not technical — it is misalignment between what the team built and what the business actually needed. Disciplined problem framing and stakeholder alignment at the start of a project prevents the majority of late-stage failures.
  • Agile methodology adapted for ML differs from software Agile in one critical way: ML work is inherently exploratory — you often do not know if a problem is solvable until you have spent weeks on it. Sprint planning must account for uncertainty by including explicit “spike” tasks for research.
  • The CRISP-DM framework (Cross-Industry Standard Process for Data Mining) provides a practical six-phase lifecycle — Business Understanding, Data Understanding, Data Preparation, Modelling, Evaluation, Deployment — that remains the most widely used project management framework for data science.
  • A data science project is not done when the model achieves target accuracy — it is done when the model is in production, monitored, and delivering measurable business value. Deployment and monitoring are as important as the modelling phase and require dedicated time allocation.
  • Technical debt in ML systems compounds faster than in software systems because it accumulates in four places simultaneously: code, data, models, and infrastructure. Regular “ML debt audits” are essential in any team that maintains multiple production models.
  • Stakeholder communication is the highest-leverage skill for career growth in data science — the ability to translate model outputs into business language and to manage expectations during uncertain exploratory phases consistently separates staff-level from senior-level data scientists.

Data science projects have a reputation for over-running timelines, under-delivering on business impact, and quietly dying in production after the initial launch excitement fades. This reputation is not undeserved — industry surveys consistently report that 80%+ of ML projects either never reach production or fail to deliver the anticipated business value. But these failures are not random. They follow predictable patterns that are well understood and largely preventable with disciplined project management. This guide covers the frameworks, processes, and practices that data science teams use to reliably deliver ML projects — from problem framing through deployment and monitoring — with particular attention to where Agile methodology works, where it must be adapted for the realities of ML work, and how to communicate progress and uncertainty to non-technical stakeholders.

The CRISP-DM Framework — A Practical ML Project Lifecycle

CRISP-DM (Cross-Industry Standard Process for Data Mining) was developed in 1996 by a consortium of industry practitioners and has remained the dominant project management framework for data science work for nearly three decades — a remarkable longevity in a field that otherwise discards frameworks rapidly. Its durability reflects the fact that it describes the actual structure of data science work rather than an idealised linear process. CRISP-DM is explicitly iterative: the arrows in its process diagram loop backwards as well as forward, reflecting the reality that findings in later phases regularly require revisiting earlier ones.

Phase 1 — Business Understanding: Define the problem in business terms before defining it in technical terms. What decision will this model inform? Who will use its outputs, and how? What does success look like — in business metrics, not model metrics? What is the cost of a false positive versus a false negative, and does that affect the choice of algorithm or the setting of decision thresholds? This phase should produce a written problem statement (one page) that is signed off by the business stakeholder. Without this document, scope creep is inevitable and the final model will be evaluated against shifting criteria.

Phase 2 — Data Understanding: Assess what data is available, what quality issues exist, and whether the available data is sufficient to solve the problem as defined. This phase regularly produces the most important finding of the entire project: the data is insufficient for the problem as stated, and either the problem must be redefined or additional data must be acquired. Discovering this in Phase 2 rather than Phase 4 (Modelling) saves weeks of wasted work. Deliverable: a data quality report and a data availability assessment against the problem requirements.

Phase 3 — Data Preparation: The most time-consuming phase in practice — typically 60–80% of total project time by most practitioners’ estimates. Cleaning, merging, transforming, and feature engineering to produce the training dataset. This phase should be version-controlled and reproducible: any transformation applied to the training data must also be applied to production data at inference time, making a clean, tested preprocessing pipeline essential.

Phases 4 and 5 — Modelling and Evaluation: Train and evaluate candidate models. Evaluation must be structured around the business problem: the right metric is not always accuracy or AUC. A fraud model with 99.9% precision but 40% recall is essentially useless — it misses 60% of fraud. These two phases iterate tightly and often loop back to Phase 3 when evaluation reveals data quality issues or missing features.

Phase 6 — Deployment: The model is serving predictions in production, integrated into the business process it was designed to support. This phase is underestimated by most teams — the engineering work of wrapping a model in a robust API, integrating with upstream data sources, implementing monitoring, and creating a rollback plan is often as large as the modelling work that preceded it.

PhaseKey DeliverableCommon FailureTime Allocation (typical)
Business UnderstandingWritten problem statement (signed off)Skipped — jumping straight to data5–10%
Data UnderstandingData quality reportOver-optimistic about data availability10–15%
Data PreparationVersioned, reproducible preprocessing pipelineNot version-controlled; not reproducible40–60%
ModellingTrained, evaluated candidate modelsWrong evaluation metric for business context10–20%
EvaluationBusiness-framed performance reportOnly reporting technical metrics5–10%
DeploymentProduction API + monitoring dashboardTreated as afterthought; no monitoring10–20%

Adapting Agile for Machine Learning Projects

person holding green paper
Photo by Hitesh Choudhary on Unsplash

Software Agile methodology assumes that requirements can be decomposed into small, independently deliverable stories with predictable effort. Machine learning work violates this assumption in a fundamental way: you genuinely do not know whether a predictive signal exists in your data until you look for it. A sprint goal of “achieve 85% accuracy on churn prediction by end of sprint” is either too easy (if the signal is strong) or impossible to commit to (if the signal is weak or the data is insufficient). Adapting Agile for ML requires acknowledging and managing this uncertainty explicitly.

The most effective adaptation is to separate research sprints from delivery sprints. Research sprints have a learning objective rather than a delivery objective: “Determine whether customer support interaction data contains a predictive signal for churn, and estimate what accuracy is achievable.” The deliverable is a written recommendation — yes, pursue this with delivery sprints; no, the signal is insufficient; or more data needed. Research sprints should be timeboxed strictly (typically 1–2 weeks) and treated as spikes in the backlog. They are not failures if the finding is negative — discovering early that a problem is not solvable saves far more time than discovering it after four delivery sprints of modelling work.

Delivery sprints, once the research phase establishes feasibility, can use standard Agile mechanics: stories, estimates, velocity tracking, and retrospectives. ML delivery stories are typically larger-grained than software stories — “build and evaluate the feature engineering pipeline” or “deploy model endpoint with latency monitoring” are appropriate sprint-sized stories. Attempting to decompose ML work into software-sized 2–4 hour stories usually produces artificial granularity that creates process overhead without adding planning value.

Sprint TypeGoalDeliverableDuration
Discovery SpikeIs this problem solvable with available data?Go/No-Go recommendation memo1–2 weeks
Data SprintBuild reproducible data pipelineVersioned dataset + EDA report2 weeks
Modelling SprintTrain and evaluate baseline + variantsModel comparison report2 weeks
Engineering SprintProductionise chosen modelAPI endpoint + unit tests2 weeks
Monitoring SprintBuild monitoring and alertingDashboard + alert rules1 week
Iteration SprintImprove model based on production feedbackUpdated model version deployed2 weeks

Stakeholder Communication and Managing Uncertainty

Communicating effectively with non-technical stakeholders is the professional skill that most directly determines a data scientist’s career trajectory. Technical competence creates a ceiling; communication ability determines how close you get to it. The core challenge is that data science work involves uncertainty that business stakeholders are not accustomed to in other domains: you might spend three weeks on data preparation only to find the data does not support the hypothesis, which is a productive finding but an uncomfortable one to communicate.

The most effective communication strategy is proactive expectation-setting at the project outset. Establish explicitly with stakeholders that (a) ML projects involve research phases whose output is a recommendation, not a guaranteed deliverable; (b) model performance will be expressed in terms of business outcomes, not just accuracy metrics; (c) the model will improve iteratively after deployment as more production data is collected; and (d) a baseline comparison will always be provided so that model improvement is measured against a realistic alternative, not against perfection. These four points, established upfront, prevent the most common stakeholder disappointments.

For ongoing communication, a weekly one-page status update is more effective than detailed technical reports. Structure it as: what we did this week (one sentence per task), what we found (one sentence, in plain language), what we are doing next week, and whether we are on track or have a blocker that requires attention. This format respects stakeholder time, keeps them informed, and surfaces blockers before they become delays. Reserve detailed technical discussions for the quarterly project reviews or for stakeholders who explicitly request depth.

✦ SUMMARIZE THIS ARTICLE WITH AI

The ML system design skills required for senior-level roles — including architecture, monitoring, and deployment — are covered in our ML System Design guide. MLOps practices that operationalise the deployment and monitoring phases of the project lifecycle are in our MLOps Interview Q&A. A/B testing methodology for evaluating model changes in production is in our A/B Testing guide. For career development and progression frameworks for data scientists, see our Data Science Career Guide 2026.

Leave feedback about this

  • Rating

Durgesh Kekare
Durgesh Kekarehttps://www.dataexpertise.in
Durgesh Kekare is a data science educator and founder of DataExpertise.in. With expertise in Python, machine learning, and analytics, he helps 10,000+ learners break into data careers.

Latest Posts

List of Categories