Data AnalyticsAnalyst career preparation and interviews

Build a data analyst portfolio around three defensible decisions

PK
Pankit Kumar
Sr. Data Scientist at Parexel (a Goldman Sachs–backed company) · 20 September 2026 · 3 min read
Technically reviewed by Ishaan Sharma
In this article (6 sections)

Build an analyst portfolio around decisions you can explain and defend: which number is trustworthy, what action the available evidence supports and whether a recurring report is ready for use. Three projects with clear reasoning and reproducible artifacts can demonstrate more than a long gallery of unexplained dashboards.

This is a proposed portfolio design, not a claim about a universal employer scoring system. Adapt it to the actual role and any stated application requirements.

Decision one: which number should the business use?

Use the synthetic commerce case to reconcile a completed-order total. The correct January amount is 104,000 paise across eight orders. Joining order headers directly to item rows produces 171,000 paise because some order amounts repeat.

Your decision is to use the source-grain calculation and repair the fan-out query. Show the metric contract, eligible IDs, wrong query, corrected query and reconciliation. Explain why SUM(DISTINCT amount) is not a valid repair when different orders can share an amount.

The source-row verification article provides a starting case. Reproduce it, then add a new boundary or mutation test of your own. Describe the work as an original extension of a teaching fixture rather than presenting the reference solution as unaided client work.

Decision two: what can incomplete evidence support?

Use the retail availability project. Two of six known snapshots are out of stock, while two of eight expected snapshots are unknown. The full-grid stockout share is bounded between 25% and 50%.

Your decision is to recover the missing states before classifying performance against a threshold inside that range. Do not claim that the data establishes lost sales, stockout duration or a causal business impact.

Show the expected product-day grid, coverage calculation, bounds and a short memo. This project demonstrates a different skill from the first: deciding what cannot yet be concluded, rather than merely correcting arithmetic.

Decision three: is the weekly report ready for review?

Use the weekly reporting assistant case. The reference pipeline handles an identical replay, includes a late-arriving eligible event and produces a 3,500-paise total across three selected events.

Your decision is whether the evidence packet is ready for human review. Demonstrate a correct run, a wrong-denominator candidate being rejected and a narrative change invalidating the packet identity. Keep the distinction between prepared, reviewed and distributed visible.

The reference uses authored deterministic narrative. If you add a live model, preserve actual run evidence and evaluate its claims; do not imply the supplied project already contains measured model performance.

Give each project a compact evidence page

Use the same navigable structure across projects without duplicating the analysis:

ElementWhat the reader should find
QuestionThe decision and who would use it
ContractPopulation, grain, period, unit and exclusions
ReproductionFiles, environment and one clear run command
ResultExpected values and a readable display
ChallengeA seeded failure, uncertainty or alternative explanation
DecisionSupported action, limitation and next evidence needed

Keep screenshots as supporting material. A screenshot of a correct total is less informative than code and data that reproduce it, especially when the reader wants to inspect a failure case.

Show ownership without inventing impact

State which components you wrote, which reference materials you used and which changes you made independently. A useful resume statement can describe implemented checks and reproducible results. It should not claim savings, revenue gains or client adoption that did not occur.

If the work is a team exercise, identify your contribution and explain the interfaces you depended on. Being able to explain a smaller owned component is preferable to implying ownership of an entire system you cannot discuss.

Exercise: write a one-sentence decision for each of your existing projects. If two projects make the same decision with different colors or tools, revise one to demonstrate a different analytical capability.

NeuraPath's Data Analytics with Generative AI course brings these skills together across SQL, BI, Python and verified AI work. Use its project practice to build evidence you can reproduce and discuss, while keeping employment outcomes separate from the quality of a teaching exercise.

Continue learning

This article is part of the Analyst career preparation and interviews sequence. Use the neighbouring tasks when you need the prerequisite or the next application.

PK
Pankit Kumar
Lead Instructor, NeuraPath Academy

Pankit Kumar has 10 years in Data Science & AI, building and shipping production systems in regulated pharma and clinical environments. He is a freelance trainer at Boston Institute of Analytics, AnalytixLabs and Scaler, and has taught this material to thousands of working professionals.

This article is part of our Data Analytics with Generative AI programme — 3–4 months. The full analyst stack — Excel, SQL, Power BI and Python pipelines — then a generative-AI layer you can prove is right.

Explore Data Analytics with Generative AI
Counselling is free · no obligation

Not sure which programme fits?

Tell us your background and we will map it to the right entry point — including saying so when a cheaper programme is the better fit. A counsellor replies within one working day.