Data Analyst Interview Questions for Freshers (2026): The Case Round Decides It

Updated August 2026

Data analyst is the most crowded entry point in Indian hiring right now, and the crowd is remarkably uniform. A large share of applicants hold the same three-month course certificate, list the same three tools, and submit the same two portfolio projects built on the same clean public datasets. If your application looks like that, nothing in it distinguishes you — and the good news buried in that sentence is how little it takes to look different, because the differentiators are cheap and almost nobody bothers with them.

The other thing worth knowing before you prepare is where candidates actually fail. It is rarely the SQL round. SQL is learnable, testable and widely practised, and our SQL questions page covers the query fundamentals this page deliberately does not repeat. What fails people is the case or scenario round — "this metric dropped last month, how would you investigate it?" — because it tests something no course teaches: whether you can reason about a business through data rather than execute queries against it. A candidate who writes decent SQL and cannot structure that investigation will lose to one whose SQL is adequate and whose thinking is organised.

So this page covers the rounds in order, gives you a working framework for the case question, adds the analyst-specific SQL that goes beyond joins, and is honest about the parts that get glossed over: whether you need Python, which BI tool to learn, why job titles in this space are so inconsistent, and what the money looks like. It is written for engineering graduates and equally for B.Com, B.Sc, BA and MBA candidates, who are hired into these roles in large numbers and are often better at the business half than the engineers are.

Frequently asked questions

What are the interview rounds for a fresher data analyst role?

Commonly: a resume and portfolio screen, where projects matter more than coursework; a SQL round, either live or as a take-home; a spreadsheet or Excel practical, often underestimated and frequently decisive; a discussion of a BI tool such as Power BI or Tableau, sometimes with a small build task; a case or scenario round on a business problem; and a final conversation covering communication, stakeholders and fit. Smaller companies compress this into two rounds and larger ones add a domain interview. Ask what the stages are at the start — it is a normal question and it lets you prepare for each round rather than meeting it cold.

A key metric dropped 20% last month. How would you investigate?

Structure beats speed here, and interviewers are listening for the order you work in. First, confirm the drop is real rather than an artefact: check whether tracking, a pipeline job or a definition changed, because a broken instrumentation event looks exactly like a business collapse. Second, check the denominator — a rate can fall because the numerator dropped or because the base grew. Third, segment before theorising: by date to find whether it was a cliff or a slide, then by geography, platform, channel, customer cohort and new versus returning. A cliff on one platform on one date is an engineering or release problem; a gradual decline across all segments is usually market or competitive. Fourth, consider external factors — seasonality, a festival, a competitor promotion, a pricing change. Then state a hypothesis and, importantly, say what data would confirm or kill it. Finish by saying what you would tell the stakeholder now and what you would come back with. Candidates who jump straight to "maybe marketing spend fell" fail this round even when they happen to be right.

How would you measure whether a new feature was successful?

Start by asking what the feature was supposed to change, because success is defined against intent rather than in general. Then name one primary metric that captures that intent, a small number of secondary metrics, and — the part that impresses — at least one guardrail metric that would tell you the feature succeeded by damaging something else. If a change increased order volume while raising cancellations and support contacts, that is not a success. Say how you would compare against a baseline, whether a controlled comparison is possible, and over what period, since early novelty effects mislead. Adding "and I would confirm the metric definition with the team before reporting" is worth a genuine point in most interviews.

What SQL do analyst interviews ask beyond joins and GROUP BY?

Three areas that come up constantly and are underprepared. Window functions — ROW_NUMBER, RANK and running totals with SUM OVER — for ranking within groups and month-on-month comparisons. Common table expressions, because analyst queries are layered and a CTE-structured answer reads as professional where a nested subquery reads as a struggle. And NULL behaviour, which is where most candidates actually get caught: aggregate functions skip NULLs, COUNT of a column differs from COUNT(*), a NULL in a NOT IN comparison silently returns nothing, and an outer join produces NULLs you then have to handle deliberately. Our SQL questions page covers the fundamentals; these are the ones that separate an analyst screen from a fresher SQL test.

You joined two tables and your row count went up. What happened?

A join fan-out — the join key is not unique on one side, so each left row matched several right rows and the result multiplied. This matters enormously because the query still runs and the number still looks plausible, which is exactly how inflated revenue figures reach a dashboard. Say how you would diagnose it: count the rows before and after, check uniqueness of the key with a GROUP BY and HAVING COUNT(*) greater than one, and decide whether you need to aggregate the many-side first or deduplicate before joining. Interviewers ask this because it is the most common way an analyst produces a confidently wrong number.

What is the difference between mean and median, and when does it matter?

The mean is the arithmetic average and the median is the middle value, and the difference matters whenever the distribution is skewed or has outliers. Salary is the standard example: a handful of very large values pull the mean well above what a typical person earns, so the median describes the typical case better. Say when you would use each — mean for symmetric data and where totals matter, median for skewed data such as income, order value, session duration or response times — and mention that reporting both, or a distribution, is often more honest than either alone. This question is really testing whether you would mislead a stakeholder without meaning to.

Explain correlation versus causation with an example from your own work.

Correlation means two things move together; causation means one produces the change in the other. The interviewer wants evidence that you would not present the first as the second. Give a concrete example: users who use a particular feature retain better, which may mean the feature causes retention, or simply that already-engaged users are the ones who find it. Then say what would move you toward a causal claim — a controlled experiment, a comparison of similar cohorts, a natural before-and-after where nothing else changed — and note the confounder explicitly. Naming a plausible confounder unprompted is the strongest answer available to you here.

What statistics do I actually need for a fresher analyst interview?

Less than most course syllabi suggest, and it needs to be genuinely understood rather than recited. Descriptive statistics — mean, median, mode, range, standard deviation and what a distribution looks like. Outliers: how to spot them and the judgement of when to exclude versus investigate them. Correlation and its limits. The intuition behind a controlled comparison or A/B test: why you need a control group, why sample size matters and why stopping early inflates false positives. If p-values come up, a conceptual answer is fine and a confidently wrong technical one is not. Depth in probability theory is rarely asked at this level, and depth in reasoning always is.

What is expected in the Excel or spreadsheet round?

More than candidates expect, and it is often where non-technical interviewers form their view of you. Lookups — VLOOKUP and increasingly XLOOKUP, plus INDEX and MATCH — and knowing why exact-match matters. Pivot tables built and rearranged quickly, since this is the fastest way to answer a question live. Cleaning functions: TRIM, text-to-columns, removing duplicates, handling dates stored as text, and IFERROR. Conditional aggregation with SUMIFS and COUNTIFS. Basic charting. Many roles, especially reporting and MIS positions, are genuinely Excel-first, so treating this round as beneath you is a common and expensive mistake.

Power BI or Tableau — which should I learn?

Learn one properly and say so honestly. The concepts transfer almost completely — connecting data, shaping it, relationships between tables, calculated measures, filters and interactivity, and designing a dashboard someone can read without you narrating it — so a candidate fluent in one picks up the other in a couple of weeks. Power BI is very widely used across Indian employers, particularly where the organisation already runs Microsoft tooling, which makes it a reasonable default. If a target employer names a tool in its posting, that settles it. What does not work is listing both after a tutorial in each; interviewers ask what you built, not what you have opened.

Do I need Python to get a data analyst job?

For most entry-level analyst roles in India, no — SQL, a strong spreadsheet ability and one BI tool cover the core, and plenty of good analysts work without Python for years. It helps in two situations: work that involves repetitive cleaning where a script saves hours, and roles that shade toward data science. So the honest sequencing is to get genuinely strong at SQL first, be quick in Excel, build real dashboards, and add pandas afterwards if your target roles ask for it. Do not delay applying until you have "finished Python" — that is a common way to spend six months not applying for jobs you were already eligible for.

What should my portfolio actually contain?

Two or three projects, each of which answers a question rather than demonstrating a tool. The structure that works: the question you set out to answer, where the data came from, what was wrong with the data and how you handled it, the analysis, and — the part almost everyone omits — what you would recommend someone do about the finding. Publish the SQL or notebook and a short written summary. Deliberately avoid the standard course datasets, because an interviewer who has seen the same Netflix or Titanic dashboard forty times cannot distinguish you from the other thirty-nine. Messy, awkward, real data from a public government portal, a local business, your college, or something you scraped yourself is worth more than a polished chart of a clean file.

How do I get data to work with if I have no job and no internship?

It is more available than people assume. Indian government open-data portals publish genuinely messy real datasets on transport, agriculture, health and census subjects. Your own college has data worth analysing — placement outcomes, attendance, library or canteen usage — and asking for it is a reasonable request that also demonstrates initiative. A small local business will often share sales data if you offer to produce something useful from it, and that becomes a project with a real stakeholder, which is the strongest kind. Public APIs and, where permitted, your own scraping produce data nobody else in the queue has. One awkward real dataset beats five clean tutorial ones.

How would you explain a technical finding to a non-technical stakeholder?

Lead with the answer, not the method. State what you found and what it means for their decision in one or two sentences, then give the supporting evidence, then the caveats and what you are unsure about. Avoid method narration — nobody needs to hear about your joins — and translate metrics into consequences: not "conversion fell 3.2%" alone but what that costs or implies. Interviewers often test this by asking you to explain one of your own portfolio projects to them as if they knew nothing, so rehearse a sixty-second version of each project out loud. Ending with a recommendation and a clear open question is what separates an analyst from a report generator.

Why are data analyst job titles so inconsistent?

Because the title describes the organisation more than the work. The same skills appear as data analyst, business analyst, MIS executive, reporting analyst, operations analyst and analytics associate, and the actual work ranges from building dashboards to running month-end reports to genuine investigation. This matters practically in two ways. When you apply, read the responsibilities rather than the title, since an MIS or reporting role can be a strong entry point that many candidates dismiss on the name alone. And in an interview, ask what the role actually does day to day — what questions it answers, who consumes the output, and which tools it uses — because the answer tells you whether you would be learning analysis or maintaining a spreadsheet.

What is the salary for a fresher data analyst in India?

Treat any figure as indicative, because it varies widely by city, by employer type and by how much of the role is genuine analysis versus routine reporting. Entry-level analyst and MIS-type roles in Indian metros are commonly quoted from a few lakh rupees a year upward, with product companies and specialised analytics teams at the higher end and reporting-heavy roles at the lower. As with any offer, ask for the structure rather than the headline — fixed versus variable, what releases any bonus, and what the appraisal cycle looks like. The larger financial point at this level is that the second job in this field usually pays substantially better than the first, so early roles are worth judging by what they let you learn and what data they let you touch.

Is a certificate course enough to get hired?

A certificate gets you literate; it does not get you hired, because a very large number of applicants hold the same one. Interviewers have learned to skip past course names to look at what you built, so the value of a course is the structure it gives your learning, not the credential at the end. If you have one, treat it as the beginning of the work: build projects on data nobody else in the queue is using, get quick in SQL and spreadsheets, and be able to defend every choice you made. If you are deciding whether to buy one, remember you can learn all of this from free material and spend the money on nothing at all.

I am from a B.Com, B.Sc or BA background. Am I at a disadvantage?

Less than you think, and in one respect you have an advantage. Analyst work is half business reasoning, and commerce, economics and statistics graduates frequently handle the case round, the metric-definition question and the stakeholder conversation better than engineering candidates who default to tool talk. What you must close is the technical half: genuinely strong SQL, quick spreadsheet work, one BI tool you have actually built with. Employers hire non-engineering graduates into these roles routinely. If you are also weighing a master's in analytics, our B.Com to business analytics guide works through when that is worth it and when it is an expensive way to delay applying.

Should I take a manual reporting or MIS job to get started?

Often yes, with one condition. These roles are real entry points, they exist in large numbers, and they get you working with actual business data and actual stakeholders — which is exactly what your portfolio and your case-round answers currently lack. The condition is that you keep building: if the job is entirely repetitive report generation, spend your own hours learning SQL depth and automating parts of your own work, which is both a portfolio project and a promotion case. What harms people is staying five years in a role that never expanded, not starting there.

What questions should I ask at the end of the interview?

Ask things that reveal what the work is: what questions the team is trying to answer this quarter, who consumes the analysis and what they do with it, what the data stack is and how clean the data actually is, and whether analysts here get to define questions or only to fulfil requests. Those answers tell you whether the role will develop you. Asking nothing reads as low interest, and asking only about salary and timing at this stage reads as low interest in the work — the compensation conversation belongs there too, just not alone.

I graduate in 2027. What should I do in the year before applying?

Four things, in order. Get genuinely strong in SQL, including window functions, CTEs and NULL behaviour, since that is the round you cannot argue your way through. Get quick in a spreadsheet, because a live practical rewards speed. Build two or three projects on data nobody else is using, each ending in a recommendation rather than a chart. And practise the case question out loud — take a metric, walk through instrumentation, denominator, segmentation and external factors until the structure is automatic. Add Python afterwards if your target roles ask for it, and start applying before you feel finished, because the interviews themselves teach you what to fix.

Don't just read Data analyst questions — get asked them

Phiny's AI interviews you on exactly these topics, follows up on weak answers, and tells you what a stronger answer looks like. Text interviews are free and unlimited.

Start a free AI mock interview

How to prepare

Where these questions get asked