Public data rarely arrives as a finished story. It usually appears as a spreadsheet, dashboard, survey, budget document, or collection of records that becomes useful only after you ask a precise question of it.
Start with a public-data source
Begin with a dataset that is relevant to your audience and specific enough to investigate. Useful sources include government open-data portals, census agencies, court records, school district reports, public budgets, election returns, environmental monitoring systems, health departments, and international organizations.
Before searching for a story, write down the basic identity of the data:
- Who collected it?
- What exactly does each row represent?
- Which dates or years does it cover?
- What geographic areas are included?
- What definitions and categories does it use?
- Is the dataset complete, provisional, or regularly revised?
- What important information is missing?
Do not treat a chart headline or search-result summary as a complete explanation. Open the documentation, codebook, methodology notes, or data dictionary. A column labeled “cases,” for example, might mean reported cases, confirmed cases, hospitalizations, or claims. Those differences can completely change the story.
A good starting point is often a dataset connected to a public decision. Look for records involving spending, access, enforcement, safety, housing, education, transportation, or environmental risk. These topics matter because changes in the numbers may affect identifiable communities and institutions.
Move from a topic to a question
A topic is broad: “housing costs,” “school attendance,” or “road safety.” A story question is narrower and testable: “Which neighborhoods experienced the largest increase in rent while household incomes remained flat between 2021 and 2025?”
Use this progression:
- Choose a public issue.
- Identify a measurable outcome.
- Select a comparison.
- Define a time period.
- Specify the people, places, or institutions affected.
- Add a reason the answer matters.
A practical question formula is:
How did [outcome] change for [group or place] during [time period], compared with [baseline or comparison], and what might explain the difference?
The phrase “what might explain” is important. Public data can reveal a pattern, but it usually cannot prove a cause by itself. Your first question should therefore separate the measurable finding from the later reporting needed to explain it.
For example, “Why did emergency-room visits increase?” may be too ambitious for one dataset. A stronger first question is: “Which age groups and districts saw the largest increase in emergency-room visits from 2022 to 2025, and did the increase exceed the regional average?” That question can be answered with data and followed by interviews, documents, and expert analysis.
Find the tension inside the numbers
Stories often emerge from a difference between what people expect and what the data shows. Look for tension, not merely large numbers.
Common forms of tension include:
- A city reports improvement overall, but some neighborhoods worsen.
- Funding increases while service levels decline.
- A policy is available statewide but used unevenly across districts.
- The average changes only slightly while the lowest-performing group changes dramatically.
- A problem appears to be shrinking, but reporting practices also changed.
- A program reaches many people, yet the people with the greatest need are least likely to use it.
- A public promise is measured differently from the result reported in official records.
Comparisons make tension visible. Compare one year with another, one location with another, or one population with a relevant benchmark. Avoid comparisons that sound dramatic but are unfair. A small rural county should not automatically be compared with a large metropolitan county without adjusting for population or explaining the difference in scale.
Pay attention to outliers, but do not assume that every outlier represents wrongdoing. An unusual value could reflect a real event, a reporting change, a small population, a classification error, or a technical problem. Treat it as a lead that requires verification.
Inspect and clean the dataset
Before calculating anything, inspect the structure of the data. Open the file in a spreadsheet or analysis program and check the column names, data types, duplicate rows, blank values, and unusual entries.
A basic inspection should answer these questions:
- Are the dates stored consistently?
- Are place names spelled the same way throughout?
- Are numbers stored as numbers, or as text with symbols and commas?
- Are totals mixed with individual records?
- Do rows represent people, events, dollars, or reporting units?
- Are there duplicate records?
- Are missing values represented by blanks, zeroes, “unknown,” or special codes?
- Did the collection method change during the period?
Never replace missing values with zero without understanding the meaning. A blank may mean that no event occurred, that the agency did not report, or that the value was withheld. Those are different conditions.
Create a short data note for yourself. Record the source URL or publication, download date, coverage period, units, known limitations, and any transformations you make. This note will make your reporting reproducible and help you explain the analysis to an editor or reader.
Choose the right comparison and calculation
The calculation should match the question. The most common mistake is reporting raw totals when rates or percentages are more appropriate.
| Story need | Useful measure | Main caution |
|---|---|---|
| Compare places of different sizes | Rate per population | Confirm the population denominator and year |
| Show change over time | Absolute and percentage change | Check whether the starting value is very small |
| Compare shares | Percentage of total | Define the total and avoid overlapping categories |
| Measure typical experience | Median | Explain what is excluded or grouped |
| Study distribution | Quartiles or range | Extreme values may distort the average |
| Compare program performance | Outcome per participant | Confirm that participants are counted consistently |
Use both counts and rates when possible. A location may have the highest number of incidents because it has the largest population, while a smaller location may have the highest rate. Neither result is automatically more important; they answer different questions.
For percentage change, use:
(new value − old value) ÷ old value × 100
Also calculate the absolute change. An increase from 2 to 4 is a 100 percent increase but only two additional events. Large percentages based on tiny numbers can mislead readers unless the underlying counts are shown.
When comparing populations, consider whether age, income, population size, or another characteristic affects the outcome. If the dataset allows it, calculate subgroup results rather than relying only on an overall average. Averages can hide unequal effects.
Turn a pattern into a reporting hypothesis
Once you find an interesting pattern, phrase it as a hypothesis rather than a conclusion. A hypothesis guides further reporting while leaving room for the evidence to change your mind.
Examples include:
- “The increase is concentrated in a small number of districts.”
- “The apparent improvement may be linked to a change in reporting rules.”
- “The program has expanded, but participation remains lower in lower-income areas.”
- “The average hides a much larger change among older residents.”
Then list possible explanations and the evidence needed to test each one. If a transit delay rate increased, possible explanations might include construction, staffing shortages, route changes, weather, or a new measurement system. For each explanation, identify records, interviews, meeting minutes, contracts, or expert sources that could confirm or challenge it.
This step prevents premature storytelling. The data may reveal where to look, but it does not give permission to state that a particular person, policy, or organization caused the result without supporting evidence.
Validate the finding before writing
Repeat the calculation using a second method or tool. For example, calculate a rate in a spreadsheet and then verify it with a short script, a separate worksheet, or a manual sample. You do not need complicated software to catch many errors.
Check the result against related sources:
- Compare the dataset with an official report or archived table.
- Ask the publishing agency whether definitions changed.
- Look for revisions or updated files.
- Check whether the same event appears in another public record.
- Review a sample of original records if they are available.
- Contact a subject-matter expert about reasonable ranges and interpretation.
Watch for survivorship bias, selection bias, denominator changes, and missing groups. If only hospitals that adopted a reporting system are included, the data may not represent every hospital. If only reported complaints are counted, the result describes reported complaints, not all incidents.
Small differences may not be meaningful. If two regions differ by a fraction of a percentage point, avoid describing them as substantially different without considering uncertainty, sample size, and data quality. For surveys, examine the margin of error. For administrative records, investigate whether the counts are complete and consistently defined.
Improve the question through interviews and documents
A strong data question becomes stronger when connected to people and decisions. Identify who is affected, who controls the relevant policy, and who can explain the process behind the numbers.
Useful sources may include:
- Residents or workers represented in the data
- Agency staff responsible for collection
- Officials who set the policy or budget
- Researchers who study the issue
- Advocates and community organizations
- Contractors or service providers
- Meeting minutes, audits, inspections, and court filings
Ask people about the pattern without presenting your interpretation as settled fact. Instead of asking, “Why did your program fail in these neighborhoods?” ask, “The records show lower participation in these neighborhoods. What factors could account for that difference?”
Request documentation, not just explanations. A press release may describe an initiative, while a budget, contract, inspection report, or performance dashboard may show how it operated in practice.
Troubleshoot common problems
If the dataset produces an implausible result, pause before publishing. Check whether you accidentally included totals, mixed monthly and annual records, or joined two tables using the wrong geographic code.
If names do not match, create a standardization table rather than changing values casually. Keep the original field and add a cleaned field. This preserves an audit trail.
If the data is incomplete, narrow the question. You may be able to report how recorded cases changed among participating agencies, but not claim that the entire population experienced the same trend.
If the pattern disappears after cleaning, that is useful information. The original result may have been caused by duplicates, missing denominators, or inconsistent categories. Document the correction and consider whether the data-quality issue itself deserves attention.
If no clear pattern appears, change the comparison rather than forcing a dramatic angle. Examine subgroups, longer time periods, geographic clusters, or differences between planned and actual results. A careful finding that the data does not support a common claim can be a valuable story question.
Write the final story question
A finished question should be specific, answerable, fair, and meaningful. It should identify the dataset or observable outcome, define the comparison, and point toward the people or decisions affected.
Use this checklist:
- Is the outcome measurable?
- Is the time period clear?
- Is the comparison fair?
- Are the units and denominators understood?
- Can the pattern be checked independently?
- Does the question avoid claiming a cause that has not been established?
- Does it lead to reporting beyond the spreadsheet?
- Would the answer matter to a real audience?
A strong final version might read: “Public transit delays rose in the county from 2023 to 2025, but the increase was concentrated on routes serving lower-income neighborhoods. Did staffing, maintenance schedules, or service reductions contribute to the uneven impact?”
That question gives you a measurable starting point, a meaningful comparison, and several avenues for verification. It turns public data from a collection of figures into a disciplined reporting plan.