Why Accurate Data Matters

Accurate data is the foundation of any successful project. It guides decisions, measures impact, and tells us whether we are truly helping the communities we serve. But what happens when the data we collect is not accurate? The data gets skewed.

Bad data doesn’t shout—it quietly leads to wrong decisions, wasted funding, and projects that look successful on paper but are failures on the ground. This article breaks down the common mistakes that skew a project’s data, with real examples from the field, and what you can do to prevent them—because donors don’t just fund activities, they fund the truth.

1. Errors on the People Collecting the Data

Skewed data often begins with the people carrying out and managing that data. When enumerators are poorly trained, they ask questions inconsistently. When supervisors pressure teams to deliver good results, data becomes inflated.

Example:

An enumerator asks questions differently depending on the respondent—sometimes using complicated language, sometimes simplifying it. Another enumerator, under pressure, adjusts a few numbers to make the project look better, adding figures that are not true and making the dataset unreliable.

One enumerator asks, “How many children are in your household?” and counts only those under 18. Another asks, “How many children live with you?” and only counts those who sleep there every night. The data cannot be compared.

How to avoid this mistake:

2. Sampling and Missing Data

This means surveying the wrong people and calling it the truth. Sampling errors happen when people collect data from groups that DO NOT represent the whole population you are trying to understand. This can occur if the sample is too small, not random, or excludes certain groups.

Example:

You are evaluating a maternal health program. You collect data from households near the clinic because they are easy to reach. However, families in remote villages who have less access to healthcare are missing from your sample. Your results show high levels of clinic visits, but this does not reflect the challenges faced by remote families. The data is overly positive and misleading.

Missing data occurs when respondents do not answer certain questions or when their responses are incomplete or lost. This happens, for example, when respondents refuse to answer or enumerators skip the question.

Example:

You are conducting a self-administered questionnaire in a community with low literacy rates. Many respondents skip questions because they do not understand the wording. Others guess or leave answers blank. Only literate respondents complete the survey properly. Your findings will reflect the views of literate people only, leaving a significant portion of the community out.

Example:

You are conducting a household income survey. Many wealthy respondents refuse to share their income, while poor households are more open. You exclude the non-respondents from your analysis. The average income you calculate is much lower than reality because the wealthy are missing.

How to avoid this:

  • Update the sampling list before every survey by cross-checking with community leaders.
  • Include all groups of people by deliberately including quotas of both genders and all community sections.
  • Reduce respondent burden by keeping the survey short and building trust around sensitive questions (e.g., about HIV & AIDS).
  • During data management, avoid missing data using digital tools that check and sync data daily, monitoring it in real time to catch problems early.

3. Inflated Numbers Through Fraud and Double Counting

This involves data collectors making up data to make the project look like it reached more people than it actually did. Fraud creates fake people, double counting creates fake reach—together they create fake impact. This is often done to please donors, meet targets, or hide poor performance.

In the end, this results in wasted resources, harm to communities because their real needs are not addressed, and continued funding for ineffective interventions.

Example:

A project needed 500 household surveys in 5 days. What happened was, 2 enumerators drove to one village and interviewed 10 households, then filed 490 forms at a guest house. The report shows 90% have clean water, but the truth is only 10 households were asked and the 3 remaining villages were never visited.

In double counting, the same person is counted many times for the same survey, either intentionally or not. The impact is the same—a project looks like it reached more people than it really did.

Example:

A vaccination campaign records each child’s vaccination. But some children are counted twice because their names appear on different clinic registers—once at a mobile clinic and once at a fixed clinic. The reported vaccination coverage is 120%, which is impossible and immediately flags the data as unreliable.

How to prevent this:

  • Create a beneficiary list with unique ID, phone number, name, and village of each respondent.
  • Emphasize honesty during training over good-looking results.
  • Do back checks by calling or visiting a sample of respondents to verify that the survey actually took place.
  • Have a third party or internal team review a sample of data for inconsistencies.
  • Validate entries in real time and compare data from different sources, like community records and program records.

4. Social Desirability Bias and Recall Errors

This occurs when respondents cannot accurately remember past events that happened weeks or months ago. They guess, estimate, or simply forget. This might also cause respondents to give answers they think are good rather than the truth.

Example:

A survey question asks, “Did your child receive a vitamin A supplement in the past 6 months?” Six months is a long time, and most mothers do not remember exactly. The data becomes inaccurate because they say yes or no while they are unsure.

How to prevent this:

  • Keep records of all periods short, and use health records if available.
  • Use indirect questioning and ensure anonymity.

Accurate data is more than a technical requirement—it is a commitment to honesty that runs through every level of a project. Whether it’s poorly trained enumerators, biased sampling, inflated numbers, or the limits of human memory, each of these four mistakes quietly distorts the truth. The good news is that each has clear, practical solutions. By training our teams consistently, sampling fairly, validating in real time, and building trust with respondents, we can protect the accuracy of our data. When we do, we not only make better decisions and use funding wisely—we honor the communities we serve and the donors who trust us. Because the truth, collected and handled with care, is the strongest foundation any project can build on.

Did you find this article useful? Support our work and download all templates.

About Elizabeth Banda Makolijah

Elizabeth Makolijah is a Capacity Building Officer at Tools4Dev, where she creates and delivers resources that help communities implement practical, sustainable solutions. She specializes in knowledge transfer, developing tools for NGOs, social enterprises, and multilateral organizations to accelerate impact. Elizabeth is committed to producing field-ready resources grounded in best practice.
Support our work ♡Download all templates
+