How do I cite a dataset using the Harvard citation style?
I need to cite a public dataset from a government repository in my research paper. What are the key elements required for a Harvard citation of a dataset, and how do I structure the reference entry and in-text citation correctly?
1 Answer
The short answer is that a Harvard‑style dataset reference should list the author (often a government agency), the year of publication, the title of the dataset (in italics), the version or edition if there is one, the publisher or repository, and a persistent identifier such as a DOI or URL, followed by the date you accessed it. In‑text you treat it like any other source: (Agency, year) for a narrative citation or (Agency, year) for a parenthetical one, adding a page‑like locator only if the dataset is divided into numbered parts. That basic skeleton keeps your citation clear and lets readers retrieve the exact data you used. The nuance comes in the details, because datasets don’t always fit the book‑or‑article template. First, the “author” can be a corporate author, a department, or even a specific project team; you should use the name that appears on the dataset’s landing page. If no individual author is listed, the agency name stands in. Second, the year is the year the dataset was released or most recently updated—not the year you accessed it. If the dataset has multiple versions, include the version number or date in parentheses after the title, and then give the retrieval date in the final element. Third, the title should be the full dataset title as it appears, italicised, and you can add a brief description in square brackets if the title is vague (e.g., [survey data]). A common misconception is that you can drop the URL because the DOI is “enough.” While a DOI is preferred, many government repositories still rely on stable URLs, and the Harvard guide usually asks for both. Another pitfall is treating the dataset like a journal article and adding volume or issue numbers—those fields simply don’t exist for most data collections, so leave them out. Trade‑offs appear when a dataset is part of a larger portal; you might need to cite the portal as the publisher and the specific collection as the title, which can feel redundant but ensures traceability. Edge cases include datasets that are continuously updated (e.g., a live API). In those cases, you note the version you used and the date you accessed it, because the content may have changed since. Imagine you’re writing about unemployment trends and you pull the “Annual Labour Force Survey, 2023” from the national statistics office. Your reference entry would look like this: National Statistics Office, 2023, Annual Labour Force Survey 2023 [dataset], version 2, National Statistics Repository, DOI 10.1234/nso.lfs2023.v2, accessed 12 May 2024. In the text you’d write something like, “The unemployment rate rose to 5.2 % in 2023 (National Statistics Office, 2023).” If you needed to point to a specific table within the dataset, you could add “Table 5” after the year in the citation. Always double‑check the latest Harvard guidelines from your institution, because small variations (e.g., placement of the URL or the use of “retrieved from”) can differ between style manuals. When in doubt, consult your school’s writing centre or the official Harvard referencing guide for the most up‑to‑date format. This approach will give you a clean, reproducible citation that satisfies most academic reviewers.
In-depth guide available
Mastering Harvard Style Citations for Complex Research Datasets →Have a similar question?
Ask the community →