Abstract
Population-level data for animals in commercial industries are rarely accessible to independent researchers. While animal-reliant industries may nominally disclose information via public registries, practical barriers such as dated interfaces, scattered sources, and reluctance to share data directly often prevent meaningful independent welfare oversight. This paper describes a novel methodology using agentic artificial intelligence (specifically, AI agents) to assemble a population-level dataset from publicly accessible registries, with greyhound racing under licensed regulation in the UK as a case study. A dataset of 31,028 individual greyhounds was assembled through cross-referencing of the Irish Coursing Club Stud Book, dog profiles, and race histories from the Greyhound Board of Great Britain (GBGB) racing database, www.Greyhound-Data.com, www.Greyhoundstats.co.uk, and the British Greyhound Stud Book, achieving 99.3% resolution of greyhound origin classifications. This provides census-level representation of the licensed greyhound racing population and data from 1,267,119 individual race entries between January 2022 and March 2026. AI agent-assisted collation under human oversight enabled data assembly at a scale and timeframe that would otherwise have been prohibitive. Dog identifiers were replaced with numerical codes in the deposited dataset, and human identifiers were removed. The dataset revealed that 85.1% of dogs racing in the UK in 2025 were Irish-bred, the median duration of active racing before exiting the industry was 30 starts across 11.9 months, and the rate at which dogs exit racing has not changed since 2022. Most inactive dogs have no traceable, publicly documented post-racing outcome. The GBGB holds unpublished data on retirement destinations, career-ending reasons, and individual-level injury and euthanasia figures—a structured visibility absence and an animal welfare concern. Where industries curate their own transparency, this AI agent-assisted collation method shows that independently assembled evidence can reveal what selective disclosure obscures. In doing so, it removes a practical barrier to the evidence-based accountability that social license and animal welfare improvement require. This approach, which includes ethical considerations and accountability, is replicable across other animal-reliant industries. Where population-level welfare data are not easily visible, this method has the potential to substantially expand the evidence base available to welfare scientists, regulators, policymakers, and the public.
1 Introduction
Oversight of animal wellbeing in commercial industries depends on access to data. Without population- and individual-level insights, animal welfare scientists cannot monitor changes, identify systemic risks, or properly inform regulators, policymakers, and the public (e.g., zoos and aquariums, DiVincenti et al., 2023). In many animal-reliant industries, from animal farming and biomedical research to zoos and animal racing, data about animal welfare are generated within industry systems (e.g., live export of Australian livestock, Fleming et al., 2020), and disclosed through limited channels, if at all. Many sectors nominally disclose welfare information through voluntary or mandatory public registries, annual reports, and certification schemes without necessarily providing honest visibility of animal experiences (Roszkowska-Menkes et al., 2024). Information may be aggregated at a level that masks local variation, reported in categories that shift between years (preventing longitudinal comparison), or held in technically accessible but difficult-to-use systems. Transparency, understood as a disclosure system, does not always produce the visibility required for effective oversight (Fung et al., 2007). Additionally, transparency that fails to demonstrate visibility of genuine safeguarding or animal welfare improvements may not sustain the social license it seeks to protect.
For many researchers, practical barriers to manually assembling industry data have proven prohibitive. Historically, accessing individual- or population-level information from fragmented public registries required navigating multiple platforms, reconciling inconsistent data formats, and manually processing voluminous archived records and publications (Tenopir et al., 2011). This exceeds what most researchers can reasonably achieve with limited time and resources. Moreover, industries may decline data-sharing when directly approached (Jakku et al., 2019; Liptovszky, 2024). Recent advances in agentic artificial intelligence (AI agents) offer a possible practical solution (Abou Ali et al., 2025; Helms Andersen et al., 2025). Unlike plain generative AI systems (e.g., ChatGPT generating text in response to human prompts), AI agents perform defined, repeatable tasks with greater autonomy under human direction (Mosqueira-Rey et al., 2023; Dwivedi et al., 2026). Applied to assembling data from public registries, they can compress months of repetitive work to days, making accurate, population-level datasets with individual-animal welfare insights achievable for non-commercial academic research, without industry cooperation.
Using agentic AI in research can impact humans or animals and thus requires ethical consideration. While regulatory and institutional frameworks are still emerging and evolving (Chandler et al., 2025; National Health and Medical Research Council, 2025a), some key ethical concepts and principles—including non-maleficence, privacy, beneficence, reproducibility, and accountability—can guide the responsible application of AI agents in research (Haibe-Kains et al., 2020; Khan et al., 2022). In AI ethics, non-maleficence requires avoiding or minimising AI-related harms (Floridi et al., 2018). AI systems may make powerful inferences and decisions that affect animal and human interests, as when AI tools reveal the hidden locations of animals (Coghlan and Parker, 2023) or sensitive private information about humans (Radanliev, 2025). The principle of beneficence requires promoting goods or benefits. Beneficence and non-maleficence entail ensuring that any risks of harm from AI tools are removed or ethically justified by likely benefits (Floridi et al., 2018). For example, some privacy risks to humans and animals in research may sometimes be justified by the promotion of significant public interests (National Health and Medical Research Council, 2025b).
Reproducibility in research studies allows result verification and error detection—vital features of accountable science (Sandve et al., 2013). In research employing AI, reproducibility requires documenting the workflow, from early steps to final outputs, so that others can verify and independently replicate the results (Sandve et al., 2013). Accountability involves taking appropriate responsibility for the effects of AI tools. Accountability can require human oversight (European Parliament & Council of the European Union, 2024) of AI outputs to ensure accuracy and fairness, mitigate human or animal harm, and ensure openness about AI risks and benefits in particular contexts (Radanliev, 2025). These ethical considerations for agentic AI in research informed this paper’s method.
This study describes a novel methodology: using AI agents to assemble a population-level dataset from publicly accessible registries to illuminate welfare insights in animal-reliant industries. Greyhounds that race in the UK offer an appropriate case study: the dog racing industry operates across multiple jurisdictions with inconsistent regulatory oversight (Scottish Animal Welfare Commission, 2023). Worldwide, welfare concerns such as racing injuries, poor living conditions, and ‘wastage’ are well-documented, but many details remain hidden from the public (Cobb et al., 2015a; Markwell et al., 2017; Thomas et al., 2017; Chang et al., 2022; Stevens et al., 2022). Dog racing’s social license to operate is eroding globally (Hampton et al., 2020; Cobb et al., 2023), and direct requests to share data for research are declined or ignored by the relevant governing bodies (Scottish Animal Welfare Commission, 2023; email communication) or deflected with responses like ‘most of what you want is already on the website’ (M. Bird, email communication, 12 January 2026). In this paper, we describe an AI agent methodology, present key findings, and examine what the process of assembly reveals about transparency, data governance, and the ethics of independent data assembly concerning animal-reliant industries. We demonstrate that information asymmetry in animal-reliant industries is not always an insurmountable barrier to welfare insights. This new method arrives at a time when animal welfare evidence is increasingly shaping legislation, and public interest in industry visibility is legitimate and urgent.
2 Background
2.1 Transparency, visibility, and animal welfare governance
Transparency has become a central mechanism in contemporary animal welfare governance (Hårstad, 2024). Regulators, industry bodies, and certifiers increasingly treat disclosure as evidence of accountability (Maciel and Bock, 2013). However, transparency and visibility are not the same thing. Transparency refers to the process or mechanism of disclosure (e.g., third-party certification audits or mandatory regulatory data releases, such as injury or fatality rates). Visibility is disclosure’s outcome: whether the information reaches stakeholders in an accessible and understandable form. Industries can obscure visibility by certifying processes rather than outcome accuracy (e.g., a certification audit may verify that a documentation process and records exist without evaluating if those records accurately represent the animal experience) or by presenting adverse welfare events in misleading formats (e.g., presenting data as rates per operational unit so that they appear smaller, rather than revealing animal-level risk). Reporting precision itself can also limit visibility; a rate published to one decimal point may appear stable over time, while the underlying value rises or falls. Transparency involving compromised disclosure fails to deliver honest visibility. Consequently, disclosure does not guarantee visibility, visibility does not guarantee accountability (the enforceable obligation to deliver animal welfare improvements or face stakeholder sanctions), and accountability does not always guarantee public trust (Fox, 2007). Transparency does affect organisational trust, but not automatically or linearly (Auger, 2014). Nonetheless, visibility and accountability are vital for authentic animal welfare assurance in animal-reliant industries.
The pathway from disclosure to animal welfare improvement involves several steps. Information must be disclosed to render animal experiences genuinely visible to relevant stakeholders. Visibility must trigger accountability mechanisms (e.g., consumer pressure, regulatory action, reputational consequences, and internal motivation) where warranted. Finally, human behaviour changes must promote animal wellbeing (Ormandy et al., 2019). Failures can occur throughout this chain (Fox, 2007). This paper is primarily concerned with the first link: the conditions under which industry disclosure does or does not produce genuine visibility for animal welfare scientists, regulators, and the public.
Stakeholders in animal-reliant industries operate differently at each stage, with various interests and capacities. Industries and certifiers control what is disclosed and how it is structured (Goodfellow, 2016; Lundmark Hedman et al., 2021). Regulators and researchers determine what can be retrieved and analysed (Holmberg and Ideland, 2012). The public determines what is trusted and what drives behaviour change, within the wider landscape of contemporary media interest and political appetite (Markwell et al., 2017; Mills et al., 2018; Ormandy et al., 2019). Where the disclosing party’s interests diverge from information recipients’ interests, disclosure systems are generally designed (deliberately or by default) to serve the former (Fung et al., 2007).
Several design features of disclosure systems can limit visibility without preventing nominal compliance. Data may be reported at aggregated national or industry-wide levels, obscuring local variation (Fox, 2007). Reporting categories may shift between years, preventing longitudinal comparison (Heald, 2006). Disclosure systems may be technically public but practically inaccessible due to interface design, data volume, or fragmented sources (Ananny and Crawford, 2018). Accreditation schemes may assess procedural compliance rather than welfare outcomes, producing certification that signals accountability without actually measuring it (Fung et al., 2007). These features allow transparency to serve as a tool of legitimacy rather than a driver of welfare improvement (Bekoff, 2022; Kline and Hooper, 2025; Mata and Marques, 2026) in a process called welfare washing: transparency that produces the appearance of accountability without measurable benefit to animals (Bjørkdahl and Syse, 2021; Cobb and Gaines, 2023).
Social license to operate—the implicit process by which communities grant or withhold approval for industry activities (Gunningham et al., 2004)—provides a complementary framework for understanding the consequences of transparency failure (Duncan et al., 2018; Breakey et al., 2025). In animal-reliant industries, public acceptance increasingly depends on welfare assurance (Hampton et al., 2020; Cobb et al., 2021; Cobb and Gaines, 2023). Where disclosure systems impede genuine visibility, the public cannot make informed judgments about social license to operate (Villeneuve et al., 2025). This matters most in industries where public support is already contested.
2.2 Greyhound racing in the UK as a case study
Greyhound racing provides a timely case study for examining transparency failure in animal-reliant industries. Its social license is actively eroding worldwide. As of April 2026, greyhound racing has been banned in New Zealand, Scotland, and Wales. Despite arguments that human entertainment justifies regulated use of animals to facilitate gambling (Campbell, 2023), political scrutiny of greyhound racing is growing, as public support declines in the remaining legal markets of Australia, the UK (England and Northern Ireland), Ireland, and one US state (West Virginia) (Hampton et al., 2020; Murphy et al., 2022; Eslake, 2025; Foxx and Carbajal, 2025; Parliament of Victoria, 2025; UK Parliament, 2025; Foxe, 2026; Mcmorrow, 2026). The industry depends on betting revenue and on political and public tolerance of self-regulation, both of which are sensitive to animal welfare concerns (Scottish Animal Welfare Commission, 2023; Parliament of Victoria, 2025; Foxe, 2026).
Data gaps are recognised as a barrier to meaningful oversight of greyhound racing. The Scottish Animal Welfare Commission’s report included evidence that the absence of whole-of-life data prevented welfare assessment across the full racing dog’s life cycle (Scottish Animal Welfare Commission, 2023). The joint submission to the Scottish Animal Welfare Commission by RSPCA and Dogs Trust explicitly called for robust traceability, suggesting that a microchip-based database would provide transparency accessible across all life stages, and noting that without it, the welfare experience of individual greyhounds would remain largely invisible (Scottish Animal Welfare Commission, 2023). In this study, direct requests for research data sharing (six requests, via email between November 2025 and January 2026) to the Greyhound Board of Great Britain (GBGB) and the Irish Coursing Club (ICC) were declined or went unanswered. The data described here were subsequently assembled by the corresponding author, without industry cooperation, using only publicly accessible sources. The process of data assembly and the limits encountered offer insights about the state of transparency as a disclosure system in UK greyhound racing.
3 Methodology
3.1 Data sources
The focal period of this study was 1 January 2022 to 31 March 2026 inclusive. The start date was chosen to begin after all pandemic-related disruptions to UK greyhound racing through 2020–2021. This period also covers the May 2022 introduction of the Greyhound Board of Great Britain’s ‘A Good Life for Every Greyhound’ welfare strategy (Greyhound Board of Great Britain, 2022). Starting with a calendar date, rather than individual dogs’ racing start dates, allowed us to observe dogs at various career stages. At times, we did pay attention to the assembled data based on when individual dogs commenced racing. The end of the focal period reflected the date of data assembly.
Five publicly available online sources were used. The ICC Stud Book (https://greyhoundregistration.ie/Webstudbook/webpages/Default.aspx) is the primary registry for greyhounds bred in Ireland. It records registrations of individual dogs, including whelp and pedigree information, and transfer records, including explicitly noting when dogs are transferred from the Republic of Ireland to race in the UK, under the regulation of the GBGB. The GBGB racing database (www.gbgb.org.uk/racing/results/) holds individual dog profiles (Greyhound’s form search function) and race histories for greyhounds registered to race in Great Britain, with each dog assigned a unique sequential identifier publicly visible in the URL of the dog’s profile page. The British Greyhound Stud Book (GSB; www.greyhoundstudbook.co.uk) is the official registry for UK-bred greyhounds, recording registered litters by their sire/dam and whelp date. Throughout this study’s focus period (greyhound racing in the UK between 2022 and 2026), GSB processing of registrations, transfers, and earmarking was subject to documented delays: in August 2025, the GBGB announced plans to develop its own parallel registration system for British-bred greyhounds, citing a backlog of work and ongoing problems with the receipt of registration documentation (Greyhound Board of Great Britain, 2005). Data assembly for this study followed that announcement, with the understanding that the information yielded from the GSB may be incomplete; as of June 2026, no new registration system had yet been publicly launched by the GBGB. Greyhound-Data.com (www.greyhound-data.com) is an international pedigree and race records database that includes country of origin classifications, dam/sire and whelp date pedigree information, and post-racing adopted owner records voluntarily reported by new owners; access requires free user registration. The website Greyhoundstats.co.uk (www.greyhoundstats.co.uk) is an independently compiled database of UK race results and trainer assignments, drawing on official results sources; in this study, it was retrieved in parallel with the GBGB race history information and used as a side-by-side cross-check on race counts, first-race dates, and trainer attribution.
All these sources are publicly available; four are accessible without registration, and the fifth (Greyhound-Data) is accessible upon free, open registration. Retrieving population-level information from any single source requires navigating interfaces designed for individual record lookup rather than systematic data extraction. Combining information from all five sources substantially compounds this complexity. These traditional barriers to data access, plus the declined data-sharing requests (refer to Section 2.2), established the methodological rationale for the collaborative AI agent-mode workflow described below.
3.2 AI agent-assisted data assembly
AI systems capable of operating in agent mode for systematic data retrieval and assembly tasks under human supervision are available from several major providers, including OpenAI, Google, and Microsoft. All stages of data extraction, parsing, cross-referencing, and structuring were performed by the corresponding author, directing an AI agent workflow powered by Claude (Anthropic, Inc.). Claude Cowork is a large language model with tool access (including file management, web retrieval, and programmatic data processing) to execute defined, repeatable data tasks under continuous human supervision. Agent-mode operation is functionally distinct from the popular conversational use of generative AI (which produces text from prompts) and involves greater task autonomy. The agent was directed to perform defined retrieval, parsing, and structuring tasks, with each action observable and subject to review by the corresponding author. The primary output was structured tabular data assembled from multiple sources in a new way, rather than generated text or replication of existing registries.
Our deployed agents performed deterministic computation (counting, date arithmetic, and gap distributions) without statistical modelling, interpreting findings, or generating our analytical conclusions. Claude Cowork was operated under Anthropic’s consumer Pro plan terms of service. The data usage was configured to opt out of model training, and session logs were subject to automatic deletion after 30 days, consistent with Anthropic’s data retention policy. All data retrieved by agents was written to local storage on a password-protected device under the corresponding author’s sole control.
The workflow was anchored in four operational constraints designed to prevent the agent from generating, predicting, or inferring information not observed in the source records, consistent with published frameworks guiding responsible human oversight of reproducible AI-assisted research workflows (Sandve et al., 2013; European Parliament & Council of the European Union, 2024). First, the agent was directed to record only fields that were locatable on identifiable source pages, with the source URL and retrieval timestamp captured for each record. Second, missing or absent fields were preserved as gaps rather than imputed, completed by inference, or filled by analogy to other records. Third, key fields used for analysis, including the country of origin and race counts, were cross-verified across multiple source registries where their data records overlapped, with discrepancies flagged for manual author review rather than silently resolved. Finally, the workflow was staged: outputs from each stage were reviewed in detail by the corresponding author before the next stage commenced.
The data assembly pipeline proceeded in four phases, each containing one or more discrete processing steps with structured logs recording inputs, outputs, and file hashes at each step documented in the workflow (see Figure 1). In Phase 1, all GBGB-sanctioned meetings between 1 January 2022 and 31 March 2026 (inclusive) were identified from the GBGB’s publicly accessible online racing records, with information about each race extracted into a structured record. Phase 2 involved identifying individual dogs and career metrics (including race counts, career duration, inter-race gaps, adverse racing events, and current racing status), which were computed deterministically from the race-level data. Each dog’s country of origin was resolved in Phase 3 by strategic consideration of public breeding registries in a fixed priority order (ICC, then GSB, and then cross-checked against Greyhound-Data), accepting the classification from the first registry to return a confirmed match (Section 3.3). In Phase 4, race count data were independently cross-validated against a separate public database (Greyhoundstats), and a random sample (n = 50) was verified by the corresponding author against the source websites. This multi-source, cross-referenced pipeline design is comparable to capture–recapture methods used in veterinary epidemiological surveillance, where independent detection systems are combined to identify cases missed by any single source (Vergne et al., 2015). Information recorded at each stage of the linkage pathway was consistent with published guidance on transparent reporting of dataset assembly using linked, routinely collected data (Benchimol et al., 2015; Gilbert et al., 2018).
Figure 1
The collaborative agent-mode retrieval was conducted by the corresponding author (located in Australia) for non-commercial academic research purposes, consistent with published frameworks for the legal, ethical, and scientific governance of web-based data retrieval in research (Brown et al., 2025). Each request by the agent was for a single record page that any human visitor could load through the source website’s own public interface, using the same standard public URLs that the source operators themselves publish. On the GBGB platform, this included direct profile URLs of the form …/greyhound-profile/?greyhoundId=N—a URL pattern that the GBGB exposes to every visitor of every dog profile. Free user registration on Greyhound-Data was completed as required by that site. No technical access controls were circumvented, no rate limits or anti-automation measures were bypassed, and no protected, paywalled or login-only data were accessed beyond the free Greyhound-Data registration. The dataset that accompanies this manuscript is an aggregated, pseudonymised and analytically transformed research output, not a copy or republication of any source registry’s underlying database.
The corresponding author monitored and checked the agents’ activity throughout, including visually verifying samples of extracted records against source pages, reviewing resolved origin classifications against primary source entries, and confirming race history data against independent records. Where discrepancies were identified, records were flagged for manual review and either resolved against a second source or excluded. The collaborative agent-mode approach reduced the time of data assembly from several months of manual work to under a week. This time-saving feature is a major methodological contribution. It makes population-level dataset assembly from publicly available information practically achievable for individual researchers.
3.3 Data governance and de-identification
The source registries contain names of individual people (trainers, owners, and breeders) in their public-facing form. During cross-registry, the agent relied exclusively on greyhound pedigree information (name, dam, sire, and whelp date) rather than any associated human identifiers. This enabled a structural separation between data linkage and data analysis. Cross-referencing records across registries requires real identifiers (e.g., dog names, sire and dam names, and associated human names also show on these sites for roles including breeder, owner, trainer, etc.), but the researcher need not encounter them. The agent was directed to generate pseudonymised identifiers at first contact with each data source and write only those identifiers to the working dataset, thereby immediately de-identifying the researcher’s analytical files. Individual greyhound names (which, along with pedigree information, served as unique identifiers linked to the publicly searchable registry profiles) were replaced with sequential codes (D00001, D00002, etc.) at the time of extraction to reduce re-identification risk in research publication (Rodriguez et al., 2025) and for the deposited dataset. An encrypted mapping key (linking codes to source identifiers) was retained securely by the corresponding author for verification purposes but was not included in any shared output. This is an agent-assisted data assembly workflow property that distinguishes the method from manual data assembly, in which the researcher necessarily handles identifiable information.
Where the source registries listed names of trainers, owners, and breeders (considered personal data under the UK General Data Protection Regulation (UK GDPR) and Data Protection Act 2018), the AI-mediated pseudonymisation described above was the principal safeguard. This accords with guidelines for compliance under the scientific-research provision of those instruments (UK GDPR, Article 89; Rhahla et al., 2021). The University of Melbourne’s Human Research Ethics and Integrity Office determined that this study fell outside the scope of human research ethics review because the data were publicly available and de-identified prior to researcher engagement. The University’s Animal Ethics Committee similarly determined that the study was exempt from animal ethics review, as no animals were directly involved. The agent-assisted assembly workflow contributed to these determinations by enabling de-identification at the point of data assembly, rather than as a post-hoc processing step.
Although all the data are publicly accessible and ethics committee review was waived, steps were taken within the study’s aims and timeframe to preserve privacy by preventing identification of individuals during and after the research. While computational and AI techniques can also reveal information about individuals through the assembly and processing of data, even when that data are already in the public domain, AI techniques can also enable data to be de-identified during assembly and hidden from researchers. Later, we discuss these and other ethical features of using AI agents in these contexts (see Section 5.2).
This governance approach is consistent with the data privacy provisions of the RECORD statement for studies using routinely collected data (Benchimol et al., 2015) (Vergne et al., 2015). The deposited dataset was structured to align with FAIR principles for scientific data management (Rodriguez et al., 2025), and the methodology described in Sections 3.1 and 3.2 provides sufficient detail for independent replication from the same public sources. Findings reported at the individual level in this paper do not use names. Where individual-level variation in greyhound welfare outcomes is described, this is presented as illustrative of system-wide patterns rather than as a characterisation of named individuals.
3.4 Data analysis
The assembled data were analysed using R version 4.6.0 (R Core Team, 2026). Descriptive summaries of dog demographics (Table 1), country of origin (Table 2), adverse events (Table 3), and dataset coverage (Table 4) were reported as counts and percentages of total dogs or races. Continuous variables were summarised with median and interquartile range (IQR) reported; mean and range were also reported where they aid interpretation. Descriptive analyses, including IQR and quartile calculations, were performed, with additional statistical tests conducted where formal inference was warranted. Between-trainer differences in median retirement age were tested using the Kruskal–Wallis test (Kruskal and Wallis, 1952) with epsilon-squared reported as the effect size (Tomczak and Tomczak, 2014). Year-on-year trend in the GBGB published fatality rate was tested using the Cochran–Armitage test (Armitage, 1955) for linear trend in proportions (Agresti, 2018) under a pre-specified one-sided directional hypothesis.
Table 1
| Characteristic | n | % of total |
|---|---|---|
| Sex | ||
| Female | 15,497 | 49.9 |
| Male | 15,529 | 50.0 |
| Not recorded | 2 | <0.1 |
| Coat colour | ||
| Black | 17,572 | 56.6 |
| Brindle (incl. combinations) | 4,731 | 15.3 |
| Blue/blue brindle (incl. combinations) | 4,376 | 14.1 |
| Black and white | 3,041 | 9.8 |
| Fawn (incl. combinations) | 1,287 | 4.2 |
| Other or not recorded | 21 | 0.1 |
| Birth year | ||
| 2018 or earlier | 3,667 | 11.8 |
| 2019 | 4,834 | 15.6 |
| 2020 | 5,521 | 17.8 |
| 2021 | 5,780 | 18.6 |
| 2022 | 5,115 | 16.5 |
| 2023 | 4,053 | 13.1 |
| 2024 | 2,056 | 6.6 |
| Not recorded | 2 | <0.1 |
Demographic characteristics of the assembled dog population (n = 31,028).
Birth month was recorded to month–year precision (e.g., Feb 2024) in GBGB records. Coat colour groupings were derived from the recorded colour in GBGB race records.
Table 2
| Year | GBGB registrations | Assembled dataset first-runners | % Assembled dataset first-runners as proportion of GBGB reported registrations | |
|---|---|---|---|---|
| 2022 | 6,307 | 6,183* | 98.0% | |
| Irish | 5,497 (87.2%) | Irish | 5,407 (87.4%) | 98.4% |
| British | 810 (12.8%) | British | 761 (12.3%) | 93.9% |
| 2023 | 5,899 | 5,560 | 94.3% | |
| Irish | 4,861 (82.4%) | Irish | 4,697 (84.5%) | 96.6% |
| British | 1,038 (17.6%) | British | 834 (15.0%) | 80.3% |
| 2024 | 5,133 | 4,917 | 95.8% | |
| Irish | 4,338 (84.5%) | Irish | 4,107 (83.5%) | 94.7% |
| British | 795 (15.5%) | British | 762 (15.5%) | 95.8% |
| 2025 | Not yet published | 4,624 | – | |
| – | Irish | 3,905 (84.5%) | – | |
| – | British | 686 (14.8%) | – |
Assembled dataset compared to Greyhound Board of Great Britain published registration, by year and country of origin.
*Reflects first-runners only in each year from the assembled dataset (i.e., excludes dogs that commenced racing prior to 1 January 2022). Dogs classified as unknown origin or from other countries are not reported here.
GBGB: Greyhound Board of Great Britain.
Table 3
| Racing intensity quartile (starts per month) | Place rate Q1 (≤43.2%) | Place rate Q2 (43.2%–51.4%) | Place rate Q3 (51.4%–59.6%) | Place rate Q4 (>59.6%) | Row median |
|---|---|---|---|---|---|
| Q1 (≤2.3) | 12.0 months (464) | 19.1 months (298) | 19.2 months (285) | 16.3 months (413) | 16.1 months (1,460) |
| Q2 (2.3-2.9) | 14.5 months (346) | 20.3 months (387) | 23.1 months (376) | 20.1 months (348) | 19.8 months (1,457) |
| Q3 (≤2.9-3.5) | 12.6 months (290) | 20.3 months (399) | 22.3 months (417) | 19.4 months (352) | 19.7 months (1,458) |
| Q4 (>3.5) | 3.9 months (363) | 18.2 months (374) | 20.6 months (376) | 11.3 months (344) | 13.3 months (1,457) |
| Column median | 10.9 months (1,463) | 19.8 months (1,458) | 21.6 months (1,454) | 17.0 months (1,457) | 17.4 months (5,832) |
Median career duration (months) by racing intensity quartile and place rate quartile, in the 2022 entry cohort (n = 5,832 inactive dogs).
Table 4
| Track | Race starts (n) | Trouble events (n) | Trouble rate (%) | Adverse events (n) | Adverse rate (%) | Adverse-to-trouble ratio (%) |
|---|---|---|---|---|---|---|
| Central Park | 62,808 | 22,212 | 35.4 | 647 | 1.03 | 2.91 |
| Star Pelaw | 17,983 | 6,784 | 37.7 | 170 | 0.95 | 2.51 |
| Oxford | 55,802 | 17,163 | 30.8 | 392 | 0.70 | 2.28 |
| Valley | 24,527 | 10,432 | 42.5 | 231 | 0.94 | 2.21 |
| Monmore | 95,283 | 41,960 | 44.0 | 903 | 0.95 | 2.15 |
| Towcester | 76,912 | 23,699 | 30.8 | 504 | 0.66 | 2.13 |
| Newcastle | 76,092 | 28,280 | 37.2 | 569 | 0.75 | 2.01 |
| Sunderland | 65,845 | 30,428 | 46.2 | 602 | 0.91 | 1.98 |
| Kinsley | 59,255 | 18,861 | 31.8 | 336 | 0.57 | 1.78 |
| Suffolk Downs | 30,559 | 10,717 | 35.1 | 189 | 0.62 | 1.76 |
| Hove | 76,315 | 41,541 | 54.4 | 713 | 0.93 | 1.72 |
| Doncaster | 67,033 | 26,789 | 40.0 | 404 | 0.60 | 1.51 |
| Harlow | 86,447 | 50,272 | 58.2 | 660 | 0.76 | 1.31 |
| Sheffield | 76,132 | 29,412 | 38.6 | 382 | 0.50 | 1.30 |
| Nottingham | 60,881 | 32,679 | 53.7 | 285 | 0.47 | 0.87 |
| Dunstall Park | 7,198 | 3,507 | 48.7 | 30 | 0.42 | 0.86 |
| Romford | 94,063 | 53,143 | 56.5 | 445 | 0.47 | 0.84 |
| Yarmouth | 49,857 | 20,697 | 41.5 | 128 | 0.26 | 0.62 |
| No longer operating | ||||||
| Swindon* | 58,278 | 21,270 | 36.5 | 427 | 0.73 | 2.01 |
| Perry Barr* | 52,202 | 25,184 | 48.2 | 481 | 0.92 | 1.91 |
| Henlow* | 13,653 | 6,989 | 51.2 | 103 | 0.75 | 1.47 |
| Crayford* | 59,997 | 37,466 | 62.4 | 491 | 0.82 | 1.31 |
| Active tracks total | 1,082,992 | 468,576 | 43.3 | 7,590 | 0.70 | 1.62 |
| All tracks total | 1,267,122 | 559,485 | 44.2 | 9,092 | 0.72 | 1.63 |
Adverse events, trouble events, and adverse-to-trouble ratios by Greyhound Board of Great Britain-licensed track, January 2022–March 2026.
*Ceased operating during the observation period.
Hove = Brighton and Hove; Star Pelaw formerly Pelaw Grange. Dunstall Park (Wolverhampton) opened in September 2025.
Several derived variables were computed from the assembled race-level data. A dog’s first race within the focal period was the earliest race entry on or after 1 January 2022; the last race was the latest entry on or before 31 March 2026. Racing duration (also referred to as career duration) was calculated as the interval in days between these. Age at first race was determined from the GBGB’s recorded date of birth (month–year precision) and the date of the first race. Racing intensity was computed as race starts per active month and grouped into quartiles for between-group comparisons. The joint relationship between racing intensity, place rate (the percentage of races that a dog placed in the top three positions, used as a proxy for the industry’s perceived performance value of a dog), and racing duration in the 2022 entry cohort was examined via cross-tabulation and Kruskal–Wallis tests of the marginal main effects.
Racing status was determined using an empirically derived 88-day retirement threshold based on the assembled dataset. Eighty-eight days was the 99th percentile of 1,234,366 observed inter-race gaps within the dataset, representing the longest gap after which a dog was still observed to return to racing. Dogs whose most recent race fell within 88 days of 31 March 2026 were classified as still active; those whose most recent race preceded that date by more than 88 days were classified as inactive. In a sensitivity analysis at alternative threshold values of 30, 60, 120, 180 and 365 days (Supplementary Table 1), the inactive-cohort proportion (68.2% to 74.6%) and median career duration (within 0.5 months) showed limited variation across 60–180 days. This indicated that the substantive findings reported are not affected by the threshold choice within this range, supporting the use of the 88-day value as the inactivity threshold throughout analyses.
Two boundary effects affected the focal cohort of racing dogs. Dogs whose racing started before 1 January 2022 (partial-history dogs) were present in the dataset, but only their races within the focal period contributed to the assembled race counts and career-level metrics; the reported racing duration for these dogs was therefore a lower bound on actual career duration (left-censoring). Dogs still actively racing on 31 March 2026 (racing-ongoing dogs) were excluded from descriptive statistics on completed careers (e.g., median career duration and retirement age). In the survivorship curves, their observation period was shorter, stopping with the focal period on 31 March 2026 (right-censoring). Cohort-level summary statistics for dogs that commenced racing in 2024 and 2025 (first-runner cohorts) reflected partial histories for some dogs; that information base will grow as more dogs from these years end racing. Differences in survivorship across entry cohorts (Figure 2) were tested using log-rank tests (Mantel, 1966; Peto and Peto, 1972); asymmetric tail-length overlap was avoided by testing at common horizons shared by the cohorts (6 and 18 months).
Figure 2
Adverse events that occurred during racing were identified from the structured remarks codes recorded in each dog’s race start comment field. These codes were recorded by track officials at race time to describe what was observed during the race. The codes describe between-dog interactions that are considered routine greyhound racing incidents (e.g., crowding and bumping) and events with direct welfare implications (e.g., falls, struck rail, and lame). The remarks codes were grouped into four adverse categories: severe (knocked over or fell, struck rail, stopped, and ran off the track), injury (came off lame and lame), medical (cramp), and did not finish. The full remarks-code dictionary is within the deposited dataset. We noted that the GBGB’s publicly accessible race results assign only the most recent trainer against every historical race for a dog, so trainer changes during a dog’s active racing are not reflected. All career durations were therefore attributed to the trainer with whom each dog finished racing.
4 Results
The findings from the assembled dataset are presented below as a case study illustrating the method (Section 3). Results demonstrate that AI agent-assisted data assembly from public registries can produce a population-level dataset of sufficient scale, coverage, and internal consistency to provide animal welfare insights. They also provide individual- and population-level findings for greyhound racing in the UK not previously available outside the governing body.
The original contributions presented in this study are publicly available. The de-identified dataset has been deposited in Zenodo at https://doi.org/10.5281/zenodo.19842376 under a Creative Commons Attribution 4.0 International (CC BY 4.0) licence; the accompanying analysis and assembly code are deposited in the same record under the MIT licence. The deposit comprises the pseudonymised dog-level and race-level datasets, the country-of-origin resolution log, the remarks-code lookup used to classify adverse events, a SQLite mirror of all tables, the data dictionary, the underlying values for the figures and tables in the manuscript, and a curated subset of pipeline and analysis code. The deposit is structured to align with FAIR principles for scientific data management (Wilkinson et al., 2016). Materials not included in the public deposit comprise the encrypted mapping key linking pseudonymised identifiers to source registry names. This is retained securely by the corresponding author and is not transferable. The underlying source registry data remain available to any reader at the URLs cited in Section 3.1. The deposited dataset is an aggregated, pseudonymised and analytically transformed research output; it is not a copy or republication of any source registry's underlying database.
4.1 The assembled population
The assembled dataset represents 31,028 dogs that raced in GBGB-regulated racing at least once during the 51-month period of 1 January 2022 and 31 March 2026 (inclusive). Table 1 summarises the demographic characteristics of the assembled cohort of dogs.
The sex ratio of female and male greyhounds was almost exactly equal. Over half of the cohort were black-coated (56.6%); together with black and white, brindle, and blue/blue brindle, these colour groups accounted for 95.8% of dogs. More dogs were born in 2021 (5,780 dogs, 18.6%); dogs born earlier than this were still racing after 1 January 2022, while younger dogs reflected those that entered the racing system during the focal period.
The 31,028 dogs were bred from 772 recorded sires and 6,252 dams. Sire representation was highly concentrated, with the five most common sires accounting for 7,205 dogs (23.2% of the dataset). Assembled data revealed that half of all dogs (51.4%) racing in the UK were produced by just 2.3% (n = 18) of sires. This pattern is consistent with the intensive use of a small number of commercially popular stud dogs and has implications for the genetic diversity of the UK greyhound racing population. Research commissioned by the GBGB, published this year, confirmed the rate of inbreeding in GB greyhounds to be one of the highest observed in dogs, offering that a ‘popular sire effect’ seen in thoroughbred racing can ‘coincide’ with increased inbreeding (Han et al., 2026). The assembled dataset provides robust evidence of the popular sire effect in greyhound racing in the UK.
Dataset coverage was assessed against the GBGB’s published annual registration figures (Greyhound Board of Great Britain, 2025a). The GBGB reported on new greyhound registrations by calendar year, while the assembled dataset identified dogs by the date of their first recorded race. The two metrics are not identical, as a dog registered with the GBGB may not yet have raced, may first race in the year after registration, or may never race competitively. However, the two figures should be broadly comparable, and the degree of agreement between them provides a practical indicator of the assembly method’s completeness. For 2022 dogs, the assembled dataset (once filtered for dogs whose racing career commenced prior to 1 January 2022, see Section 4.5) identified 6,183 new dogs, reflecting 98.0% agreement with the GBGB’s 6,307 registrations. For 2023 and 2024, the dataset represented 94.3% and 95.8% of the GBGB’s reported new registrations (5,899 and 5,133, respectively). The GBGB has not yet published registration figures for 2025 (at 01 June 2026); the assembled dataset recorded 4,624 new dogs entering racing in that year. The number of new greyhounds entering GBGB-regulated racing has declined in each successive year, from over 6,150 in 2022 to fewer than 4,650 in 2025. At the end of the focal period (31 March 2026), 8,456 (27.3%) of the assembled cohort of dogs were still actively racing, with the remaining 22,572 (72.7%) having exited the GBGB-regulated system.
4.2 Country of origin
The traceability and welfare assurance of greyhounds across international borders are a key concern for greyhounds that race in the UK (Scottish Animal Welfare Commission, 2023; Baga, 2026). The governing bodies’ visibility of canine welfare across all life stages is limited due to the international supply of greyhounds into the UK. This risks dogs falling between regulatory bodies when they are bred in one jurisdiction and raced in another. Breeding and rearing practices in the exporting country can be largely invisible and beyond the control of regulators in the importing country. While the GBGB provides an annual aggregated breakdown of the country of origin for newly registered dogs (Greyhound Board of Great Britain, 2025a), this information is not available or traceable at the individual dog level on its publicly facing website. Establishing where each dog was born required cross-referencing against external registries through the strict stepwise verification process (Section 3.2). Origin resolution for each dog was both a substantive finding and a test of the assembly method’s capability to build evidence not present from a single existing source.
Of the 31,028 dogs, the majority (85.1%) were classified as Irish-bred (Table 2). The ICC Stud Book confirmed Irish origin for 14,114 dogs (45.5% of the cohort) in the first step of the process, comprising 14,076 authoritative matches where name, sire, dam, and whelp date were confirmed, and 38 high-confidence matches (a name typo, or name and parents confirmed, or whelp date within 2–3 months of GBGB documented date). The British Greyhound Stud Book step confirmed the UK origin for 3,641 dogs (11.7%), of which 3,095 were matched on sire, dam, and whelp date, 41 on sire and dam with a different litter date, and 505 on dam only. A further 13,061 dogs were unresolved after these first two steps and were cross-checked against Greyhound-Data, which placed 12,305 as Irish-bred and 752 as UK-bred. Only 4 (<0.1%) dogs were identified as born in countries outside of the UK or Ireland (Romania and the United States). The origin of the remaining 212 dogs (0.7%) could not be resolved at any step.
The proportion of Irish-bred dogs in the assembled dataset (83.5%–84.5% across the 2023–2025 first-runner cohorts) is consistent with the GBGB’s published figures (Greyhound Board of Great Britain, 2025a) of 82.4% (2023) and 84.5% (2024). For reference, the ICC registered 12,438 non-coursing greyhound births in 2021 and 11,698 in 2022 (PQ 721 and 722, Houses of the Oireachtas, 2023). Allowing for an approximate 12–18-month interval between birth and first race, these figures indicate that approximately 40%–45% of ICC-registered non-coursing pups that were bred in the Republic of Ireland in 2021–2022 entered GBGB racing in the UK the following year. Irish-bred dogs accounted for 84.6% (n = 1,072,267) of all race starters in the assembled dataset.
4.3 Racing activity
The 31,028 dogs in the assembled dataset participated in 1,267,122 individual race starts across 218,009 races at 22 licensed tracks, drawn from 18,174 race meetings between 1 January 2022 and 31 March 2026. A further 40,815 vacant trap entries (indicating empty start boxes, typically due to late dog withdrawals) were identified and excluded during data assembly. Flat races accounted for 99.7% of starts, with hurdle races comprising the remainder. A small proportion of records (0.5%) related to non-competition individual trials; these were identifiable and can be filtered for, but are retained in all figures reported here. The number of race starts per year was 295,223 in 2022, 307,321 in 2023, 309,248 in 2024, and 288,048 in 2025, with 67,282 starts from the first quarter of 2026.
Among dogs that raced for the first time between 1 January 2022 and 31 March 2026, the median age at first race was 21 months (IQR 19–25; mean 22.4, SD 5.3), with the youngest dog racing at 15 months. This was consistent across entry cohorts (median of 21 months in each year, from 2022 to 2025), suggesting that the typical age for trainers to start racing greyhounds is approximately 1 year 9 months.
At the individual level, dogs recorded a median of 32 race starts within the focal period (IQR 14–59, mean 40.8, range 1–237). Racing duration (measured as the interval between a dog’s first and last recorded race, right-censored at the assembly date for the 8,456 dogs still actively racing) had a median of 363 days (IQR 156–630 days; mean = 422 days). Among dogs classified as no longer racing on 31 March 2026 (n = 22,572), median career duration was 358 days (IQR 152–610, mean 410 days) with a median of 30 race starts (IQR 13–56). These figures indicate that the typical greyhound exited racing after approximately 30 starts run over a period of just under 12 months. Applying the 88-day retirement threshold (defined in Section 3.4), 22,572 dogs (72.7%) were classified as inactive and 8,456 (27.3%) as still actively racing on 31 March 2026.
Year-on-year attrition (the proportion of dogs active in a given calendar year whose last recorded race also fell in that year) was 38.8% in 2022, 38.7% in 2023, 39.6% in 2024, and 43.0% in 2025. The 2025 proportion may be somewhat inflated by the proximity of the end of the data assembly’s focal period, but the pattern appears consistent: for every five dogs racing in a given year, two will exit racing in that same year.
Figure 2 presents survivorship curves of dogs that commenced racing in each year. This shows the proportion of first-time racers, grouped by the year that they started racing, still actively racing over time after their first race. Each cohort line was trimmed consistently by 4 months to allow for latency in record updates. Within our assembled dataset, the 2022 cohort (n = 6,183) had the longest observation window. Among these dogs, one in six (16.7%) had ceased racing within 6 months of their first race, one in three (33.5%) had exited racing within 12 months, and half (50.0%) had exited racing within 18 months. By 24 months, two-thirds (65.2%) had exited, and by 36 months, 89.1%. Only 273 dogs from the 2022 cohort (4.4%) remained active on 31 March 2026. Among inactive dogs in this cohort, the median time from first race to exit was 17 months.
The trajectories in Figure 2 are closely comparable among all four entry cohorts. The 2022 and 2023 groups, whose longer observation periods allow for better comparison, remained within 1 percentage point of each other at 12 months (33.5% vs 33.9%) and 18 months (50.0% vs 50.9%). The 2024 and 2025 cohorts are limited by their shorter observation periods but track the same trajectory at earlier time points. Log-rank tests at common follow-up horizons confirmed this, returning no detectable difference across cohorts (all four cohorts compared at 6 months: χ2 = 0.58, df = 3, p = 0.90; the 2022, 2023, and 2024 cohorts compared at 18 months: χ2 = 1.78, df = 2, p = 0.41). The consistency of these curves across the four entry cohorts (2022–2025), spanning the May 2022 introduction of the GBGB’s ‘A Good Life For Every Greyhound’ welfare strategy, reveals no detectable change in the rate at which dogs exit racing (Armitage, 1955). Whether these attrition patterns differ from those before the introduction of the welfare strategy cannot be determined from this study.
The relationship between racing intensity and time in active racing was examined with finer granularity within the 2022 entry cohort (n = 5,837 inactive dogs with more than one race start at the end of the focal period). This year was chosen because it offered the longest observation window. Dogs were grouped by intensity quartiles based on the number of race starts per month. Dogs that raced at moderate intensity (quartiles 2 and 3; 2.3–3.5 race starts per month) had the longest median time in racing (19.8 and 19.7 months, respectively) and demonstrated the most total starts (median 51 and 63). Dogs in the highest intensity quartile (more than 3.5 race starts per month) had shorter racing durations (median 13.3 months) and, despite being raced more frequently, were noted to accumulate no more total starts (median 51) than moderately raced dogs. By comparison, dogs in the lowest intensity quartile showed the second-shortest active racing durations (median 16.1 months) and the fewest total starts (median 27). This pattern suggests that higher racing intensity is associated with an earlier exit from racing, while dogs that raced at moderate intensity had the longest racing durations. Higher racing intensity was not associated with improved finishing outcomes.
Extending this analysis, the joint relationship between racing intensity, the industry’s perceived performance value of each dog, and time in active racing was examined. Place rate (percentage of races finishing in the top 3) was used as a proxy for the industry’s perceived performance value and the consequent decision to continue racing, rather than as a welfare-relevant indicator of canine experience. Racing duration in months was cross-tabulated by intensity quartile and place-rate quartile (Table 3). Both factors contributed to career duration, with place rate being the stronger predictor (Kruskal–Wallis H = 492.3, df = 3, p < 0.001, with moderate effect size ϵ2 = 0.084) than intensity (H = 207.6, df = 3, p < 0.001, with small effect size ϵ2 = 0.035). The shortest racing duration occurred where the highest racing intensity coincided with the lowest place rate: median racing duration of 3.9 months, noted as substantially shorter than the 18.2-month median across all other cells (Mann–Whitney U = 431,271, z = −18.07, n1 = 363, n2 = 5,469, p < 0.001 one-sided, with large effect size r = 0.57). This shows that dogs that raced most intensively with the lowest place rate were exited from racing roughly five times faster than other dogs. The descriptive analyses presented here cannot identify whether trainer practices, adverse racing events, injury, or individual dog characteristics underlie this finding of accelerated attrition among racing’s most intensively used but least successful animals, but the assembled dataset enables further targeted investigation at the population scale.
4.4 Adverse racing events
Across the dataset’s 1,267,122 race starts, 666,690 individual adverse and trouble events were reported in the assembled dataset, with 9,092 (0.72%) race starts involving at least one adverse event. The majority were classified as severe (8,944 events), of which 5,910 were falls, and 2,940 were instances of a dog striking the inside rail. A further 108 events reported cramp, 41 involved lameness, and two dogs were recorded as did not finish. These adverse events were not directly comparable to the GBGB’s summary of clinical injury data (Greyhound Board of Great Britain, 2025b), which is shared annually, and reported higher incidence rates (1.07%–1.20% of runs, 2022–2024) based on post-race veterinary examination.
At the individual dog level, nearly a quarter of dogs (7,349, 23.7% of the dataset) experienced at least one adverse event during this study’s observation period. Falls affected 5,069 (16.3%) of dogs; of these, 4,349 experienced a single fall, 613 experienced two, and 107 dogs experienced three or more. Struck-rail events affected 2,599 (8.4%) dogs. Adverse event rates were stable across years: 0.8% of starts in 2022, 0.7% in 2023, 0.7% in 2024, and 0.7% in 2025. This suggests that in an average racing year, one in 11 dogs will fall at least once, and one in 22 will strike the rail at least once. The GBGB reports annual track fatalities at the aggregate level: 99 in 2022, 109 in 2023, and 123 in 2024 (Greyhound Board of Great Britain, 2025b). The GBGB-published fatality rate was reported as a stable 0.03% across 2022, 2023, and 2024 (Greyhound Board of Great Britain, 2025b; Greyhound Board of Great Britain, 2025c). Calculated to higher precision (three decimal places) from the same source figures, after noting that more individual dogs died in 2024 than in 2022 despite fewer total runs, we found the fatality rate actually rose each year from 0.027% (2022) to 0.030% (2023) to 0.035% (2024). This represents a 30% increase in the rate of on-track fatalities across this 3-year period (Cochran–Armitage trend test, z = 1.76, one-sided p = 0.039) within our study’s focus, noting that 2025 data are yet to be released. These events are not captured at the individual or track level in public records or race-comment remarks codes and are therefore absent from the assembled counts presented in this section.
Adverse event rates varied substantially across racetracks. Using the assembled dataset, Table 4 presents the number of race starts, observed races involving events classified as trouble (crowding and bumping), observed races involving adverse events, and an adverse-to-trouble ratio (adverse events as a proportion of trouble events). A higher ratio indicates tracks where severe incidents occurred more often relative to racing contact that is considered routine. The ratio does not imply that specific trouble events led to specific adverse events. Rather, it compares the relative frequency of adverse incidents to that of ‘routine’ racing incidents across tracks and can be interpreted as an indicator of risk. The broader concept of comparing observed severity levels within monitored populations at the site-level for risk indication is supported by contemporary racing-injury surveillance literature and safety science (Palmer et al., 2021; Gibson et al., 2024; Bennet and Parkin, 2026).
GBGB, Greyhound Board of Great Britain.Central Park was noted to report the highest adverse-to-trouble ratio (2.91%); this shows that for every 100 trouble events recorded at that track, approximately three adverse events were also recorded. By comparison, Yarmouth (0.62%) and Romford (0.84%) reported substantially lower adverse-to-trouble ratios. Trouble rates were noted to vary by a ratio of roughly two across tracks (e.g., 30.8% at Oxford compared to 62.4% at Crayford or 58.2% at Harlow), but the adverse-to-trouble ratio varied by a factor of nearly five. Romford’s pattern showed one of the highest trouble rates of any track (56.5%) but one of the lowest adverse rates (0.47%) and subsequently low adverse-to-Trouble ratios (0.84%). This suggests that either the track configuration produced frequent but low-severity contact or adverse events may have been under-reported relative to other venues. These between-track insights offer a level of welfare visibility not previously available in publicly reported data, where injury and fatality figures have been presented as aggregated across all licensed stadia rather than by venue (Scottish Animal Welfare Commission, 2023; Greyhound Board of Great Britain, 2025b). The differences between tracks made visible from the assembled public-facing data merit further exploration.
4.5 Trainers, racing endpoints, and post-racing traceability
The 31,028 dogs were managed by 635 trainers, with an uneven distribution noted. Just 33 (5%) trainers are responsible for a quarter of all racing greyhounds in the UK, with one-third of trainers (n = 213) managing 80% of the dog population. Trainer load during the focal period varied from one to 531 dogs (median 23, IQR 6–68). A small proportion of trainers accounted for the majority of dogs in the GBGB-licensed system, with implications for welfare governance.
How long dogs actively raced for (racing career duration) before exiting racing (retirement) varied substantially. Racing duration across the 22,572 inactive dogs at the end of the focal period ranged from a single race (552 dogs) to just under 4 years of active racing, with a median of 11.7 months (Section 4.3). Some trainers retained dogs in racing for longer than others. Median racing career duration varied markedly across trainers. Among the 163 trainers with at least 50 retired dogs (16,016 dogs; 71% of the inactive cohort), median career length ranged from 4.8 to 19.8 months, a 15-month spread that was statistically significant (Kruskal–Wallis H = 815.1, p < 0.001) with a small-to-moderate effect size (ϵ² = 0.041).
The destinations of retired dogs are only partially documented by the governing body. While the GBGB does publish aggregate retirement data annually (e.g., 78), their reporting cannot be verified from our assembled data due to the lack of traceability at the individual level. In 2024, the GBGB released their most recent summary, indicating that 93.8% of the 6,181 exiting dogs that year were rehomed or retained (53.9% were placed with charitable organisations for rehoming, 26.2% retained by owner or trainer, 10.6% rehomed by owner or trainer, and 3.1% redirected to breeding or classified as ‘Other’). It is unclear what life looks like for the 1,618 dogs that were recorded as retained by owners/trainers or how this retention may be verified by the GBGB. The remaining 6.2% of dogs died with causes attributed to sudden death, veterinary euthanasia (at or away from the track), euthanised because no home was found or the dog was deemed unsuitable for rehoming, or natural causes. The GBGB’s reported fatalities between 2022 and 2024 represented a 24% increase in absolute on-track dog fatalities within this study’s focal period, again noting that 2025 data are yet to be released.
The lack of individual-level traceability has repeatedly been noted by independent reviews. The Scottish Animal Welfare Commission (2023) stated in 2023 that the pooled industry figures provide no basis for comparing dog welfare outcomes across venues or trainers. A recent advocacy report estimates that at least 1,000 retired greyhounds are unaccounted for each year (Baga, 2026). No public registry currently enables independent verification of individual dogs’ post-racing welfare. This traceability gap constrains what can be verified about greyhound welfare after racing ends and presents a data-access issue rather than a limitation of the assembly method described.
4.6 Data quality and cross-validation
The GBGB’s published ‘2018–2024 Injury and Retirement Data’ summary (Greyhound Board of Great Britain, 2025b) includes annual race start totals. Table 5 compares these counts with our assembled dataset. The assembled data cover between 81.5% (2022) and 86.9% (2024) of the GBGB’s reported totals, a level that meets the 80% completeness threshold commonly applied in epidemiological surveillance (European Centre for Disease Prevention and Control, 2014). The difference reflects race entries that the GBGB’s internal database does not expose via the publicly accessible online race results. The rates reported here are likely indicative of the true rate, but any variation from them cannot be determined without the GBGB providing full public disclosure of the complete underlying dataset. To assess if the missing 13%–19% of race-start records may be concentrated at particular locations or months (which could bias venue-level findings), we examined the temporal distribution of assembled race starts. Findings were consistent both month-by-month within each year and across tracks. The only high-variability tracks are related to explainable operational changes (e.g., closure). The missing records appear evenly distributed across months and locations and are unlikely to bias the findings (Agresti, 2018).
Table 5
| Year | GBGB published | Assembled dataset | Coverage (%) |
|---|---|---|---|
| 2022 | 362,427 | 295,223 | 81.5 |
| 2023 | 364,981 | 307,321 | 84.2 |
| 2024 | 355,682 | 309,248 | 86.9 |
| 2025 | Not yet published | 288,048 | – |
Annual race starts in Greyhound Board of Great Britain-regulated racing: published industry figures compared with the assembled dataset.
Note. Coverage (%) reflects the proportion of the assembled dataset’s race starts relative to GBGB’s reported ‘Total Dog Runs’ for the corresponding year.
Greyhound Board of Great Britain.
Race count accuracy was assessed at the population scale by cross-referencing the assembled dataset with public-facing data on Greyhound-Stats, an independent website that indexes race results from licensed GBGB tracks. All 31,028 dogs were queried, and of these, 99.9% (30,991 dogs) were found, and 37 (0.1%) returned no record. The two sources agreed precisely for 13,258 (42.7%) of the dogs. Greyhound-Stats was higher for 12,053 (38.8%, median excess 13 races), which was expected for some dogs given they had raced prior to 1 January 2022. Finally, Greyhound-Stats was lower for 5,680 (18.3%, median difference of three races), consistent with gaps in the website’s coverage rather than errors in the assembled data, as confirmed by subsequent sample verification, described below.
Further verification of the assembled dataset was conducted using a stratified random sample of 50 dogs drawn from the list, proportional to origin classification (Ireland 39, UK 6, and unknown or other 5), and verified by the corresponding author against the live GBGB results interface on 16 April 2026. All 50 dogs returned matched records. The live race count met or exceeded (excesses expected for dogs racing prior to 2022) the assembled count for every dog, confirming that the focal period bounded the data correctly. No systematic errors were noted in the fields checked. Together, the population-scale coverage check, independent source cross-validation, and race record level sample verification establish the assembled dataset as being fit for the welfare insights presented here at both population- and individual-level dog scales, extending the granularity available from the governing body’s existing aggregated published figures. Further, these findings demonstrate that AI agent-assisted assembly of data from public registries can produce meaningful data of sufficient quality, coverage, and visibility for animal welfare research applications.
5 Discussion
5.1 What the methodology achieved
This study demonstrated that AI agent-assisted data assembly from publicly accessible registries can produce a dataset of sufficient scale, coverage, and granularity for novel welfare insights to aid research into animal-reliant industries. The method’s case study explored greyhound racing in the UK and was able to assemble data representing 31,028 individual greyhounds, 1,267,122 race starts, and 666,690 individual adverse and trouble events recorded in race data across a 51-month focal period. The governing body, the GBGB, currently publishes welfare-relevant information as national aggregates, with no salient granularity for visible animal welfare insights (i.e., individual-dog, per-track, or per-trainer level) publicly available. The assembled dataset enables such insights, for example, documenting the track-specific adverse event patterns (Table 3) and survivorship curves by year (Figure 2). These analyses relied on information that is publicly available but was not previously feasible to assemble in practice. Our method, using an AI agent-assisted assembly, converted latent public-facing information into analysable research data to offer novel animal welfare insights. While public understanding of animal welfare partly depends on disclosure of relevant information in the first place, this method offers a practical way to generate independent welfare evidence about animal-reliant industries that refuse direct data access but maintain public-facing registries.
5.2 AI agent-assisted research: ethical considerations
While AI tools can improve animal welfare visibility, agentic AI research methodologies present important ethical issues. Below, we discuss and suggest ways to address ethical concerns using well-recognised ethical principles, research ethics frameworks, and regulation (Floridi et al., 2018; National Health and Medical Research Council, 2025a; National Health and Medical Research Council, 2025b). This section aims to both illuminate ethical issues encountered in this case study and provide general guidance for researchers intending to employ AI agent methodologies for animal-reliant industries.
Unlike standard AI technologies, AI agents can exhibit greater adaptability, context awareness, and autonomy (Acharya et al., 2025). This enables agents to work more independently on problems, including breaking tasks down into manageable subgoals executed with reduced human input (Dwivedi et al., 2026). While effective, this agency can make it harder to trace and document all steps in the process (Brohi et al., 2025). Furthermore, AI agents often incorporate familiar AI capabilities, such as computer vision, image generation, and large language models (LLMs), which operate with deep neural networks and processes like next-token prediction and reinforcement learning. This gives AI agents remarkable abilities of pattern recognition, natural language processing, computational reasoning, etc. Nevertheless, without strict guardrails and active human supervision, AI agents can suffer from well-known problems of bias, opacity, and inaccuracy or hallucination (confident but false outputs) (Acharya et al., 2025). Harm can sometimes result. For example, an AI agent tasked with making inferences from a dataset with information about greyhound owners or trainers may misrepresent the true situation. Investigators must therefore cross-check outputs for accuracy, especially when the exact basis of outputs is opaque. Such human oversight, as described earlier, is essential for responsible AI agent use (Brohi et al., 2025).
Another possible source of harm concerns privacy (Danilevskyi et al., 2025). As observed, AI agents can enhance privacy by de-identifying data, giving researchers a prima facie ethical reason to consider their use. However, AI tools can also process aggregated data to reveal otherwise hidden or obscured insights about individuals or groups (e.g., concerning the treatment of greyhounds by industry participants). Here, privacy-preserving steps may be taken, including ensuring secure on-site data storage with restricted access and collecting only necessary data. Further privacy risks stem from data entering commercial offshore AI systems, where the data may be used to train AI systems, passed on, or even sold to third parties (National Health and Medical Research Council, 2025a). To minimise data leakage or misuse, investigators should understand how various AI companies use data and choose arrangements that include data anonymisation, deletion, strong cybersecurity, and minimising data sent to company servers.
When risk of harm can be reduced as the ethical principle of non-maleficence demands, but not completely avoided, researchers are advised to weigh the magnitude of the harm (e.g., the sensitivity of personal information; impacts on vulnerable populations) and ensure that the research promises sufficient benefit according to the principle of beneficence (Beauchamp and Childress, 2026). Such a benefit may, for instance, involve satisfying a pressing public interest (e.g., accurately understanding greyhound welfare) but not various less serious interests. Using a relatively autonomous AI agent does not negate research obligations, and researchers retain final responsibility for data handling and published findings.
Privacy is influenced by contextual norms (Nissenbaum, 2004). For example, the public may have no right to know about someone’s religious views, but has every right to know about animal welfare outcomes in animal-reliant industries. Industries like greyhound racing depend on a social license to operate (Duncan et al., 2018), which may be rescinded and invite political and legal responses when animal treatment is considered unacceptable (Douglas et al., 2022). A further contextual privacy norm in our case study is that the UK greyhound industry is legally obligated to publish certain welfare-related facts. When such transparency fails to provide visibility of the animal experience, the public, policymakers, regulators, academic researchers, investigative journalists, and advocacy groups may justifiably promote visibility in the public interest, consistent with relevant ethics requirements (Drolet et al., 2023; National Health and Medical Research Council, 2025b). Another privacy consideration is whether relevant data are already publicly available. Research ethics guidelines and regulations generally waive the need for ethics committee review in such cases (National Health and Medical Research Council, 2025b), as in this study. Nevertheless, once again, AI tools can combine data (e.g., from social media or other platforms) to reveal new insights. Contextual awareness in data governance thus remains essential when applying AI agents even to non-confidential or public-facing data where privacy impacts are possible (Hassanpour and Yang, 2026).
Accountability in research with agentic AI requires attention to legal and regulatory requirements (e.g., privacy law and institutional ethics review) (Rhahla et al., 2021; Australian Government, 2022) and the diligent application of principles, like non-maleficence, beneficence, and privacy. It also requires recording and documenting the workflow pipeline to allow fellow scholars and peer reviewers to inspect management of ethical issues, assess methodological validity, and reproduce the research (Stoudt et al., 2024). Notably, the nature of AI agents and changes in the AI tool and/or the relevant data (e.g., published industry information) may affect study replicability (Brohi et al., 2025). However, this does not obviate the need for careful documentation and explanation of AI agent-assisted research methods.
5.3 What the findings reveal about visibility
The AI agent-assisted assembly method revealed welfare-relevant findings not visible in the governing body’s existing public reporting on greyhounds. Three of these observations are particularly consequential and address information gaps noted in formal inquiries (Welsh Parliament, 2025). First, adverse events varied substantially across tracks (Table 3), providing an insight absent from the GBGB’s national-aggregate injury data. The Scottish Animal Welfare Commission noted in 2023 that providing pooled data in this manner obscured stadia-specific evidence for whether participating in some races is more hazardous than in others (Scottish Animal Welfare Commission, 2023).
Second, specific patterns can emerge from assembled data that aggregate annual reporting misses. One pattern arises from comparing data sources, and another from tracking outcomes across time. Across sources, our assembled race-comment adverse event rate (0.73% of starts) was approximately two-thirds of the GBGB’s clinical injury rate (1.15% of runs) for the same period (Greyhound Board of Great Britain, 2025b). The two sources capture different events: the comments capture in-race incidents, while clinical diagnoses from post-race veterinary examinations report injuries. Musculoskeletal injuries make up most of the GBGB’s reported injuries (i.e., hock, wrist, fore, and hind limb muscles) but often emerge after the post-race veterinary examination or do not produce visible racing incidents (Gower, 2021; Palmer et al., 2021; Welsh Parliament, 2025). Welfare insights from race-comment data alone likely understate the incidence of clinical injuries. Considered together, the two data streams reveal that a substantial proportion of injuries occur without observable signs during the race.
Over time, the survivorship curves across the 2022–2025 entry cohorts (Figure 2) were effectively indistinguishable, differing by no more than 1 percentage point at 12 and 18 months after commencing racing. This provided no evidence that the rate of annual attrition has changed since the introduction of the GBGB’s ‘A Good Life for Every Greyhound’ welfare strategy in 2022 (Greyhound Board of Great Britain, 2022). The GBGB’s published fatality rate (Greyhound Board of Great Britain, 2025b), reported to two decimal places, appears stable across 2022–2024 at 0.03%. Revealed at a higher precision (3 decimal places), we showed (Section 4.5) that the underlying fatality rate rose by 30% over the same period, accompanied by a 24% increase in the number of individual dogs killed on track. This is another instance of self-regulatory disclosure that satisfies a regulatory transparency obligation (i.e., the UK’s Department for Environment, Food and Rural Affairs (DEFRA)) while concealing a deteriorating safety outcome. A reporting practice that lets a worsening trend serve as evidence of safety fails the welfare purpose that those disclosure obligations exist to achieve: protecting the dogs whose lives are subject to regulation. These patterns of adverse incidents, attrition, and fatalities are consistent with the inherent risks and physical demands of racing greyhounds at speed, in groups around a curved track, but are not apparent from any single source considered alone. Attrition at this rate and scale constitutes what the animal welfare science literature terms ‘wastage’—the systematic removal of animals from a sport/operation at the population level (Branson et al., 2010; Cobb et al., 2015a; Cobb et al., 2015b; Cobb et al., 2021)—and represents a core concern for animal welfare and racing’s social license (Hampton et al., 2020; Stevens et al., 2022).
Finally, post-racing destinations remain markedly opaque. The GBGB’s most recent (2024) retirement commentary reports that 1,618 greyhounds were retained by their owner or trainer, but only 299 (18.5%) were confirmed as living ‘permanently as pets’ (Greyhound Board of Great Britain, 2025c). The status of the remaining 1,319 dogs is unclear and cannot be verified individually in any public record. Furthermore, these dogs will not be reported on again by the GBGB in future updates, as the GBGB has classified them as ‘successfully retired from the sport’ (Greyhound Board of Great Britain, 2025c) after they exited the licensed racing system. By treating racing retirement as the endpoint of its welfare oversight, the GBGB’s reporting effectively transfers responsibility for post-racing welfare outside any public accountability framework (Stevens et al., 2022). The kennel audits that the GBGB cites as evidence of their independent welfare oversight (Greyhound Board of Great Britain, 2026) are limited to actively racing dogs. The British Standards Institution’s (BSI’s) (the UK’s National Standards body) Publicly Accessible Specification (PAS) for greyhound trainers’ residential kennels (BSI PAS251:2017, British Standards Institution, 2017), which underpinned those audits, explicitly excluded retired greyhounds. The PAS was withdrawn on 11 April 2025 (British Standards Institution, 2017), without a publicly available replacement (typical practice to maintain standard currency), making it unclear against which standard the current audit findings are measured. Continuing to cite the withdrawn PAS as the basis for UK Accreditation Service (UKAS-accredited) inspections creates a credibility gap where the regulator claims the authority of the BSI without maintaining the BSI’s quality that standards remain current. Advocacy reporting has estimated that at least 1,000 retired greyhounds in the UK are unaccounted for each year (Baga, 2026). No public registry systematically verifies individual dogs’ post-racing welfare outcomes; voluntary self-reporting on Greyhound-Data captures only a small fraction of rehoming events.
5.4 Social license, public interest, and independently generated evidence
The AI agent-assisted data assembly method presented here can gather welfare evidence in a form that has been explicitly identified as lacking (Houses of the Oireachtas, 2023; Scottish Animal Welfare Commission, 2023; Welsh Parliament, 2025). In both Scotland and Wales, parliamentary processes examining greyhound racing would have benefited from individual-level welfare evidence at the population scale informing legislative decisions. The method is also transferable beyond this case study. It could be used to increase welfare visibility in other animal-reliant industries with public-facing registries (e.g., horse racing, working animals, zoos, and animal research laboratories); other applications may extend beyond animal-specific contexts.
More fundamentally, AI agent-assisted data assembly methods can change the terms of welfare policy debates (Nurse, 2016; Muhammad et al., 2022). Public accountability of animal-reliant industries depends on welfare evidence not wholly controlled by the operating industry (Goodfellow, 2016). This study has demonstrated a new methodological contribution to assembling such independent evidence from public sources, supporting public interest and informed animal welfare policy discussion.
5.5 Limitations, accountability, and future directions
This methodology assembled what public registries contain. Because it cannot retrieve what they lack, various animal welfare issues may escape clear public view. The public data available for this study, for example, did not contain information about greyhound living conditions and training methods, which directly affect welfare. Still, the absence of welfare-relevant data, such as missing information about the fate of hundreds of retired greyhounds, is itself a revealing visibility and accountability gap.
Endpoint outcome data rely on voluntary reporting by new owners via www.Greyhound-Data.com and significantly undercount the true rehoming burden often carried by shelters and rescue groups. Causes for dogs exiting racing are not publicly recorded: shorter careers at particular tracks, or under particular trainers, cannot be attributed to specific causes from this dataset alone. Dogs that commenced racing in 2024 and 2025 include many that are still racing; cohort-level statistics for these years will change as more dogs retire. Observer-driven variation in race-comment coding cannot be excluded as a contributor to the reported patterns of between-track adverse events, as these are entered in real time by track officials rather than a standardised instrument. This analysis covers GBGB-licensed racing in the UK only; independent tracks operating in the UK outside GBGB regulation had closed by late 2022 and were not included. This represents a source of selection bias, as dogs raced exclusively at independent venues before their closures, or those that transitioned to GBGB-regulated racing within the focal period, are not captured in the assembled dataset.
Our paper’s analyses of career duration, attrition, and adverse-event exposure represent only the GBGB-licensed racing population and cannot be generalised to the dogs whose careers included independent venues. This bias is likely small given that independent tracks effectively ceased operations by late 2022, but it cannot be quantified from publicly available data. This methodology depends on the continued public accessibility of registry data. Should industries respond to AI agent-assisted data assembly by restricting public access, the resultant more acute transparency failure would strengthen the case for mandatory, independent disclosure of granular welfare data.
The deposited data warrant further analyses to bring additional welfare insights beyond the scope of the case study reported here. These could include regression-based modelling of trainer, track, season, racing intensity, and cohort effects on welfare outcomes. Periodic repetition of the data assembly would allow monitoring to determine if welfare strategy interventions produce detectable effects as new data become available. The deposited dataset and accompanying scripts provide a fixed reference point for reproducibility. Independent re-assembly using the same public sources may require method revision over time. An equally important direction is research to explore how independently generated welfare evidence is received, trusted, and acted upon by different stakeholder groups, including industry participants, regulators, policymakers, welfare organisations, and the public.
6 Conclusion
The novel AI agent-assisted data assembly method presented in this paper can efficiently produce de-identified population-level welfare datasets in animal-reliant industries where direct data sharing is refused. The case study documents welfare-relevant patterns in the UK greyhound racing population that were previously invisible to independent researchers, policymakers, or the public. These contributions address a structural problem in animal welfare governance—the information asymmetry between animal-reliant industry operators and everyone else. Disclosure does not necessarily produce visibility, and visibility does not guarantee accountability. The GBGB publishes national aggregate welfare figures and describes its approach as transparent. The dataset assembled here demonstrates what that curated transparency does not make visible. Approximately one in seven dogs racing in the UK experiences at least one adverse event every year, and a substantial number of retired dogs will not have their post-racing status verified in any public record. Greyhound racing is losing its social license to operate across multiple jurisdictions simultaneously. Whether improved access to welfare data will alter that trajectory or inform a responsible industry exit is unclear. Understanding how improved visibility can activate accountability for animal welfare assurance, especially when industry transparency has been lacking, is the necessary next step. By demonstrating how AI agents can help assemble data to generate insights about greyhound racing in the UK in an ethically responsible way, this study offers guidance for researchers seeking to use agentic AI tools to enhance welfare visibility in animal-reliant industries.
Statements
Data availability statement
The original contributions presented in the study are included in the article and Supplementary Material. Further inquiries can be directed to the corresponding author.
Ethics statement
This study used publicly accessible data from greyhound racing registries. No animals were handled, observed, or otherwise directly involved in the research. The dataset contains information about individual greyhounds derived from public registry records; the study was exempt from animal ethics approval. The dataset removed names of individual trainers, owners, and breeders drawn from public-facing registry sources during the assembly process. A formal determination on whether this constitutes human subjects research was sought from the University of Melbourne’s Office of Research Ethics and Integrity prior to submission. All human identifiers in the deposited dataset were removed or replaced with sequential numerical codes before researchers engaged with the data. An encrypted mapping key is held securely by the corresponding author.
The use of AI agent-assisted data collation tools was conducted in accordance with the University of Melbourne’s Research Integrity and Misconduct Policy and its guidance on the use of digital assistance tools in research. The use of Claude Cowork agents (Anthropic, Inc.) was limited to data collation tasks. No AI tools were used for data interpretation or manuscript generation. AI tool use has been documented in the corresponding author’s research diary as required by local institutional guidelines.
Author contributions
MC: Writing – original draft, Writing – review & editing, Data curation, Methodology, Conceptualization, Formal analysis, Project administration, Validation, Investigation, Visualization. SC: Writing – original draft, Writing – review & editing, Methodology.
Funding
The authors declared that financial support was received for this work and/or its publication. MC is supported by the Chaser Innovation Research Fellowship (The Chaser Initiative and The University of Melbourne) and the John McKenzie Fellowship (The University of Melbourne). This study did not receive additional project-specific funding.
Conflict of interest
The authors declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Generative AI statement
The authors declared that generative AI was not used in the creation of this manuscript. The use of AI agent-assisted data collation tools was conducted in accordance with the University of Melbourne’s Research Integrity and Misconduct Policy and its guidance on the use of digital assistance tools in research. The use of Claude Cowork agents (Anthropic, Inc.) was limited to data collation tasks. No AI tools were used for data interpretation or manuscript generation. AI tool use has been documented in the corresponding author’s research diary as required by local institutional guidelines.
Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.
Publisher’s note
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
Supplementary material
The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fanim.2026.1868726/full#supplementary-material
References
1
Abou AliM.DornaikaF.CharafeddineJ. (2025). Agentic AI: a comprehensive survey of architectures, applications, and future directions. Artif. Intell. Rev.59, 11. doi: 10.1007/s10462-025-11422-4
2
AcharyaD. B.KuppanK.DivyaB. (2025). Agentic AI: Autonomous intelligence for complex goals—A comprehensive survey. IEEE Access13, 18912–18936. doi: 10.1109/access.2025.3532853
3
AgrestiA. (2018). An introduction to categorical data analysis, 3rd edition (Hoboken, NJ, USA: John Wiley & Sons).
4
AnannyM.CrawfordK. (2018). Seeing without knowing: Limitations of the transparency ideal and its application to algorithmic accountability. New Media Soc.20, 973–989. doi: 10.1177/1461444816676645
5
ArmitageP. (1955). Tests for linear trends in proportions and frequencies. Biometrics11, 375–386. doi: 10.2307/3001775
6
AugerG. A. (2014). Trust me, trust me not: An experimental analysis of the effect of transparency on organizations. J. Public Relations Res.26, 325–343. doi: 10.1080/1062726x.2014.908722
7
Australian Government (2022). The Australian Privacy Principles. Office of the Australian Information Commissioner. Available online at: https://www.oaic.gov.au/privacy/australian-privacy-principles.
8
European Parliament & Council of the European Union (2024). Regulation (EU) 2024/1689 of the European Parliament and of the Council laying down harmonised rules on artificial intelligence (Artificial Intelligence Act), Article 14 (Human Oversight). Official Journal of the European Union. Available online at: https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=OJ:L_202401689.
9
BagaP. (2026). Reaching the finish line: time to end dog racing in the UK (London: GREY2K USA Worldwide; The League Against Cruel Sports).
10
BeauchampT. L.ChildressJ. F. (2026). Principles of biomedical ethics (9th edition) (London: Oxford University Press).
11
BekoffM. (2022). Time to stop pretending we don’t know other animals are sentient beings. Anim. Sentience6, 2. doi: 10.51291/2377-7478.1699
12
BenchimolE. I.SmeethL.GuttmannA.HarronK.MoherD.PetersenI.et al. (2015). The REporting of studies Conducted using Observational Routinely-collected health Data (RECORD) Statement. PloS Med.12, e1001885. doi: 10.1371/journal.pmed.1001885
13
BennetE. D.ParkinT. D. H. (2026). Novel risk factors associated with fatal musculoskeletal injury in Thoroughbreds in North American racing (2009–2023). Equine Vet. J.58, 20–30. doi: 10.1111/evj.14503
14
BjørkdahlK.SyseK. V. L. (2021). Welfare washing: Disseminating disinformation in meat marketing. Soc. Anim.32, 37–55. doi: 10.1163/15685306-bja10032
15
BransonN.CobbM.McGreevyP. (2010). Australian working dog industry survey report: 2009. Canberra, ACT, Australia: Australian Government Department of Agriculture, Fisheries and Forestry.
16
BreakeyH.WoodG.SampfordC. (2025). Understanding and defining the social license to operate: Social acceptance, local values, overall moral legitimacy, and ‘moral authority’. Resour. Policy102, 105488. doi: 10.1016/j.resourpol.2025.105488
17
British Standards Institution (2017). Specification for greyhound trainers' residential kennels [withdrawn 11 April 2025]. PAS 251:2017 (London: BSI).
18
BrohiS.MastoiQ.JhanjhiN.PillaiT. R. (2025). A research landscape of agentic ai and large language models: Applications, challenges and future directions. Algorithms18, 499. doi: 10.3390/a18080499
19
BrownM. A.GruenA.MaldoffG.MessingS.SandersonZ.ZimmerM. (2025). Web scraping for research: Legal, ethical, institutional, and scientific considerations. Big Data Soc.12, 20539517251381686. doi: 10.1177/20539517251381686
20
CampbellM. L. H. (2023). Ethical justifications for the use of animals in competitive sport. Sport Ethics Philosophy17, 403–421. doi: 10.1080/17511321.2023.2236798
21
ChandlerG. S.SalasL.Le SerraB.SethiS. (2025). AI guidelines (Sydney, Australia: The Research Society).
22
ChangV.DescovichK.HenningJ.AllavenaR. (2022). Greyhound morbidity and mortality in Australia: A descriptive analysis of reported data from regulatory racing agencies. Front. Vet. Sci.9. doi: 10.3389/fvets.2022.925948
23
CobbM.BransonN.McGreevyP.BennettP. C.RooneyN.MagdalinskiT. (2015a). Review & Assessment of best practice rearing, socialisation, education and training methods for greyhounds in a racing context. 110, 96–104. doi: 10.13140/RG.2.1.1822.7043
24
CobbM.BransonN.McGreevyP.LillA.BennettP. (2015b). The advent of canine performance science: offering a sustainable future for working dogs. Behav. Processes110, 96–104. doi: 10.1016/j.beproc.2014.10.012
25
CobbM.GainesS. (2023). Animal welfare assurance in the 21st century: Navigating the landscape of welfare-washing, regulatory capture, and social license to operate [Conference presentation]. International Animal Welfare Conference. Universities Federation for Animal Welfare.
26
CobbM. L.LillA.BennettP. C. (2023). Not all dogs are equal: perception of canine welfare varies with context. Anim. Welfare29, 27–35. doi: 10.7120/09627286.29.1.027
27
CobbM. L.OttoC. M.FineA. H. (2021). The animal welfare science of working dogs: Current perspectives on recent advances and future directions. Front. Vet. Sci.8, 666898. doi: 10.3389/fvets.2021.666898
28
CoghlanS.ParkerC. (2023). Harm to nonhuman animals from AI: A systematic account and framework. Philos. Technol.36, 25. doi: 10.1007/s13347-023-00627-6
29
DanilevskyiM.Perez-TellezF.BuscaldiD. (2025). Implementing ethical principles in AI: an initial discussion. AI Ethics5, 3549–3555. doi: 10.1007/s43681-025-00710-y
30
DiVincentiL.McDowellA.HerrelkoE. S. (2023). Integrating individual animal and population welfare in zoos and aquariums. Anim. (Basel)13. doi: 10.3390/ani13101577
31
DouglasJ.OwersR.CampbellM. L. (2022). Social licence to operate: what can equestrian sports learn from other industries? Animals12, 1987. doi: 10.3390/ani12151987
32
DroletM.-J.Rose-DerouinE.LeblancJ.-C.RuestM.Williams-JonesB. (2023). Ethical issues in research: Perceptions of researchers, research ethics board members and research ethics experts. J. Acad. Ethics21, 269–292. doi: 10.1007/s10805-022-09455-3
33
DuncanE.GrahamR.McManusP. (2018). ‘No one has even seen … smelt … or sensed a social licence’: Animal geographies and social licence to operate. Geoforum96, 318–327. doi: 10.1016/j.geoforum.2018.08.020
34
DwivediY. K.HelalM. Y. I.ElgendyI. A.AlahmadR.WaltonP.SuhA.et al. (2026). Agentic AI systems: What it is and isn't. Global Business Organizational Excellence45, 253–263. doi: 10.1002/joe.70018
35
EslakeS. (2025). The financing of greyhound racing in Tasmania. Eslake S. The financing of greyhound racing in Tasmania. Hobart, Australia; 2025.
36
European Centre for Disease Prevention and Control (2014). European centre for disease prevention and control (Stockholm: ECDC).
37
FlemingP. A.WickhamS. L.Dunston-ClarkeE. J.WillisR. S.BarnesA. L.MillerD. W.et al. (2020). Review of livestock welfare indicators relevant for the Australian live export industry. Anim. (Basel)10. doi: 10.3390/ani10071236
38
FloridiL.CowlsJ.BeltramettiM.ChatilaR.ChazerandP.DignumV.et al. (2018). AI4People—An ethical framework for a good AI society: Opportunities, risks, principles, and recommendations. Mines Mach.28, 689–707. doi: 10.31235/osf.io/2hfsc_v1
39
FoxJ. (2007). The uncertain relationship between transparency and accountability. Dev. Pract.17, 663–671. doi: 10.1080/09614520701469955
40
FoxeK. (2026). “ Greyhound Racing Ireland sought State support as welfare and rehoming costs rose,” in The irish times. Dublin, Ireland: The Irish Times.
41
FoxxV.CarbajalS. (2025). H.R.5017: Greyhound Protection Act of 2025 [Legislative bill]. US House of Representatives, 119th Congress. Available online at: https://www.congress.gov/bill/119th-congress/house-bill/5017/text.
42
FungA.GrahamM.WeilD. (2007). “Full disclosure: The perils and promise of transparency,” Cambridge: Cambridge University Press City.
43
GibsonM. J.LeggK. A.GeeE. K.SmetA.MeddJ.McMullenC.et al. (2024). Incidence and risk factors for limb fracture in greyhound racing in Western Australia. Aust. Vet. J.102, 543–549. doi: 10.1111/avj.13377
44
GilbertR.LaffertyR.Hagger-JohnsonG.HarronK.ZhangL. C.SmithP.et al. (2018). GUILD: GUidance for Information about Linking Data sets. J. Public Health (Oxf)40, 191–198. doi: 10.1093/pubmed/fdx037
45
GoodfellowJ. (2016). “ Regulatory capture and the welfare of farm animals in Australia,” in Animal law and welfare-international perspectives (City: Cham, Switzerland: Springer), 195–235.
46
GowerS. (2021). “ Greyhound Board of Great Britain website,” in Greyhound board of great britainLondon: Greyhound Board of Great Britain. Available online at: https://www.gbgb.org.uk/veterinary-blog-injury-detection-in-racing-greyhounds/.
47
Greyhound Board of Great Britain (2005). “ GBGB to launch own registration system for British-bred puppies,” in GBGB website (London: GBGB). Available online at: https://www.gbgb.org.uk/gbgb-to-launch-own-registration-system-for-british-bred-puppies/ (Accessed April 2026).
48
Greyhound Board of Great Britain (2022). A good life for every greyhound. Animal welfare strategy (London: Greyhound Board of Great Britain).
49
Greyhound Board of Great Britain (2025a). Delivering ‘A good life for every greyhound’ Progress report: october 2025 (London: Greyhound Board of Great Britain).
50
Greyhound Board of Great Britain (2025b). Injury and retirement data summary, 2018–2024 (London: Greyhound Board of Great Britain).
51
Greyhound Board of Great Britain (2025c). Licensed greyhound racing: independently verified track injury and retirement data for 2024 (Commentary report) (London: Greyhound Board of Great Britain).
52
Greyhound Board of Great Britain (2026). With multiple kennel inspections each year, racing greyhounds receive far higher protection than domestic pets [LinkedIn post]. LinkedIn. Available online at: https://www.linkedin.com/posts/the-greyhound-board-of-great-britain-limited_greyhoundcommitment-greyhoundracing-activity-7462795923174146048-M-eD
53
GunninghamN.KaganR. A.ThorntonD. (2004). Social license and environmental protection: why businesses go beyond compliance. Law Soc. Inq.29, 307–341. doi: 10.1111/j.1747-4469.2004.tb00338.x
54
Haibe-KainsB.AdamG. A.HosnyA.KhodakaramiF.Wendell4. D. J. 4. F. C. 4. M-A.WaldronL.et al. (2020). Transparency and reproducibility in artificial intelligence. Nature586, E14–E16. doi: 10.1038/s41586-020-2766-y
55
HamptonJ. O.JonesB.McGreevyP. D. (2020). Social license and animal welfare: Developments from the past decade in Australia. Animals10. doi: 10.3390/ani10122237
56
HanH.BlackettT. A.CampbellM. L. H.HoltbyA. R.McGivneyB. A.HillE. W. (2026). Genomic diversity and selection in the racing greyhound of Great Britain. Proc. R. Soc. B. Biol. Sci.293. doi: 10.1098/rspb.2025.2552
57
HårstadR. M. B. (2024). The politics of animal welfare: A scoping review of farm animal welfare governance. Rev. Policy Res.41, 679–702. doi: 10.1111/ropr.12554
58
HassanpourA.YangB. (2026). Contextual integrity in large language models: A review. J. Cybersecurity Privacy6, 74. doi: 10.3390/jcp6020074
59
HealdD. (2006). “ Varieties of transparency,” in Oxford university press for the british academy. Oxford: Oxford University Press for The British Academy.
60
Helms AndersenT.MarcussenT. M.TermannsenA. D.LawaetzT. W. H.NørgaardO. (2025). Using artificial intelligence tools as second reviewers for data extraction in systematic reviews: A performance comparison of two AI tools against human reviewers. Cochrane Evidence Synth. Methods3, e70036. doi: 10.1002/cesm.70036
61
HolmbergT.IdelandM. (2012). Secrets and lies: "selective openness" in the apparatus of animal experimentation. Public Underst Sci.21, 354–368. doi: 10.1177/0963662510372584
62
Houses of the Oireachtas (2023). Written answers nos. 721–740: greyhound industry (Dublin, Ireland: Houses of the Oireachtas).
63
JakkuE.TaylorB.FlemingA.MasonC.FielkeS.SounnessC.et al. (2019). If they don’t tell us what they do with it, why would we trust them? Trust, transparency and benefit-sharing in Smart Farming. NJAS Wageningen J. Life Sci.90-91, 100285. doi: 10.1016/j.njas.2018.11.002
64
KhanA. A.BadshahS.LiangP.WaseemM.KhanB.AhmadA.et al. (2022). “ Ethics of AI: A systematic literature review of principles and challenges”, in: Proceedings of the 26th International Conference on Evaluation and Assessment in Software Engineering. New York, NY, United States: Association for Computing Machinery (ACM). doi: 10.1145/3530019.3531329
65
KlineC.HooperJ. (2025). Saving animals or saving face? Avoiding welfare-washing in animal tourism partnerships. J. Ecotourism, 1–16. doi: 10.1080/14724049.2025.2559985
66
KruskalW. H.WallisW. A. (1952). Use of ranks in one-criterion variance analysis. J. Am. Stat. Assoc.47, 583–621. doi: 10.1080/01621459.1952.10483441
67
LiptovszkyM. (2024). Advancing zoo animal welfare through data science: scaling up continuous improvement efforts. Front. Vet. Sci.11. doi: 10.3389/fvets.2024.1313182
68
Lundmark HedmanF.AnderssonM.KinchV.LindholmA.NordqvistA.WestinR. (2021). Cattle cleanliness from the view of Swedish farmers and official animal welfare inspectors. Anim. (Basel)11. doi: 10.3390/ani11040945
69
MacielC. T.BockB. (2013). Modern politics in animal welfare: The changing character of governance of animal welfare and the role of private standards. Int. J. Sociology Agric. Food.20, 219–235.
70
MantelN. (1966). Evaluation of survival data and two new rank order statistics arising in its consideration. Cancer Chemother. Rep.50, 163–170.
71
MarkwellK.FirthT.HingN. (2017). Blood on the race track: an analysis of ethical concerns regarding animal-based gambling. Ann. Leisure Res.20, 594–609. doi: 10.1080/11745398.2016.1251326
72
MataF.MarquesM. R. (2026). Animal welfare washing in agriculture supply chains: Regulatory gaps, trade incentives, and ethical risks. World7, 48. doi: 10.3390/world7030048
73
McmorrowC. (2026). “ UK bans reignite debate over greyhound racing in Ireland,” in Raidió Teilifís éireann (RTÉ). Dublin, Ireland: Raidió Teilifís Éireann (RTÉ).
74
MillsK. E.HanZ.RobbinsJ.WearyD. M. (2018). Institutional transparency improves public perception of lab animal technicians and support for animal research. PloS One13, e0193262. doi: 10.1371/journal.pone.0193262
75
Mosqueira-ReyE.Hernández-PereiraE.Alonso-RíosD.Bobes-BascaránJ.Fernández-LealÁ. (2023). Human-in-the-loop machine learning: a state of the art. Artif. Intell. Rev.56, 3005–3054. doi: 10.1007/s10462-022-10246-w
76
MuhammadM.StokesJ. E.MorgansL.ManningL. (2022). The social construction of narratives and arguments in animal welfare discourse and debate. Animals12, 2582. doi: 10.3390/ani12192582
77
MurphyB.McKernanC.LawlerC.ReillyP.MessamL. L. M.CollinsD.et al. (2022). A qualitative exploration of challenges and opportunities for dog welfare in Ireland post COVID-19, as perceived by dog welfare organisations. Animals12, 3289. doi: 10.3390/ani12233289
78
National Health and Medical Research Council (2025a). Guide for assessing research involving Artificial Intelligence, Machine Learning and Large Language Model Technology (Canberra, Australia: National Health and Medical Research Council).
79
National Health and Medical Research Council (2025b). National statement on ethical conduct in human research (Canberra, Australia: National Health and Medical Research Council, Australian Research Council and Universities Australia).
80
NissenbaumH. (2004). Privacy as contextual integrity. Wash. L Rev.79, 119. doi: 10.2139/ssrn.3244876
81
NurseA. (2016). Beyond the property debate: animal welfare as a public good. Contemp. Justice Rev.19, 174–187. doi: 10.1080/10282580.2016.1169699
82
OrmandyE. H.WearyD. M.CvekK.FisherM.HerrmannK.Hobson-WestP.et al. (2019). Animal research, accountability, openness and public engagement: report from an international expert forum Vol. 9 ( Animals (Basel).
83
PalmerA. L.RogersC. W.StaffordK. J.GalA.BolwellC. F. (2021). Risk-factors for soft-tissue injuries, lacerations and fractures during racing in greyhounds in New Zealand. Front. Vet. Sci.8. doi: 10.3389/fvets.2021.737146
84
Parliament of Victoria (2025). Policy costing: Shut down the greyhound racing industry in Victoria. Victorian Parliamentary Budget Office. Available online at: https://pbo.vic.gov.au/response/6841.
85
PetoR.PetoJ. (1972). Asymptotically efficient rank invariant test procedures. J. R. Stat. Society: Ser. A. (General)135, 185–198. doi: 10.2307/2344317
86
RadanlievP. (2025). AI ethics: Integrating transparency, fairness, and privacy in AI development. Appl. Artif. Intell.39, 2463722. doi: 10.1080/08839514.2025.2463722
87
R Core Team (2026). R: A language and environment for statistical computing Vol. 4.6.0 (Vienna, Austria: R Foundation for Statistical Computing).
88
RhahlaM.AllegueS.AbdellatifT. (2021). Guidelines for GDPR compliance in Big Data systems. J. Inf. Secur. Appl.61, 102896. doi: 10.1016/j.jisa.2021.102896
89
RodriguezA.WilliamsL. J.LewisS. C.SinclairP.EldridgeS.JacksonT.et al. (2025). Evaluating re-identification risks scores in publicly available clinical trial datasets: Insights and implications. Clin. Trials22, 649–666. doi: 10.1177/17407745251356423
90
Roszkowska-MenkesM.AluchnaM.KamińskiB. (2024). True transparency or mere decoupling? The study of selective disclosure in sustainability reporting. Crit. Perspect. Accounting98, 102700. doi: 10.1016/j.cpa.2023.102700
91
SandveG. K.NekrutenkoA.TaylorJ.HovigE. (2013). Ten simple rules for reproducible computational research. PloS Comput. Biol.9, e1003285. doi: 10.1371/journal.pcbi.1003285
92
Scottish Animal Welfare Commission (2023). Report on the welfare of greyhounds used for racing in Scotland by the Scottish Animal Welfare Commission (Edinburgh, Scotland: Scottish Government).
93
StevensE. G.BakerT.LewisN. (2022). Dealing with sentient surplus: A moral economy of greyhound rehoming. Environ. Plann. E: Nat. Space5, 2033–2051. doi: 10.1177/25148486211054843
94
StoudtS.JerniteY.MarshallB.MarwickB.SharanM.WhitakerK.et al. (2024). Ten simple rules for building and maintaining a responsible data science workflow. PloS Comput. Biol.20, e1012232. doi: 10.1371/journal.pcbi.1012232
95
TenopirC.AllardS.DouglassK.AydinogluA. U.WuL.ReadE.et al. (2011). Data sharing by scientists: practices and perceptions. PloS One6, e21101. doi: 10.1371/journal.pone.0021101
96
The Greyhound Board of Great Britain (2026). With multiple kennel inspections each year, racing greyhounds receive far higher protection than domestic pets. The Greyhound Board of Great Britain. With multiple kennel inspections each year, racing greyhounds receive far higher protection than domestic pets.
97
ThomasJ. B.AdamsN. J.FarnworthM. J. (2017). Characteristics of ex-racing greyhounds in New Zealand and their impact on re-homing. Anim. Welfare26, 345–354. doi: 10.7120/09627286.26.3.345
98
TomczakM.TomczakE. (2014). The need to report effect size estimates revisited. An overview of some recommended measures of effect size. Trends Sport Sci.21, 19–25.
99
UK Parliament (2025). Government response to Petition 731016: Ban greyhound racing. UK Government and Parliament. Available online at: https://petition.parliament.uk/petitions/731016.
100
VergneT.Del Rio VilasV. J.CameronA.DufourB.GrosboisV. (2015). Capture-recapture approaches and the surveillance of livestock diseases: A review. Prev. Vet. Med.120, 253–264. doi: 10.1016/j.prevetmed.2015.04.003
101
VilleneuveJ.-P.HeideM.MugelliniG. (2025). Understanding transparency failures: When sunshine causes sunburns. SAGE Open15, 21582440251352638. doi: 10.1177/21582440251352638
102
Welsh Parliament (2025). Prohibition of Greyhound Racing (Wales) Bill: Stage 1 Report. Culture, Communications, Welsh Language, Sport and International Relations Committee, Senedd Commission. Available online at: https://laiddocuments.senedd.wales/cr-ld17610-en.pdf.
103
WilkinsonM. D.DumontierM.AalbersbergI. J.AppletonG.AxtonM.BaakA.et al. (2016). The FAIR Guiding Principles for scientific data management and stewardship. Sci. Data3, 160018. doi: 10.1038/sdata.2016.18
Summary
Keywords
animal welfare, artificial intelligence, data governance, dogs, ethics, greyhound racing, social license to operate, transparency
Citation
Cobb ML and Coghlan S (2026) Using AI agents to assemble population-level data for visibility and animal welfare insights: a case study of greyhound racing in the UK. Front. Anim. Sci. 7:1868726. doi: 10.3389/fanim.2026.1868726
Received
29 April 2026
Revised
21 May 2026
Accepted
22 May 2026
Published
15 June 2026
Volume
7 - 2026
Edited by
Christine Janet Nicol, Royal Veterinary College (RVC), United Kingdom
Reviewed by
Ana Maria Quessada, Universidade Paranaense, Brazil
Geert De Meyer, Mars Petcare, United Kingdom
Updates
Copyright
© 2026 Cobb and Coghlan.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Mia L. Cobb, mia.cobb@unimelb.edu.au
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.