← J·Index — The Journalism AI Index

Methodological Note

How J·Index maps AI adoption in news media — method, sources and limits behind the Journalism AI Index.

Version 1.1 · July 2026 · Katya Gorchinskaya · Data licensed CC BY-NC 4.0

1.Purpose of this note

J·Index positions itself as an evidence base for media and newsroom executives, tech and product leaders, media developers and funders, researchers, journalists and educators, and executives in other industries watching one sector navigate deep disruption. This note documents precisely what the index measures, how the data is collected, verified and transformed, and — crucially for responsible academic or policy use — what it does not measure. It applies the cautions of the research literature on newsroom AI to the actual state of the index, without embellishing what has not been done.

As of July 17, 2026, the index holds 2,002 verified cases across 1,310 organizations in 131 countries, spanning 2012–2026, of which 157 are flagged as controversies.

2.Conceptual framing

The index builds on the empirical tradition begun by Beckett (2019), whose 71-newsroom survey established the pre-generative baseline for AI in journalism, and continued by Beckett and Yaseen (2023) and Simon (2024), who documented how AI "retools, rationalizes and reshapes" news work. Where those studies sample and survey, J·Index inventories: it is an atlas of publicly documented adoption events, closest in spirit to a cartography project rather than a monitor.

Two of its analytical scales are original but grounded:

The five-level Human Involvement scale (Fully Automated → AI-Assisted Human → Collaborative Human–AI → AI-Led with Human Oversight → Human-Only) operationalizes the augmentation-versus-replacement question the field debates (Diakopoulos, 2019; Marconi, 2020) at the level of individual implementations.

3.Unit of analysis and inclusion rules

One row = one publicly documented, source-linked instance of AI adoption by a news organization. "Adoption" covers five event types: implementations (tools and workflows in use or in pilot), formal policies and governance frameworks, content-licensing and technology partnerships with AI companies, landmark litigation, and controversies arising from newsroom AI conduct.

A candidate enters the index only if it clears all three bars:

  1. A named organization — a newsroom, publisher, broadcaster, news agency, journalist-built tool or media-development organization. (Vendors appear only when their newsroom adoption is documented, and are tagged "Tech vendor.")
  2. A specific, concrete implementation — a named tool, workflow, policy or deal doing an identifiable thing. Statements that an organization is "exploring," "investing in" or "committed to responsible" AI do not qualify.
  3. A live, verifiable source — every row carries a source URL a reader can check.

Standing exclusions, applied consistently: commentary and punditry about AI; vendor marketing without a newsroom adopter; an outlet reporting on someone else's AI (coverage is not adoption — a recurring misattribution pattern in AI-assisted research); contested allegations not established by admission, investigation or sanction; and general-purpose AI products without a journalism deployment. The index logs landmark news-media AI litigation but does not attempt to be a complete litigation tracker; dedicated trackers are linked from the dashboard.

4.Structure of a record

Each case carries 18 structured fields: Organization, Country, Region, Use Case, Tool/Technology, Super-Category, Function Type, Human Involvement Level, Adoption Stage, Outcome/Metric, Controversy Flag, Source URL, Funder, Year, Publisher Scale, Year_Quality, Last_Verified and Link_Status. Nine fields draw on closed, controlled vocabularies (regions, macro-domains, function types, involvement levels, stages, flags, funder categories, publisher scales, link statuses); a deterministic validator enforces them, and the full database passes with zero violations as of this version.

Deduplication uses a normalized key — organization plus the opening of the use-case description — backed by content-level review, since the same tool can surface under variant organization names in different sources.

5.Data collection (actual state, July 2026)

The corpus is assembled through five channels, in declining order of volume:

  1. Manual research by the author — including country-by-country sweeps (Eastern Europe, MENA, Sub-Saharan Africa, South/Southeast Asia, China in Chinese-language sources, Latin America, Australia/NZ) and systematic gap-filling of program cohorts (JournalismAI, WAN-IFRA accelerators, CUNY AI Journalism Lab, Lenfest AI Collaborative and others).
  2. Human-guided AI research — research assistants (Claude) working to set parameters, with every resulting candidate verified against a live source before entry.
  3. An autonomous monitoring layer (since June 2026) — a deterministic, token-free feeder polls about 20 trade-press, industry, accelerator and research sources at weekly, monthly and quarterly cadences and surfaces new items as candidates. It gathers; it does not judge. Every candidate passes human review.
  4. Manually outsourced searches to other AI models where they hold an access advantage — for example, Gemini for spoken-word content in YouTube podcasts and videocasts. These results are treated strictly as unverified leads: in one audited batch, roughly three quarters of AI-supplied "findings" carried fabricated or mismatched sources and were discarded. Nothing from this channel enters the index without independent verification.
  5. Community submissions via the public "Add a Missing Entry" form, verified like any other lead.

Sources are classified into 11 types (trade press, industry body/awards, accelerator programs, depositories, outlet self-reports, research/academic, media-development organizations, vendors, PR wires, government/legal, news reports). About 11% of cases rest on outlet self-reporting and under 1% on PR wires; the majority rest on independent trade, industry and academic reporting.

Regional research mappings each anchor multiple entries: the Thomson Foundation/MJRC study of Visegrád-Four newsrooms for Central Europe, DW Akademie's mapping of Ukrainian regional outlets' AI use, IMS's AI, Journalism, and Public Interest Media in Africa, UNESCO's Journalism and Artificial Intelligence in Latin America, the KAS/CINIA South Africa survey and Harb & Arafat's study of Arab newsrooms (see References).

6.Verification pipeline

7.From master table to dashboard

The public dashboard is not hand-drawn: it is generated deterministically from the master table by a single build step. Every chart, map, treemap, stat and Case Explorer entry is recomputed from the current data on each publication, so the visualizations and the underlying rows can never drift apart. The build is reproducible — the same table always produces the same dashboard — and it is version-controlled alongside the data.

Derived dimensions are computed at build time; the master table's 18 fields are never altered. The main derivations, each grounded rather than cosmetic:

Narrative blurbs are data-bound. The short interpretive sentences beside several widgets are generated from the computed figures (leading region and its count, funding shares, stage clustering), not written by hand per release, so they update with the data and cannot silently contradict the charts.

The map is a country choropleth; a small number of geopolitical conventions are applied explicitly (for example, Crimea is rendered within Ukraine), and are documented in the build rather than left implicit.

Curated layers are labelled as such. A few elements are editorial selections rather than computed aggregates — the "curiosities" and "unusual tech stacks" highlights, and the list of reusable open-source tools. Each curated item is pinned to a specific row and guarded by an automatic organization-check: if the underlying table is reordered or a row changes, the build fails loudly rather than mislabelling an example. These layers are clearly presentational, not statistical claims.

Self-test. Before each release the dashboard runs an internal suite (43 checks as of this version) exercising every interactive feature — navigation, filters, every chart's render, the Case Explorer, the stage stepper — so a broken widget is caught prior to publication, not by readers.

8.Known limits and biases

Disclosure bias — the fundamental one. The index records publicly documented adoption. Organizations that communicate about innovation are over-represented; quiet adopters are invisible, and so are quiet abandonments. The index cannot see undisclosed AI use — and, as the 2026 detection controversies showed, no reliable tool exists that could (AI-text detectors measure resemblance to model output, not provenance).

Coverage bias. Sources skew toward the English-language trade press and the award/accelerator circuit (WAN-IFRA, INMA, JournalismAI, ONA and peers), which over-represents award-winners and program cohorts. Targeted sweeps in local languages (including Chinese, Spanish, Portuguese, Arabic, Ukrainian and others) mitigate but do not eliminate this; organizations documented only in less-indexed languages are under-represented.

Currency. A case documented at time T may since have been discontinued. Last_Verified records when the source was last checked, not whether the implementation still operates.

Counts measure documentation, not intensity. An organization with 15 rows versus one with a single row partly reflects communication culture, not necessarily a 15-fold difference in adoption.

Outcomes are mostly self-reported. The Outcome/Metric field preserves claimed results (time saved, traffic, subscriptions) as stated by the organization or its case-study author; the index does not independently audit them.

No probabilistic sampling. Unlike survey research (Beckett, 2019; Beckett & Yaseen, 2023; the Reuters Institute's audience work), this is an inventory of visible actors. It supports questions of the form "what exists, where, documented how" — not prevalence estimates of the form "X% of newsrooms use AI."

Year uncertainty is flagged, not hidden. 36 rows carry no year (tagged Missing); a further small share is tagged Estimated, Approximate or Inferred rather than asserted as confirmed.

9.Classification conventions

10.Descriptive, not evaluative

Presence in the index signals that an adoption event is documented — not that it worked, created value, or meets any ethical or legal standard. The index deliberately includes failures, retractions, union disputes and even AI content farms impersonating local news (flagged as controversies), because the phenomenon of AI in news includes its pathologies. No row is an endorsement; no count is a quality ranking.

11.Update cadence and correction mechanism

The autonomous feeders run on a fixed schedule (weekly Mondays, monthly and quarterly on the 1st, with automatic retry when the host machine is offline at fire time); their candidates are processed in reviewed sessions, typically within days. The dashboard is rebuilt from the master table on every update, with the build date shown on the masthead. Corrections arrive through the public comment form, the submission form and continuous re-verification; erroneous rows are corrected or removed with a documented backup trail. The database is versioned by timestamped backups at every modification.

12.What J·Index is not

13.Roadmap

Planned extensions, stated as intentions rather than accomplished facts: a natural-language query layer over the index; standalone organization pages; a dedicated deals/licensing lens; and continued expansion of the source registry as AI adoption spreads into less-covered markets and languages. Any future composite indicator built on this data (for example, adoption-intensity measures by country or scale) will be documented in a revision of this note before publication, following the sequence: framework → indicator choice → normalization → aggregation → robustness.

References

  1. Beckett, C. (2019). New Powers, New Responsibilities: A Global Survey of Journalism and Artificial Intelligence. LSE Polis / JournalismAI.
  2. Beckett, C., & Yaseen, M. (2023). Generating Change: A Global Survey of What News Organisations Are Doing with AI. LSE JournalismAI.
  3. Brigham, N. G., Gao, C., Kohno, T., Roesner, F., & Mireshghallah, N. (2024). Developing Story: Case Studies of Generative AI's Use in Journalism. University of Washington.
  4. Diakopoulos, N. (2019). Automating the News: How Algorithms Are Rewriting the Media. Harvard University Press.
  5. The Economist (2026, January 29). GenAI v Gen Z (Season 3, Episode 4). The Economist podcasts.
  6. DW Akademie (2025). Mapping Ukrainian Media Outlets' AI Use: How AI Pioneers Are Changing Ukrainian Regional Media. Deutsche Welle Akademie.
  7. Harb, Z., & Arafat, R. (2024). The Adoption of Artificial Intelligence Technologies in Arab Newsrooms: Potentials and Challenges. Emerging Media, 2(4).
  8. International Media Support (2023). AI, Journalism, and Public Interest Media in Africa. IMS, June 2023.
  9. Konrad-Adenauer-Stiftung Media Programme Sub-Saharan Africa / CINIA (2026). Navigating Risks and Rewards: How South African Journalists Use AI in the Newsroom.
  10. Kioko, P. M., Booker, N., Chege, N., & Kimweli, P. (2022). The Adoption of Artificial Intelligence in Newsrooms in Kenya: A Multi-case Study. European Scientific Journal, 18(22), 278.
  11. Kyomugisha, E. (2025). Artificial Intelligence Adoption and Journalistic Practices: A Case Study of Nation Media Group–Uganda (Daily Monitor) (Unpublished thesis). Aga Khan University, Graduate School of Media and Communications.
  12. Li, F., Zhang, L., & Roy, S. K. (2026). What Drives Effective AI Use in the Newsroom? Communication Barriers, Organizational Support, and Journalist Performance in China. Journalism and Media, 7(2), 105.
  13. Marconi, F. (2020). Newsmakers: Artificial Intelligence and the Future of Journalism. Columbia University Press.
  14. Newman, N., et al. (2026). Digital News Report 2026. Reuters Institute for the Study of Journalism, University of Oxford.
  15. Reuters Institute for the Study of Journalism (2025). Generative AI and News Report 2025: How People Think About AI's Role in Journalism and Society.
  16. Rogers, E. M. (2003). Diffusion of Innovations (5th ed.). Free Press.
  17. Roy, N. (2024). AI in the Newsroom case-study series. Online News Association, AI in Journalism Initiative.
  18. Sánchez Esparza, M., & Palella Stracuzzi, S. (2026). AI in the newsroom: A case study of investigative journalists in Spain. Online Journal of Communication and Media Technologies, 16(2), e202614.
  19. Simon, F. M. (2024). Artificial Intelligence in the News: How AI Retools, Rationalizes, and Reshapes Journalism and the News Industry. Tow Center for Digital Journalism, Columbia University.
  20. Thomson Foundation & Media and Journalism Research Center (2024). How Artificial Intelligence Is Changing Media and Journalism in Central Europe: A Study Mapping the Use of AI by Newsrooms in the Czech Republic, Hungary, Poland and Slovakia. Marko, D. (research coordinator), Dragomir, M. (editor). June 2024.
  21. UNESCO (2023). Journalism and Artificial Intelligence in Latin America. UNESCO, Paris (unesdoc PF0000388124).
  22. Wang, Y. (2021). The Application of Artificial Intelligence in Chinese News Media. ICAIIS '21: 2nd International Conference on Artificial Intelligence and Information Systems. ACM.

Corrections and challenges to any row are welcome via the forms on jindex.ai. This note will be versioned; substantive changes to method will increment the version number and be dated.

← Back to the index