JavaScript is required to use this site. Please enable JavaScript in your browser settings.

September 29, 2026

GPTKB 2.0: Making Large Language Model Knowledge Searchable and Auditable

GPTKB 2.0: Making Large Language Model Knowledge Searchable and Auditable
Research

Researchers at ScaDS.AI Dresden/Leipzig have released GPTKB 2.0, a large-scale knowledge base built entirely from a large language model (LLM), GPT-5.1. The knowledge base contains 38 million factual beliefs with automated entity disambiguation, offering unprecedented introspection into the beliefs of frontier AI models.

About GPTKB 2.0

How much structured knowledge does a large language model contain — and how can this knowledge be made accessible and human-readable? GPTKB 2.0, developed at ScaDS.AI Dresden/Leipzig, addresses this question with a new large-scale knowledge base constructed entirely from a large language model.

GPTKB 2.0 contains 38 million factual triples about 2.3 million entities, extracted for $7000 from the GPT-5.1 LLM. A triple represents a simple, machine-readable statement such as “Marie Curie — occupation — physicist.” Such structures allow users to browse entities, search facts, download data, and formulate their own analytics queries.

The key advance in version 2.0 is automated entity disambiguation and deduplication. Large language models may refer to the same entity using different names or surface forms—for example, “USA,” “United States,” and “United States of America.” GPTKB 2.0 identifies and resolves such variants, reducing duplication and helping users distinguish genuinely different entities with similar names. The result is a substantially cleaner and more reliable resource than earlier versions.

This process also makes the knowledge base more transparent. For each fact, the GPTKB browser exposes a disambiguation chain from the original surface form to the canonical entity and its provenance batch. Researchers can therefore not only inspect what information is contained in the resource, but also trace how it was consolidated.

Evaluation results indicate a high quality of the extracted knowledge. GPTKB 2.0 reaches 96% entity precision, and resolves 95% of all disambiguations correctly. These improvements are particularly important when LLM-generated knowledge is used for data analysis, knowledge exploration, or downstream AI applications.

Screenshot. Research results for "Dresden" in GPTKB 2.0.

Publications

The project is accompanied by two scientific publications:

  • Yujia Hu, Tuan-Phong Nguyen, and ›Simon Razniewski. “Direct Construction of Disambiguated Knowledge Bases from Large Language Models.” AACL 2026 / arXiv-2608.03729.
  • Yujia Hu, Tuan-Phong Nguyen, and Simon Razniewski. “GPTKB 2.0: Browsing, Querying, and Auditing a Disambiguated LLM-Derived Knowledge Base.” EMNLP 2026 / arXiv-2608.06992.

They will be presented at EMNLP in Budapest at the end of October, and at AACL in Hengqin at the beginning of November.

Access

GPTKB 2.0 is openly accessible here. Visitors can explore entities, inspect facts, run queries, and download the dataset. Further information on the project is available here.

Previous Entry Back to Overview Next Entry

Welcoming Dr. Marina Litvak as 3rd Ada Lovelace Distinguished Research Fellow

ScaDS.AI Dresden/Leipzig

ScaDS.AI Dresden/Leipzig welcomes Dr. Marina Litvak as the third Ada Lovelace Distinguished Research Fellow. […]

ScaDS.AI Dresden/Leipzig at KR 2026 in Lisbon, Portugal

Research

From July 20-23, 2026, researchers from our Knowledge Representation and Methods Group attended the […]

PriME-LLM: Privacy Preserving Masking for LLM Interaction

Research

At the beginning of the year, researchers from Leipzig University and ScaDS.AI Dresden/Leipzig were […]

ACL 2026: Outstanding Paper Award for Zhan Qu and Michael Färber

Events

From July 2–7, 2026, the 64th Annual Meeting of the Association for Computational Linguistics […]
funded by:
Gefördert vom Bundesministerium für Bildung und Forschung.
Gefördert vom Freistaat Sachsen.