Knowledge extraction from Chinese wiki encyclopedias

Zhi Chun Wang*, Zhi Gang Wang, Juan Zi Li, Jeff Z. Pan

*Corresponding author for this work

Research output: Contribution to journalArticlepeer-review

24 Citations (Scopus)


The vision of the Semantic Web is to build a 'Web of data' that enables machines to understand the semantics of information on the Web. The Linked Open Data (LOD) project encourages people and organizations to publish various open data sets as Resource Description Framework (RDF) on the Web, which promotes the development of the Semantic Web. Among various LOD datasets, DBpedia has proved a successful structured knowledge base, and has become the central interlinking-hub of the Web of data in English. However, in the Chinese language, there is little linked data published and linked to DBpedia. This hinders the structured knowledge sharing of both Chinese and cross-lingual resources. This paper deals with an approach for building a large-scale Chinese structured knowledge base from Chinese wiki resources, including Hudong and Baidu Baike. The proposed approach first builds an ontology based on the wiki category system and infoboxes, and then extracts instances from wiki articles. Using Hudong as our source, our approach builds an ontology containing 19 542 concepts and 2381 properties. 802 593 instances are extracted and described using the concepts and properties in the extracted ontology and 62 679 of them are linked to equivalent instances in DBpedia. As from Baidu Baike, our approach builds an ontology containing 299 concepts, 37 object properties, and 5590 data type properties. 1 319 703 instances are extracted from Baidu Baike, and 84 343 of them are linked to instances in DBpedia. We provide RDF dumps and SPARQL endpoint to access the established Chinese knowledge bases. The knowledge bases built using our approach can be used not only in Chinese linked data building, but also in many useful applications of large-scale knowledge bases, such as question-answering and semantic search.

Original languageEnglish
Pages (from-to)268-280
Number of pages13
JournalJournal of Zhejiang University: Science C
Issue number4
Early online date3 Apr 2012
Publication statusPublished - Apr 2012


  • Knowledge base
  • Linked Data
  • Ontology
  • Semantic Web


Dive into the research topics of 'Knowledge extraction from Chinese wiki encyclopedias'. Together they form a unique fingerprint.

Cite this