Data Descriptor: Disambiguation of patent inventors and assignees using high-resolution geolocation data

Morrison Greg<sup>*</sup>; Riccaboni Massimo; Pammolli Fabio

doi:10.1038/sdata.2017.64

摘要

Patent data represent a significant source of information on innovation, knowledge production, and the evolution of technology through networks of citations, co-invention and co-assignment. A major obstacle to extracting useful information from this data is the problem of name disambiguation: linking alternate spellings of individuals or institutions to a single identifier to uniquely determine the parties involved in knowledge production and diffusion. In this paper, we describe a new algorithm that uses high-resolution geolocation to disambiguate both inventors and assignees on about 8.5 million patents found in the European Patent Office (EPO), under the Patent Cooperation Treaty (PCT), and in the US Patent and Trademark Office (USPTO). We show this disambiguation is consistent with a number of ground-truth benchmarks of both assignees and inventors, significantly outperforming the use of undisambiguated names to identify unique entities. A significant benefit of this work is the high quality assignee disambiguation with coverage across the world coupled with an inventor disambiguation (that is competitive with other state of the art approaches) in multiple patent offices.

出版日期2017-5-16

全文

访问全文

收藏分享被引(13) 浏览

更新时间：2021-01-16 11:11

Data Descriptor: Disambiguation of patent inventors and assignees using high-resolution geolocation data

摘要

全文

产品服务

站内浏览

服务支持

联系方式

科研之友