
One of the main reasons is that patent data are collected to identify the novelty of a invention and to make it public as a new stock of knowledge for further inventors, whatever the applicants. The process of the quality of information is then, first of all, focused on the data regarding scientific and legal information. Thus for instance, every patent has an unique ID number in order to identify perfectly the invention and to be able to establish scientific links between patents. By contrast, no quality check is applied to applicants’ names and address-es, and in some cases (US data for instance) apart from the country even address data are unavailable. Thus identifying (called afterwards disambiguating) existing applicants by name or address, in order to build up a unique identifier for each patenting entity, is not a simple task.