Exploring Patent Networks Using U.S. Patent and Trademark Office Patent Records

Project Summary

This data expedition explores the local (ego) patent citation networks of three hybrid vehicle-related patents. The concept of patent citations and technological development is a core theme in innovation and entrepreneurship, and the purpose of these network explorations is to both quantitatively and visually assess how innovations are connected and what these connections mean for the focal innovations and the technologies that draw on those patents in the future. The expedition was incorporated as part of the Sociology of Entrepreneurship class, where students are thinking about the emergence and diffusion of innovations.

Themes and Categories

Graduate Students: Josh Bruce (joshua.bruce@duke.edu) and Molly Copeland (molly.copeland@duke.edu)

Faculty: Dr. Martin Ruef

Course: Sociology of Entrepreneurship (SOCIOL/MMS 359)

Guiding Questions

The main research question students answered with this expedition was: What are the characteristics of innovation networks? To address this, how can we conceptualize and visualize patent citations as evidence of innovation network structures? Then, how can we characterize features of innovation ego-networks in patents with common ego-network measures (such as size or density)?

To answer these questions, groups of students worked with different patent ego-networks in the data that are relevant to broader topics they cover in the class and class projects. Students adapted visualizations to reflect characteristics and attributes determined to be of interest within the data (e.g., patent technology classes), chose the descriptive measures that provide insight into broader questions in other innovation projects in the class, and groups of students compared descriptive measures across networks using common network measures. Additionally, after constructing the ego-networks, students brainstormed potential pitfalls of conducting and reporting descriptive network analysis (e.g., consistency in visualization, comparing across networks, interpreting network measures, coding and R-specific issues that commonly arise).

The specific techniques and teaching goals include:

  • Introduction to R and RStudio (with a focus on the network-specific package igraph)
  • Introduction to RStudio workflow components (e.g., installing and loading packages, directories and workflow, what is coding and why do we do it)
  • Introduction to network basics
    • Conceptually, we discuss:
      • What is a network?
      • How can citations make a network (i.e., ego-networks of citation) and what other data could be conceptualized as networks?
      • What do these networks tell us?
    • Practically:
      • Identifying network components (e.g., nodes, edges, attributes)
      • Constructing network objects
      • Visualizing networks
      • Calculating typical network descriptive measures
  • Meaningful interpretation of networks related to innovation

Below is an example of the network students learned to create, based on a single focal patent in the center of the network. Nodes are sized according to their number of technological claims and shaded by their primary technology class.

Example of network

The Dataset

The data for this Expedition consist of patent records from the US Patent and Trademark Office (USPTO). The raw data are individual patents granted by the USPTO, which include information on the nature of the technology being patented, the year the patent is granted, the patents each focal patent cites as “prior art,” and numerous other data points. The version of the data used for this Expedition is publicly available from the PatentsView project (www.patentsview.org), a joint effort between the USPTO, American Institutes of Research, NYU, UC Berkeley, and other stakeholders. PatentsView.org exists to make the bulk USPTO data files accessible for research and practitioner use.

The dataset used in the Expedition is a small subset of the total patent database. It was collected by identifying three focal patents (the egos in the example networks used by students) and then all patents that cited those three focal patents up to 2017. The citations among these patents were also collected, creating three distinct patent ego-networks. For both focal and citing patents, we have metadata on the primary technology class, patent title, abstract, number of claims, and year granted by the USPTO.

Course Materials

Networks of Innovation (Powerpoint presentation)

Patent Nets Code.R







Related Projects

A large and growing trove of patient, clinical, and organizational data is collected as a part of the “Help Desk” program at Durham’s Lincoln Community Health Center. Help Desk is a group of student volunteers who connect with patients over the phone and help them navigate to community resources (like food assistance programs, legal aid, or employment centers). Data-driven approaches to identifying service gaps, understanding the patient population, and uncovering unseen trends are important for improving patient health and advocating for the necessity of these resources. Disparities in food security, economic stability, education, neighborhood and physical environment, community and social context, and access to the healthcare system are crucial social determinants of health, which studies indicate account for nearly 70% of all health outcomes.

We led a 75-minute class session for the Marine Mammals course at the Duke University Marine Lab that introduced students to strengths and challenges of using aerial imagery to survey wildlife populations, and the growing use of machine learning to address these "big data" tasks.

Most phenomena that data scientists seek to analyze are either spatially or temporally correlated. Examples of spatial and temporal correlation include political elections, contaminant transfer, disease spread, housing market, and the weather. A question of interest is how to incorporate the spatial correlation information into modeling such phenomena.


In this project, we focus on the impact of environmental attributes (such as greenness, tree cover, temperature, etc.) along with other socio-demographics and home characteristics on housing prices by developing a model that takes into account the spatial autocorrelation of the response variable. To this aim, we introduce a test to diagnose spatial autocorrelation and explain how to integrate spatial autocorrelation into a regression model



In this data exploration, students are provided with data collected from remote sensing, census, and Zillow sources. Students are tasked with conducting a regression analysis of real-estate estimates against environmental amenities and other control variables which may or may not include the spatial autocorrelation information.