acl-org/acl-anthology

Extract abstracts from PDF

Open

#395 opened on Jun 6, 2019

View on GitHub
 (28 comments) (2 reactions) (1 assignee)Python (386 forks)github user discovery
enhancementhelp wanted

Repository metrics

Stars
 (733 stars)
PR merge metrics
 (PR metrics pending)

Description

The anthology currently only shows the abstracts if there is an authoritative version in the XML. It would be nice if we could scrape the PDF using some off-the-shelf software to extract the abstracts and dump them into a different file (to not tamper with handcrafted information). Having an abstract on the web pages makes quickly searching through literature much faster.

Contributor guide