P6, MVP, and Beta Perseus oh my, or Keeping the Flame of P4 Alive
Background
For almost a year now the Perseus development team has been hard at work on Minimum Viable Perseus (which we redundantly call Perseus MVP), now viewable at: https://beta.perseus.tufts.edu.
Perseus MVP builds upon decades of work by Lisa Cerrato, Alison Babeu
and others, creating and curating structured data that can be then
recast into different formats. David A. Smith built the first web-based
version of the Perseus Digital Library in 1995. In 2003, David Mimno
created the Perseus “Hopper,” the system which has, in various forms,
served audiences for decades – the Hopper is older than many (if not
most) current undergraduates. Developers such as Gabe Weaver, Adrian
Packel, Rashmi Singhal and Bridget Almas maintained and enhanced the
Hopper for a decade.
Further, years of research and development
on the Scaife Viewer by Jake Wegner and James Tauber have provided the
foundation for Perseus MVP. Clifford Wulfman took the lead in designing a
Minimum Viable Perseus. Crucially, Charles Pletcher with his pioneering
work on the Ajax Multicommentary
and the New Alexandria Commentary Platform, demonstrated that it was
actually possible to build a complex philological system on a
conservative, serverless architecture. Charles has been responsible for
implementing Perseus MVP, drawing upon Cliff’s outline and his own
experiences.
MVP has been greatly inspired by years of
requests from dedicated users whose devotion to the current version of
the Perseus Digital Library (P4) is matched only by the volume of email
we receive to Perseus webmaster whenever P4 goes down. For over 20
years, P4 has served as the flagship digital library for Perseus’s
various collections. For a number of years now it has not been possible
to update anything on the P4 site due to aging infrastructure and
complex software dependencies. In addition, age made it vulnerable to
scraping by LLMs (Large Language Models), which has led to numerous
outages and the dreaded “Error 503 Backend fetch failed.” While we have
made various efforts to mitigate this vulnerability it has become
increasingly clear it is not possible to easily maintain a robust P4,
especially in the face of aggressive and what increasingly feels like
nonstop LLM scraping.
The Design and Logic Behind Perseus MVP
Development on P6 has come in fits and
starts depending on project funding and staffing (for more on the
history of Perseus and its different instances, please see our recent
article here).
MVP provides a structurally conservative foundation to which we can add
new services (such as support for treebanks and translation
alignments). We had planned to use the Scaife Viewer as the foundation
for Perseus 6 and were able to create a working prototype for next
generation Perseus features such as treebanks, alignments and bilingual
searching. But the Scaife Viewer was developed almost ten years ago and
included versions of older JavaScript libraries that could not be
trivially replaced. Its architecture also depended upon the host
institution running a database. While computation has become less
expensive and the per-cycle cost of managing server side databases is
certainly much lower than it was when David A. Smith created the first
web version of Perseus in 1995, swarming bots in the age of LLMs can,
unchecked, swamp any normal system and mitigating their impact requires
ongoing adjustments.
Perseus MVP does not depend on the host
managing a database. Instead, Perseus MVP compiles the TEI corpora
maintained by Perseus into static HTML pages that are delivered to users
from a file system via a simple static web server. All computation
takes place in the reader’s browser. In addition to providing a
functional replacement for P4, MVP has been designed to not impose
technical burdens on maintainers and future developers. This design also
seeks to establish a foundation to which treebanks, translation
alignments and other new features developed for the Scaife Viewer and Beyond Translation may be included.
The first version of Perseus MVP aims to
provide core features familiar from P4. One advantage of Perseus MVP is
that the data is new and, in most cases, more extensive than that
available in P4. There are still texts that need to be updated but LLMs
have made it easier to update the format and metadata for XML (with a
human in the loop) and we hope to have straggler texts updated as work
continues. Readers will see a growing body of new content. And, of
course, where P4 has been a frozen library, Perseus MVP will include the
latest updates and corrections.
Perseus MVP does not (yet) support all
features familiar from P4 (i.e., that’s the “minimum viable” part).
Features not yet implemented include:
Site search, i.e., look for all documents in Perseus that contain the word “Thucydides.”
Greek and Latin word searching: for now, we recommend that users take advantage of Perseus under Philologic.
This version of Perseus is not quite the same as what we have (though
it is close) but it does an excellent job of providing classic search
services.
Cross-references: This needs to wait until we have
more of the various reference works from P4. Our initial focus is
including source texts, translations and core dictionaries.
URL mapping from P4 and the Scaife Viewer: Before
we retire P4 and the Scaife Viewer, we will produce a system that
converts URLs to P4 and Scaife into URLs that point to the appropriate
pages in Perseus MVP. The goal is to maintain, insofar as possible,
backward compatibility with the vast body of links into Perseus that
have evolved over decades.
New Classes of Content Found in Perseus MVP
Perseus MVP, however, also contains a
growing body of content that is available neither in P4 nor in Scaife.
Optical Character Recognition (OCR) and LLMs have allowed us to create
draft versions of complex documents such as critical editions,
commentaries and dictionaries. Going forward there will be three core
classes of content.
Curated Digital Resources-A core of digital
resources that we have been able to review and curate. Given the need to
make substantial amounts of data available, our goal is to bring these
resources to a reasonable level and then augment them as members of the
community identify issues and/or offer corrections and enhancements.
OCR Drafts-A larger body of documents with rich TEI
XML markup that we have produced, with minimal manual intervention by
using LLMs to add TEI XML. We offer these on an as-is basis with the
label “OCR draft.” Over time, some of these will be curated by Perseus
staff or by members of the community, but we make them available now for
those who wish to use them. We are very much in the process of
understanding how to use rapidly improving systems to create these
documents but we feel that they have reached a point where we find them
useful and hope that others will as well.
For examples of “OCR draft” resources, look at the contents for
Homer, Aeschylus, Sophocles. To filter out these resources, you can
click on a search filter “Status” on the Collections/Texts screen and select “Curated” vs. “Experimental”.
Raw textual data/Massive corpora-We now have access
to six hundred million words of Classical Greek and billions of words
of Classical Latin. Powerful as new systems are, we simply do not have
the computational resources to apply LLMs to textual data at this scale –
the LLMs are too slow or (if we were to work with scalable commercial
enterprises) too expensive to apply. We need to develop ways to make
this content visible and accessible. The traditional method is to search
and analyze the uncorrected, automatically produced text and then to
display the pages (image front processing) and this will probably be our
approach for the foreseeable future.
And introducing Perseus MVP/Beta Perseus!
Home Page
The design has been heavily influenced by
P4 and seeks to ultimately create a feature for feature duplication of
the P4 reading experience. The Home page still includes
the familiar navigation bar above, news and updates (found on this
trusty blog), and links to Art & Archaeology images at their new
online home.
Figure 1: Homepage for Beta Perseus
Collections/Texts Page
The Collections/Texts
page below does not include links to all of the collections currently
found in P4. Beta Perseus does contain links to all the public domain
works in P4 Greek, the Open Greek and Latin (OGL) First Thousand Years of Greek collection, Latin from both P4 and OGL, and initial Hebrew and Italian collections. The default Collections/Texts page view opens with all Collections and authors fully expanded, with the Greek
collection at the top (illustrated below). The user can then scroll
down through a fully expanded author list for each collection. Clicking
on the downward facing arrow to the left of a Collection (Greek) or
an author name (e.g. Achilles Tatius) will minimize the collection or a
given author. The Italian, Hebrew and Latin collections can be found
by scrolling down (not illustrated).
There are also two search boxes on the
left. The top level search box allows the user to search for a specific
title, author or editor and produces a drop down list under the box. The
second search box is an Author filter search that will produce a
filtered list of search results on the right based on the author search.
The current Collections/Texts page is very much a work
in progress and the current visual below could change in the near
future. Due to the addition of a large number of “OCR draft” resources, Perseus and First1KGreek editions are labeled to assist the user in selecting a text.
Figure 2: Current Collections Page in Beta Perseus
Reading Texts in Beta Perseus
The reading interface in Beta Perseus
offers an experience similar to P4, although some features are still
under development. As illustrated below, the Table of Contents
navigation can be found on the left, the selected study text in the
center, and related texts and tools on the right (e.g. notes and
commentaries will appear here in the future). Currently, the user can
expand the Translations tab on the right to choose the available translation for Herodotus Histories. Clicking on the translation name in the panel loads the related passage from the translation and clicking on focus (not illustrated) will load the entire translation in the central reading panel.
Figure 3: Initial reading environment for Beta Perseus-Greek text of Herodotus Histories
As noted above, while some texts will have
commentaries and notes available, not all texts have all of their
related reference works from P4 available yet, but work to add them is
ongoing.
Looking up Words and the Word Study Tool in Beta Perseus
One key feature in Beta Perseus is the ability to look up a word from within a text by clicking on it. The Word Study Tool has been updated with new data first published in the Scaife Viewer.
For example, clicking on ἱστορίης
(first image below, selected word in text is underlined) launches a
search with the result illustrated in the second image below.
Figure 4: Clicking on a word in Beta Perseus
Figure 5: Entry page for ἱστορίης in Beta Perseuswith LSJ and Middle Liddell entries expanded.
The entry page for this word includes
short definitions provided courtesy of Logeion and entries from relevant
lexicons (LSJ and Middle Lidell for Greek and Lewis & Short for
Latin). While the citation abbreviations are not yet linked to their
instances in other texts as in P4, this feature will become available as
work progresses.
Let us know what you think!
We would like to ask all those who love P4
to help us make Beta Perseus better, please use it, test it, and most
importantly tell us about your experiences with it! Email us at perseus_webmaster@tufts.edu
Aside from the features, mentioned above, that we have not yet implemented, known issues include:
The majority of reference works and commentaries found in P4 are not
yet available in Beta Perseus and linked to their source texts-but work
on this is progressing.
Broken links and non-existent pages. Many content pages such as Research/About/Help and other pages have yet to be written.
This is still very much a work in progress,
we wanted to let our dedicated user community know that we hear their
frustration with how this new age of AI is impacting the ability to use
P4 and we are exploring solutions such as Beta Perseus to try and find a
sustainable way forward.
The AWOL Index: The bibliographic data presented herein has been programmatically extracted from the content of AWOL - The Ancient World Online (ISSN 2156-2253) and formatted in accordance with a structured data model.
AWOL is a project of Charles E. Jones, Tombros Librarian for Classics and Humanities at the Pattee Library, Penn State University
AWOL began with a series of entries under the heading AWOL on the Ancient World Bloggers Group Blog. I moved it to its own space here beginning in 2009.
The primary focus of the project is notice and comment on open access material relating to the ancient world, but I will also include other kinds of networked information as it comes available.
The ancient world is conceived here as it is at the Institute for the Study of the Ancient World at New York University, my academic home at the time AWOL was launched. That is, from the Pillars of Hercules to the Pacific, from the beginnings of human habitation to the late antique / early Islamic period.
AWOL is the successor to Abzu, a guide to networked open access data relevant to the study and public presentation of the Ancient Near East and the Ancient Mediterranean world, founded at the Oriental Institute, University of Chicago in 1994. Together they represent the longest sustained effort to map the development of open digital scholarship in any discipline.
No comments:
Post a Comment