Tuesday, February 26, 2008

The Oxford Text Archive

The Oxford Text Archive was founded more than 30 years ago and is currently a depository of more than 2000 electronic works of literature, some classic reference books (such as the Bible and dictionaries), and "linguistic corpora", in 20 languages.

Selection Principles
OTA makes it clear that they do not do the digitizing themselves. Rather, they rely on their contributors to follow the contribution and digitization guidelines and standards provided. Though up until now the selection criteria have been quite vague, in the future, "priority will be given to digital resources of interest to those working in the literary and linguistic disciplines (including modern and ancient languages), whilst continuing to take materials from any literary genre, period, or language. Other disciplines will be supported on a best-effort basis."

Object Characteristics
OTA has very clunky design and functionality. Texts were available only as "downloads" -- and the only format options were TEI Header Information (SGML), ASCII Text, or DOS ASCII Text. Before one downloaded a file, one had to click on the Terms & Conditions statement, which stated that you would not duplicate or sell the text, and that it was only to be used for educational or scholarly purposes, and then you had to enter your email address. After going through all this and clicking "download", it didn't download after all but opened up on the screen. The text wasn't in a particularly readable form -- the text was squished on the left side and was full of unicode in place of certain punctuation. There seemed to be no way to actually save or download the text. Perhaps the SGML search mechanism is tied to the DTD associated with it and therefore downloading it would negate that aspect of it? I'm not sure. If you clicked the request button, this took you to a page to request a hard copy version, which had to be returned to OTA when you were finished with it. How quaint. I wonder if that system actually works!

Metadata
As the texts were SGML tagged, everything was full-text searchable by numerous TEI headings. The search functionality accommodates advanced Boolean operators.

Audience
This archive was intended purely for scholars and researchers of literature, linguistics, history and the like, and this is reflected in OTA's quite non-user-friendly interface (tons of redundant clicking of tiny boxes and links!) and architecture. I doubt very much that a lay person, one, would find this archive in the first place and two, assuming she did, would have the patience or the know-how to effectively use it. This is definitely not a site that one would utilize for some light reading before bed.

No comments: