Cataloging Confusion

For class I just read the book Catalog It!: A Guide to Cataloging School Library Materials by Allison G. Kaplan.  I’m by no means an organized person in my personal life, but oh how I want to be. As an archaeologist I spent much time learning and thinking about how to document and arrange little objects with lots of information attached to them. In archaeology there isn’t really  a standard that translates easily from site to site, but whatever system and nomenclature you’re using to catalog finds it is imperative that the standards are employed consistently so that you can easily group and analyze different types of data. This means that I very much understand the need to do cataloging “correctly”, but in my mind “correctly” mostly means “internally consistent and captures as much information as possible even if you don’t think you need it, but in a way that makes sense.” 

Fun aside: at one of the projects I worked on, a few museum workers had digitized all of the entirely handwritten artifact catalogs from excavations in the 1960s. In an attempt at historical fidelity (laudable), they transcribed every single spelling mistake and, even more maddening, treated every  break in a line as a new cell in excel. This made the document completely unsortable and fragile. I spent weeks standardizing and copy-pasting, so that the excel documents were usable spreadsheets of archaeological data rather than artifacts of mid-century formatting choices.  I guess my point is, as long as a system actually works and lets you find what you’re looking for, that sounds great to me.  If it’s an actual database that lets you link and relate different fields to each other? Oh brother, now we’re talking.

Library cataloging, I’m learning, is somehow both way more standardized and way more arbitrary than I originally thought. Most of what I’ve learned about library cataloging up to this point  has been on-the-job, as the result of generous mentors and messing around.  It was great to have a systematic overview of all the intricacies I didn’t even know to ask about. 

 So, in no particular order, here are the things that made me go “huh!” or “ok…” or “but what about…?”

Historical eccentricities

I know I probably “shouldn’t”, but I kind of love how many things are holdover from the card catalog days. 

The factoid that has stayed with me the most, for better or worse, answers a question I’ve had since I started working in a library: why are the words in titles not all capitalized? The answer - because utilizing the shift key on a manual typewriter can be sticky and slow, especially when you’re typing hundreds of cards.

Even MARC records, the standard way of encoding catalog information, is mostly still used because that’s what we’ve been using.  In a world of so many different kinds of databases and increasingly AI-ed search engines, it’s both fun and a little confusing to be using a format that was created in-house for the Library Congress, and that reflects the efficiency needs of earlier computing. 

That being said, the quirks and formats that have held on for decades are why we can do “copy cataloging” - where we can easily get records from other libraries and publishers that will interact with a catalog system in the same way, regardless of the item. It’s why we don’t have to build every record from scratch, for which I am grateful. 

MARC is even more confusing than I thought

]So, I have looked at MARC records before, and they make my head swim. But, I figured, there must be a coding language here that I just don’t understand. Surely it becomes more simple once you know what all the symbols and letters mean. 

Turns out…only kinda? 

Subfields are marked with a $ symbol. $a, $b, etc.  mean different things depending on what field they are in. Ok, that’s annoying but makes sense, you can at least look up specific fields to see what the different labeled subfields correspond to. But there are constantly little one-off rules that apply only to specific contexts. Let’s look at subject headings again. 

As mentioned, subject headings will come from a few different places, but all highly standardized. In order to tell which list a heading is from, the MARC record will include an indicator number. All MARC records include a space for two indicators, but again, what those numbers are indicating  depends on what field you’re looking at. For the 600’s (the topic fields), it’s the second indicator number that corresponds to the original list where the heading was found. Library of Congress Headings are indicated with _0, the LC Annotate Card (children’s) Headings are indicated with _1, and the Sears Headings are indicated with _7 and then ALSO $2sears. I gather the _7 indicator just says “this is from another controlled list,” and then subfield $2 is where you put the name of the specific list.  When I first read this in Kaplan, I was stumped. Looking it through a second time, it makes more sense. Is this how all of MARC is? A winding maze that makes no sense until it does?

What still truly baffles me are all the punctuation standards, which feel entirely random, even if they aren’t. For example, field 245 refers to the title and “statement of responsibility” (basically, the creators). There are 12 valid subfields, the most common of which are $a, $h, $b, and $c. Good so far. 

However, each of the subfields have their own punctuation standards.

“The $h (if there is one) is separated from the $a with a space. The $b is separated from the $a (or $h if there is one) by a space and a colon. The $c, if is separated from the rest by a space and slash (/)” (Kaplan 2016, 108). 

I’m sorry, what?! Why does $b need a colon? Why does $c need a slash? I don’t think this is true for subfields with the same labels in other fields…is it?

It feels like learning a language that has evolved with grammatical rules that include many seemingly random exceptions. But unlike, say,  English, which has grown,  morphed, and incorporated other linguistic influences over the past 1500 years, MARC has been around for less than 100. And computers, as we know, now have the capacity to understand many new kinds of coding languages.

It’s hard for me to understand why we are still so attached to this architecture, although I understand that change is hard, especially when we’re dealing with a network of institutions that want to share information. Still, in a world where we need to explicitly teach students that not all search engines will understand their spelling mistakes, and where Google insists on giving us AI answers, surely we can figure out something a little more intuitive. Of course, the issue isn’t “figuring it out”, the issue is implementing it.

The end of the book has a section on BIBFRAME, a new program created by the Library of Congress with hopes to replace MARC. It is “designed to integrate with and engage in the wider information community” (Kaplan 2016, 181), and operate within the context of the internet. Kaplan says that “BIBFRAME is probably a few years off in terms of full implementation, and it most certainly will hit university and larger institutions first” (Kaplan 2016, 182). Ten years later and our school library, at least, is still reliant on an OPAC that runs on MARC. I’m curious if experts are still expecting to see this shift reflected in smaller libraries, and if so, what kind of timeline we’re looking at. 

Why subject headings are so weird

My district uses Follett Destiny and subject headings have been very opaque to me.  On the one hand, being able to click a button and see all the books with the same headings is a useful way to explore the catalog. On the other hand, they have been applied inconsistently and often seem unusably broad. Sometimes they have the format listed (“juvenile literature”, more often than not) , sometimes they do not. Sometimes we get records with tons of slightly different subject headings, sometimes we get records with very few. When adding subject headings myself, I’ve been at a bit of a loss. Before now, I’ve not known where these headings really came from. I knew they would be linked to other titles in the collection with the same headings, so almost always just pick from already existing headings. 

Now I know that these subject headings are supposed to be coming from an external list. The three big examples Kaplan gives are, as mentioned above, the Library of Congress Headings, the LC Annotate Card (children’s) Headings, and the Sears Headings. 

Curious as to where most of our headings were coming from, and also how Follett handles manual entry of subject headings (in the “easy edit” view, not the MARC view), I did a few experiments. First, I noticed that there are a lot of subjects with the indicator of “0”, meaning they came from the Library of Congress Headings list. We also have a lot of subjects that say they are from the Sears list. Take a look at the easy-edit view and MARC view of the book Shut Up, This Is Serious by Carolina Ixta. 

Now look at what happens when I add a heading manually. 

Adding “teen pregnancy” under topical heading and “test” under format, leads to a change in the MARC record that identifies this heading as coming from the Sears list. Uh-oh. It most certainly does not, it comes from my head.

I thought perhaps it identified my input as from Sears because so many other headings for this book were supposedly from that list, so I found another new book that had no pre-existing headings from Sears. 

The same thing happened. So now I know: if I make up a heading, it will alter the MARC record to erroneously tell catalogers that it came from the Sears list. Except wait…I also added the heading Cults - - Juvenile Literature. However, that one I chose from a list of previously used headings, so it had probably been initially added via another record.

Beyond searching through the default settings, I’m not sure how to go about cleaning this up or even beginning to have a system for updating topic headings moving forward. We rely heavily on copy cataloging, and it sure is nice to just import a record, give it a once over, assign it a call number and sublocation, and get the item circulating.

I’m also just not sure we need subject headings that match all the libraries in the country - we need subject headings that let us search for topics that are relevant to our students. How many students are going to go “Ah yes! Bildungsroman! Exactly what I’m looking for, what else we got”. What we do  need is for the topic headings to be internally consistent, not externally consistent. I’m curious about implementing a more systemic tagging system for our own collection. It would be a lot of work, and shouldn’t be in the same field, but it could be really helpful. Once we set it up, as long as we had a documented operating procedure for incoming materials, I think it would be manageable. I’m thinking of something a little more robust than resource lists and collections. Finding out if this is possible  will require more poking around in Follett, since the ability for students to click them would be important. 

It’s probably impossible to make sure your data is entirely consistent (but you should still try?)

This all leads me to the biggest question I had while reading this book (not to mention all the other times I’ve tried to think about cataloging): 

How do I (realistically) keep a catalog both consistent and up to date?

Kaplan says that it’s unrealistic to go back and correct all old records, and I have to agree. It’s one thing to go back and edit a dataset from an archaeological project that will last a few years at most; that might take a long time but at least it’s somewhat contained. A library collection, on the other hand, grows and shifts over the course of decades.  She does say it’s important to know what current cataloging standards are, so that we can make sure any NEW books we’re adding are up to date (even if they come attached with old records). 

I confess I’m pretty stumped on how to systemetize any of this. What needs to be prioritized and codified? What does it look like to be a new school librarian trying to implement consistent cataloging practices for hersef and her staff?

Lingering Thoughts and Questions

So, this blog post is already too long. But, here are some of the other questions and topics my brain has been chewing on since reading Catalog It! 

WEMI - This acronym kind of changed my life. I love having vocabulary to articulate things I’ve been confused about! WEMI stands for Work, Expression, Manifestation, and Item. When does something get its own record, and when do we include it as a copy of the same title (work or expression)?  It’s not clear-cut. I was also relieved to know that I can add multiple ISBNs to MARC records because I have previously been feeling vaguely guilty every time I put the hard back and paperback versions under the same Title. It’s less clear that it’s “correct” to put different editions of, say, the Great Gatsby all together, but especially for core texts that we have many copies of, I am ok with doing a bit of fudging. 

How is my OPAC even working? What changes the MARC record, what changes are just for the Follett interface? This will involve much more messing around, I suspect.

Classification, Dewey, call numbers, and sublocations. This topic definitely needs its own blog post (TBD), but the gist is that it’s all kind of a colonial mess and needs new categories and/or an entire overhaul. On the other hand, the actual system of being able to build decimals based on individual library needs? Pretty cool.

Erica Lockwell

Queer artist and founder of Our Back Pockets.
Likes archaeology, crafts, and cats.

Previous
Previous

Adventures in Koha

Next
Next

Fun & Games in the Library