On Alt Text Image Description as Literary Practice
M. Leona Godin Considers the Mundane, Profound Act of Translating a Picture
My first experience with AI image description was both mundane and profound. I was reclining on the bed in our East Village sublet, writing and drinking afternoon coffee—”Prousting,” as my partner Alabaster and I call it—when a new app landed in my inbox. I downloaded it and snapped a picture of the room, crammed with the owner’s things. Seconds later, I received a reply. Like an art historian pointing out details in a Renaissance painting, the AI told me about the red paisley blanket in the foreground, the door with an orange curtain beyond, and above that, a “decorative piece with the Om symbol.”
I couldn’t remember what the Om symbol looked like, so I used the “Ask more” function, and learned that it “resembles a numeral 3 with a tail swooping inwards and a small curve with a dot above it.” A vague image floated before my mind’s eye—not so vivid as the room with its tangible objects—but it made enough of an impression to recall another sense memory. I’ve taken my share of yoga classes that sometimes finish with the chanting of “Om,” and remembered how the sound vibrates bodily tissue as it moves from belly to lips. Maybe this is what translating between the senses and language feels like: a spark that takes on new associations as it travels perceptual byways.
That a machine could “see” and describe was striking, and it reinforced for me the power of description—something I knew well from a lifetime of seeing the world through literature. The sudden access to the visual world was unprecedented yet familiar; it also challenged my identity as a blind person. After all, in my first book, There Plant Eyes, I critiqued ocularcentrism—the often unconscious, sometimes tyrannical sight bias. And here I was, awed by pictures of paisley fabric and an Om symbol.
Though totally blind now, I spent most of my life with varying degrees of sight and remain intensely visual. At some point, a good description turns from being a collection of words to an image in my mind’s eye.
Two perfectly sighted people do not see the same image in the same way.
I had the idea to write about the anonymous blind people in famous photographs way back in 2022—when AI was barely on my radar. My original goal was to find and interpret these images from a blind perspective. The photographers were famous, but little to nothing was known about the subjects—I wanted to nestle them in historical context and lived experience. I also wondered what image description might bring to the photos—even to photography itself.
Alabaster’s description of Paul Strand’s iconic 1916 photo, Blind Woman, was my first glimpse of her. I downloaded the image from the New York Public Library and emailed it to him—sitting over on the couch.
“She’s looking to her left with her left eye open,” he said, “her rather large iris jammed into the left corner and her right eye nearly closed and appearing injured.”
Two perfectly sighted people do not see the same image in the same way. So I got another description—this time from disabled artist Finnegan Shannon, who recognized something familiar in the pale face. “She looks a little like me,” Finnegan told me in an email, and described the wisps of grey or blond hair peeking out from under the black headscarf.
Alabaster saw “a small mole on her left cheek,” and “a slight frown with a touch of anger and longing across her face.”
Finnegan saw “wrinkles at her brow,” but found her face relaxed. “I don’t read a particular expression or emotion from it.”
Perhaps her sign is the most striking feature of the photo. She wears it around her neck. The word “blind” in black capital letters on a white placard against her black dress. Very stark, very high-contrast. Finnegan noted that “it’s maybe the size of a hand but feels really big and loud.”
It appears to be hand-painted—by her own hand or that of another, I’ll never know. But for whose eyes is obvious: the SIGHTED.
In one of her many essays on photography, Susan Sontag notes that “what a photograph is of is always of primary importance.” If the subject of the photograph is “visually ambiguous,” we don’t know how to react “until we know what piece of the world it is.”
But writers thrive in the ambiguous spaces. We understand that description is never neutral. Perhaps this explains the hesitancy when it comes to description as access. Author friends often hand their books to Alabaster so he can describe the covers for me—as if he holds the key to getting things “right.” We always laugh because, though he has to describe stuff a lot, that does not make it any easier.
Hesitancy becomes exclusion in literary newsletters and Instagram posts, where alt text is often absent. Almost daily, I’m confronted with images of covers that don’t exist for me. One of the many ironies of my writing life is that my book, There Plant Eyes, won an award for its cover design—something I myself had a hand in. I wanted the vivid violet speckles to evoke the parts of the spectrum that are not visible to the average human eye. Yes—I’m blind, and I appreciate a good cover as much as anyone.
The thoughtful description of a cover—like the design itself—might be considered a kind of mini review, which does its duty if it provokes the reader/viewer to investigate further. Getting it “right” is not really the point. Alt text and image description are not meant to close the book on the definitive, but to open it onto the vast realms of meaning-making.
In an Alt Text as Poetry workshop led by co-creators Finnegan Shannon and Bojana Coklyat, we began with self-descriptions. It’s an opportunity to give people who can’t see you an idea of your physical appearance, and those who can see you a sense of how you’d like to be seen. Even in self-description, it’s often easier to name clothes than identities—surfaces that offer themselves up with little risk. It’s no wonder people struggle to describe others, where biases and assumptions easily sneak in.
Rather than resolving the tensions, Bojana and Finnegan urge poetic, playful, and experimental approaches to description. After all, looking at images is a creative act—no more neutral than producing images. As John Berger famously put it, “Every image embodies a way of seeing.” And image description forces that embodied seeing into words.
Though it is generally quite bold with its descriptions, when it comes to naming (or guessing at) identity, AI can be as timid as a human—something I’m acutely aware of as someone who relies on metadata to find blind representations in the photographic archive. It is being trained to be careful, but its lack of stakes—skin in the game—can be useful when you need to inspire humans to describe. This was brought home to me when it came time to practice what we learned on historical photographs in the Wallach Division of the New York Public Library.
Besides Strand’s Blind Woman, there was a jazz singer with an incredibly expressive face, little people in a bizarre moon-fantasy exhibit from the 1901 Buffalo World’s Fair, and an early calotype (circa 1848) by pioneer British photographer William Henry Fox Talbot—the one my little group received.
Image description is not mystical or compensatory—not about transcendence. Rather, it’s a kind of translation—from one sense modality to another.
A moment of silence befell us as we contemplated how to begin. My two companions—one of whom was Alabaster—seemed a bit tongue-tied. Just a few months earlier, I might have felt useless and frustrated. Instead, I got us started by snapping a photo of the printed-out sepia-toned image and reading aloud the AI description of “two shelves lined with delicate porcelain and ceramic objects” such as: a tea cup with a matching saucer, both adorned with a floral motif, a small, bottle-shaped object, possibly a perfume or scent bottle, highly decorated with filigree-like patterns, and so forth.
The human tongues were immediately loosened by AI’s boldness and began correcting, nuancing, and leveling up the detail.
Alabaster suggested that the image was some kind of advertisement. I disagreed. Sure, he had direct access to the visual stuff in Talbot’s photo, but I had access to historical context—another way of knowing a picture. I was pretty certain that the photograph of teacups and vases was simply an experiment, made in a moment when almost everything in the world had yet to be captured by the camera for the first time.
What seeing through language loses in immediacy, it gains from time and attention. Image description is not mystical or compensatory—not about transcendence. Rather, it’s a kind of translation—from one sense modality to another. As with all translations, something will be lost and something gained.
Last summer I received support for my photography project from the Yaddo Artist residency, where I met May-Lan Tan, writer of beautiful and twisted stories, including her collection, Things to Make and Break. On our way from dinner to the studio of artist Clayton Merrell, I told her about my book and AI-generated image description to interact with the photos. As we were stepping over the threshold, she asked, “Do you want me to be your AI?”
It was the beginning of a friendship that continues today through voice messages across continents.
When I asked May-Lan what she remembered of Clay’s paintings nearly a year later, fragments reappeared: a photographically rendered sky fractured by silhouetted trees and landscapes superimposed with motifs almost like a “visitation or neon signage.”
What I remember best is the feel of walking through the energized space arm-in-arm with a very cool new friend, beers in hand. Conversation blossomed out of each new detail and question. In a way, we created the images we discussed—and maybe they created us too. As May-Lan put it, paintings are not just visual or surface experiences; “they really talk to you—and they observe us back.”
M. Leona Godin
M. Leona Godin is the author of There Plant Eyes: A Personal and Cultural History of Blindness (Pantheon, 2021). Her writing has appeared in The New York Times, Playboy, Catapult, Electric Literature, Literary Hub, and others. Her online magazine, Aromatica Poetica, is an arts and culture laboratory for the advancement of smell and taste. As a 2023-24 New York Public Library Diamonstein-Spielvogel Fellow, she is currently working on a series of essays about photography, blindness, and image description.



















