Skip to main content

Handwriting Recognition - A 2026 status check

Background: For years, I had a habit of writing whenever I traveled. No laptop, no apps, just a pen, a notebook, and pages of cursive observations about airports, flights, fellow travelers, and the emotions that come with being away from home. Two decades later, I decided to digitize those notebooks and make them available on Kindle, eventually publishing them as Transit Tales: Between Airports and Emotions: Travel Blog 20 Years Too Late (View on Amazon). As a first step, I assumed modern LLM-powered handwriting recognition would effortlessly convert the pages into text. Instead, the journey led to an unexpected discovery: handwriting recognition is still far from "solved."

The apparent success of modern handwriting recognition is increasingly a combination of imperfect OCR and powerful language-model reconstruction. What users see is often not what the system read, but what the system inferred.

Handwriting Recognition: Better Than It Looks, But Not For The Reason You Think

a hand written page which has been scanned


Modern handwriting recognition demos can be surprisingly impressive. Feed a handwritten page such as the one below into a state-of-the-art OCR system and the resulting transcript may appear almost perfect at first glance. However, a closer examination reveals that the raw recognition quality is often much poorer than it seems. The magic comes not from character recognition alone, but from the language model sitting behind it.

Consider several examples from this handwritten page. The handwritten phrase corresponding to"airport" may initially be recognized as something closer to"surpat". Similarly,"vehicle" can become"valuele","takes a U-turn" becomes"tales a Utum","X-ray machine" becomes"A-vay machine", and"immigration card" appears as"immerdian card". Taken literally, these outputs are poor transcriptions and would be difficult for a human reader to understand without context.

Large language models change the game. Given the surrounding sentence, the model realizes that a traveler arriving at a terminal is far more likely to write"airport" than"surpat","vehicle" than"valuele", and"X-ray machine" than"A-vay machine". The model effectively performs error correction using context, world knowledge, and statistical expectations about language. In other words, it is not necessarily reading the handwriting correctly; it is often inferring what the author probably meant.

This distinction becomes even more obvious with longer phrases. The raw transcription contains fragments such as"want continues - the grove length increase" which are difficult to interpret. Yet an LLM can confidently reconstruct the phrase as"wait continues - the queue length increases". Likewise,"dot Mupose" can be corrected to"what purpose","quese" to"queue", and"lanton" to"London". These corrections are plausible and often accurate, but they are driven at least as much by language understanding as by handwriting recognition itself.

There is an important lesson here. Traditional OCR performance metrics focus on recognition accuracy: did the system correctly identify the handwritten characters? Large language models introduce a second layer that asks a different question: given a noisy transcription, what is the most likely intended text? For users, the final output looks dramatically better, creating the impression that handwriting recognition has become nearly solved. In reality, the underlying recognizer may still be making substantial mistakes, many of which are quietly repaired by contextual reasoning.

The result is both impressive and potentially misleading. If the document contains common travel terms such as airport, boarding pass, London, or immigration card, contextual correction works remarkably well. But when the document contains uncommon names, technical terminology, passwords, identifiers, code snippets, or domain-specific jargon, the same mechanism can "correct" the text into something that is fluent but wrong. Put differently, today's systems often excel not because they perfectly read handwriting, but because they are extraordinarily good at guessing what the handwriting was intended to say.

Summary: Looking at the sample page below, the raw transcription produced words such as"surpat" for airport,"valuele" for vehicle,"A-vay machine" for X-ray machine, and "immerdian card" for immigration card. Yet the final output often looks remarkably accurate because the language model uses context to infer what was probably intended. In many cases, what appears to be excellent handwriting recognition is actually a combination of imperfect character recognition and powerful language understanding. The real breakthrough is not that the machine can read every handwritten word, but that it can often guess the right one.

 Extra Reading:

Confidence is the LLM estimate of how likely the "best guess" matches the intended handwritten word. 
Handwritten textBest guessConfidence
bleambecause / blame20%
surpatairport90%
valuelevehicle95%
canncame85%
tales a Utumtakes a U turn95%
1000per10:00 pm90%
frightflight99%
instid enquiryinitial enquiry80%
want continueswait continues95%
grove lengthqueue length90%
tweekningscreening60%
whoumhmm / whomm30%
bequinbegin95%
mere slowlymove slowly95%
A-vay machineX‑ray machine98%
boarding fan the west itemboarding pass, the next item35%
had baggagehand baggage98%
lantonLondon99%
yhyes95%
dot Muposewhat purpose90%
fan et aromfrom her arm / from a folder20%
gilt the boarding tamgives the boarding pass85%
should Day are of the ben pleasant girlsshould say one of the best pleasant girls70%
tack your bagslock your bags75%
20 1TSA lock? / combination lock?10%
immerdian cardimmigration card99%
quesequeue99%
sigafare vise pageSingapore visa page80%
Monty me in 3001cost me Rs 300/-65%
hundleyhundred85%

Places where context helps.



Comments

Popular posts from this blog

Visualizing Speech Processing Challenges!

Often it is difficult to emphasize the difficulty that one faces during speech signal processing. Thanks to the large population use of speech recognition in the form of Alexa, Google Home when most of us are asking for a very limited information ("call my mother", "play the top 50 international hits" or "switch off the lights") which is quite well captured by the speech recognition engine in the form of contextual knowledge (it knows where you are; it knows your calendar, it know you parents phone number, it knows your preference, it knows your facebook likes .... ). Same Same - Different Different:   You speak X = /My voice is my password/ and I speak Y= /My voice is my password/. In speech recognition both our speech samples (X and Y) need to be recognized as "My voice is my password" while in speaker biometric X has to be attributed to you and and Y has to be attributed to me! In this blog post we try to show   visually   what it means to pro...

Paying Property Taxes Online - Government of Andhra Pradesh

When my father received a SMS stating that you can pay your property tax online. I was thrilled. Why? My father stays with me in Mumbai and paying property tax for a small flat in   Proddatur Municipality was always a pain. The best was to request someone to pay it on behalf of him which meant it was at the time and convenience of the person we requested. Mind you this is no easy task, asking someone to pay on your behalf. A quick search on the web got me to  Commissioner & Director of Municipal Administration  and I it does have an online payment of taxes tab. And boy this was a breeze. As soon as you press the online payment tab you see a neat selection of District | Muncipality | Tax Type. For my purposes I choose Ysr Kadapa (it would be nice if they changed it to read "YSR Kadapa") and then "1014-Proddatur" for Municipality and I chose Tax Type is "Property Tax" (the other option is Water tax) Once you fill in these details. You are directe...

BITS Pilani Goa Campus - Some Useful Information

You have cleared the BIT Aptitude Test and have got admission to BITS Pilani Goa Campus. Congratulation . Well Done. This is how the main building looks! Read on for some useful information, especially since you are traveling for the first time to the campus and more or less you will face the same scenario that we faced! We were asked report on 29-Jul-2018 (Sunday) to take admission on, 30-Jul-2018.  We reached Madgoan (we traveled by train though the airport is pretty close to the BITS campus, primarily to allow us to carry more luggage!)at around 0700 hours (expect a few drizzles to some good rain - so carry an umbrella) on 29-July-2019. As you come out you will be hounded by several taxi drivers, but the best is to take the official pre-paid taxi. It should cost you INR 700 to reach the BITS campus. We had booked a hotel in Vasco (this is one of the closest suburb from BITS campus, a taxi should charge you around 300-350 INR; you will make plenty of trips!) ...