Google eyes filing cabinets

Paper files next for the great data hoover

Tue 5 Sep 2006 // 10:47 UTC

Google has revealed plans to help convert the world's paper filing cabinets, in Tron-like fashion, into mere nodes in the great hive mind.

The firm will be using an optical character recognition program called Tesseract that was found gathering dust in Hewlett Packard's garage.

"In a nutshell, we are all about making information available to users, and when this information is in a paper document, OCR is the process by which we can convert the pages of this document into text that can then be used for indexing," Google uber techie Luc Vincent said on the firm's code blog today.

Once recognised as one of the three most accurate OCRs on the market, Tesseract had been out of action since 1995.

HP decided it was better out than in if it wasn't making any money and punted it to the Information Science Research Institute at the University of Las Vegas to have it restored for an open source release. The uni gave it to Google, where it was quickly assimilated.

The software has some limitations, Vincent said. Comparatively speaking, it's not that accurate any more, it will only read English, does not like multiple columns or fancy layouts, and baulks at greyscale and colour documents. But, he said it was better than any other open source OCR software.

"Google currently "reads" almost every web page in the world. Come help us read all the printed material as well!" the firm said in an advertisement for OCR engineers. ®

More about

TIP US OFF

Send us news

Topics

Special Features

Vendor Voice

Resources

Channel

Google eyes filing cabinets

Paper files next for the great data hoover

More about

TIP US OFF

Other stories you might like

Japanese and Singaporean devs battle over gamified crowdsourced telco maintenance app

China's mega-telcos are spending billions on AI servers

Senate passes law forcing ByteDance to sell off TikTok – or face a US ban

Protecting distributed branch office environments from ransomware

US government reportedly ponders crimping China's use of RISC-V

White House tweaks HIPAA to shield medical files of those seeking reproductive care

Intel Foundry ticks another box in quest to fab mil-spec chips for US DoD

Using its own sums, AMD claims it's helping save Earth with Epyc server chiplets

Waymo robotaxi drives down wrong side of street after being alarmed by unicyclists

Banned Nvidia GPUs sneak into sanction-busting Chinese servers

Miles of optical fiber crafted aboard ISS marks manufacturing first

Seagate joins the HDD price hike party, blames AI for spike in demand

About Us

Our Websites

Your Privacy