Problem
Develop a optical character recognition (OCR) engine that could convert financial statements and invoices into a PDF format.
Challenge
The system needed to convert several kinds of documents, but the client did not have enough samples of each type to easily train a model.
Solution
Fusemachines experts started the project by ensuring the end product would match the client’s strategic goals. The client’s stakeholders shared their vision for the tool with the team and performed a joint discovery of their data. Once the problem was fully scoped out, Fusemachines deployed computer vision (CV) and natural language processing (NLP) engineers to begin building the tool.
Fusemachines NLP and CV engineers built a MVP to automatically recognize and extract alphanumeric characters from tables within their financial statements and invoices. Despite the lack of samples, the engineers were able to successfully organize the extracted information in PDFs.
Tell us how we can help and we’ll take care of the rest.