What is multimodal AI?

For its first couple of years, using AI meant describing things in words. If you wanted help with an invoice, you typed out what it said. Multimodal models removed that step, and it changed which jobs are worth using AI for more than most feature announcements have.

What "modal" means here

A modality is a kind of input: text, images, audio. A text-only model reads and writes words. A multimodal model takes several kinds in the same conversation and reasons across them.

So you can upload a photo of a delivery docket and ask "which items are missing compared to this order" — with the order pasted as text and the docket as an image. It works across both.

What this unlocks for a business

The common thread is removing a translation step. Previously a person had to convert the real-world thing into text before AI could touch it — and that conversion was often most of the work.

An Australian small-business example

A Gold Coast building company has site supervisors submitting daily reports. The old process: take photos, then sit down and write up what happened, usually at the end of a long day and usually thinly.

Now they upload the day's photos with a few voice-noted lines. The assistant drafts the report describing what's visible, matched against the schedule. The supervisor corrects and submits.

Reports got noticeably more detailed — not because anyone tried harder, but because the friction that made them thin was in the writing, not the observing.

Where to be careful

Why it matters more than it sounds

A lot of small-business work is stuck in formats AI historically couldn't reach: paper dockets, photos, PDFs of scanned forms, whiteboards. Multimodal capability is what brings that work into range.

If you assessed AI a while back and concluded it didn't fit because your work isn't text-based, that conclusion is worth revisiting. The constraint that made it true has largely gone.

Frequently asked questions

Can it read handwriting?
Often, and better than traditional OCR — it uses context to resolve ambiguous characters. Neat handwriting reads well; poor handwriting still causes errors. Worth verifying anything where a misread digit matters.
Is uploading an image different from uploading text, privacy-wise?
No — same questions apply. An image of a document is still the document. A photo of an invoice contains everything the invoice contains, including the client's details.
Can it look at a spreadsheet?
It can read one, but a screenshot of a spreadsheet is a picture of numbers. For anything involving calculation, give it the actual data rather than an image, and check the arithmetic either way.
Does it work with video?
Support is emerging and varies by tool. Audio transcription is well established; genuine understanding of video content is less mature and worth testing on your own material before relying on it.
Does it cost more?
Images consume more tokens than an equivalent description would, so on usage-based pricing they cost more per request. On flat subscriptions you generally won't notice.

Put this to work

Ad On Group runs AI training and enablement for Australian teams through Ad On AI — a three-month, self-paced program that takes non-technical staff from their first prompts to working AI agents.

Talk to us →

Keep reading

← All resources