Processing health reports without seeing PII

This post is part of the My Health Log series. The idea is to build a privacy-first health data app in public. Code on github.

In order to extract and store the information from a health report, my first thought was to directly send the report to an LLM and get the results. But that would inadvertently share PII with the LLM since reports almost always have information like name, contact details, patient IDs etc. The challenge is to strip this information before normalisation can happen.

Proposals

The idea is to use a separate PDF parsing step.

  • Locally run small size LM: A local run model gives us more control over the document parsing and the data shared with the model provider but also introduces inconsistencies in parsing.
  • PDF parsing libs: There are a bunch of Python libs like PyPDF, tesseract etc that might be useful but each has a different success rate based on the report type. This option does give me the best control over the data being sent to the language model for normalisation. From my reading, I found Docling to be the best option here.
  • Online parsing services: Services like Upstage or LlamParse are a good and reliable option too. It would reduce the load on the local machine and might have better accuracy. But again, user’s PII would be shared.

Solution design

To get started with the MVP, the best approach seems to be Docling. The next steps would be to

  • Add another module for PDF parsing to our monorepo. This would have to be a Python module. (since Docling is a Python lib)
  • Add a service to the server for PDF parsing and register local Python server as a provider. Also add the expected types for the output of the service. This way we can easily add another provider and ensure the rest of the system is agnostic to the change.
  • The flow would like
    • User uploads their PDF and it gets added to queue. (We need a queue since the performance of the local server might differ based on environments)
    • Server picks up the task from the queue and uses the local Docling provider to get the results of the parsing.
    • Run the output through some pre-defined regex to strip the user’s PII. The info to be stripped would include
      • Names, date of birth
      • Phone numbers, email addresses
      • Patient IDs, insurance numbers
      • Physical addresses
    • Store this output to DB to act as an audit trail. (It might be a good feature to show the user what output is being sent to the LM for normalisation ensuring we didn’t send PII)
    • Continue with the next normalisation step.

Next steps

With this pipeline, health data gets extracted and sanitised locally before any LLM sees it. We can even enable users to review the sanitised output to verify personal wasn’t shared. Next up is the implementation step for this setup.

Leave a Reply

Discover more from Ayush Pahwa

Subscribe now to keep reading and get access to the full archive.

Continue reading