Six Weeks to Launch: How Nava built Document AI in response to H.R. 1
E111

Six Weeks to Launch: How Nava built Document AI in response to H.R. 1

Ryan Koch (00:21)
Lawrence, Sophia, thank you both for joining us here on Civic Tech Chat. Could each of you introduce yourselves and tell us a bit about what you do?

Sophia Philip (00:32)
Great being here. I'm Sophia Phillip, a product manager at NAVA, specifically in its R and D R

Laurence Goolsby (00:40)
And I am Lawrence Goolsby, senior software engineer, also part of Nava Labs working with Sophia.

Ryan Koch (00:49)
And for each of you, what would you say is your personal why? The thing that drives you to get up every morning and do what you do?

Laurence Goolsby (00:58)
Yeah, I think for me it's solving challenging problems, right? And helping people and delivering technology that has positive impacts on people's lives. I know that sounds kind of you know, rehearsed, but that's the absolute truth, right? So doing something that I love, programming all day, every day, and also helping people. That's that's my why.

Sophia Philip (01:26)
I think building off of what Lauren said, definitely plus one, but I think also being able to work with a variety of individuals, organizations that are mission driven to not only just providing services to people, but how might we improve the experience of the delivery of that service? How might we ensure we are cutting right through to the outcome that we're building toward and

Just as Lauren said, getting to engage in those messy problems, but also getting to collaborate across not just specializations, whether in software engineering and product, but also expertise, whether that's deep knowledge or policy, those responsible for the program to collaborate and and find the solutions that will work best for the problem we're aiming to solve.

Ryan Koch (02:16)
The mention of expertise there is a nice segue to my next question, which is about the fact that y'all work in practices that are well known for rapid change, you know, software engineering for product, especially these days. Sophia, with product management, what's something that has changed and what's something that's been a consistent foundation as you've kind of lived through that?

Sophia Philip (02:41)
Yes, I think with product management, something that I'm grateful for is by inherently leaning on the pillars of methodologies like Agile, we really have it built into our ways of approaching the work to adapt, to respond to sudden change, whether it's expected or unexpected. So I think it complements really well with waves of new technologies available, whether it's

I think we're all very familiar with the integration of AI into both workflows but also what we're building. I think as well, this something like using more AI capabilities or building it for customers really helps everyone tackle the blank page problem. I think oftentimes we hear demos over memos, and this just accelerates our ability to once we apply our expertise on honing in what is the problem to solve.

using prototyping, using quick turnaround to validate that understanding and also get it into users' hands as soon as possible to learn from them, to understand are we solving their problem? That is is the most important to prioritize. I think no matter what, never compromising that data is always going to drive our decisions, no matter how much we are, our velocity has increased. It's always ensuring we have the feedback loops from the qualitative and quantitative to understand.

Are we actually moving in the right direction? So for me that's that's really been something that I've embraced in the surprise of how product management has changed in its speed. But at the same time, we we've got some really some really time tested true methods and fundamentals that are are setting us all up for success to move through it in acclimate.

Ryan Koch (04:27)
And Lawrence, I'd wanna ask you the same regarding software engineering as well.

Laurence Goolsby (04:31)
Yeah, I mean, rapid change. You can't really talk about that in software engineering without mentioning AI. and it's been huge, it's everywhere, right? But I think the important thing there is, you know, still understanding the fundamentals through things, being, you know, at the behind the wheel and driving a lot of the decisions and not handing everything over to AI and just say, hey, build this thing and don't make any mistakes. I think that's one of the jokes that that we use internally. right. Because if it was that easy, then everyone would do it.

but it d definitely takes a critical eye to look at some of the things that AI does and, you know, to have conversations and, you know, bounce things off of back and forth, I think is what's really accelerated. You know, whether it's a new framework that's out there or, you know, something, you know, a new tool that maybe, you know, one's not aware of, like that's where AI can come into place. And especially if you're asking, you know, how should I handle this or I implemented it this way. Do you see any issues with it? And I think that's probably the the biggest thing. And it's really sped up the development life cycle.

And I think people expect things to be delivered faster, right? Like taking a year to deliver something is a long time. And I think the expectation is now with AI that it's going to be certainly significantly faster than that. And I think Sophia had a great point there, right? Demos over memos, right? Like show me what you're building. Let me see the thing, you know, as soon as possible. So that's kind of how, in my opinion, things have accelerated software engineering wise.

Ryan Koch (05:59)
You each mention AI in some way. Sophie, you mentioned it with prototyping. Lawrence, you mentioned with with speed and kind of the ability to change fast. that's in the context of your own practices. I'd also be curious if each if either of you have a take on how it's affected how you'd work together kind of across practice.

Sophia Philip (06:17)
Yeah, I think an experience that immediately comes to mind is one thing that we'll be talking about shortly is this service document AI that we're working with Maryland to and have recently launched in a pilot. But on the journey there, something that we've realized and learned again from the data and talking to stakeholders who are trying to understand what the future looks like with a service-like document AI is

How are a bunch of people in offices going to use it? How do you manage it? How do you scale it? And so something that using AAI enabled us to do is in ensuring Lawrence was able to focus on delivering for Maryland in, I think just one morning, also with support of access to tools like Claude and ChatGPT in Navas workspaces and tool suite, I was able to spin up an interactive front-end prototype of.

How might we create a management layer? How might we help configure tenants? And even though it was not something that could immediately go from zero to one, it helped address that blank page problem and gave Lawrence and I not only just a shared understanding. Yes, we have, we always differ back to what's the core problem statement, what's the scope, who are the users we're solving for, but it actually allowed us to play out in real time what are the features we want to prioritize building first.

That help achieve this outcome. And it was by saying, hey Lawrence, I know you've been deep in the weeds of helping stand this up for Maryland at 1 p.m. today. Can you just here's a URL, can you just click on this, poke around? Where are the gaps? What's some does this give you enough to work off of? And so I think it also really helps quickly close the back, the gap between abstraction and concrete implementation.

And frankly, it's also a ton of fun building things. I think yes, we have our respective capabilities, but some of the favorite work we've done, my favorite work we've done around evaluation, we've had designers, software engineers, product managers, and evaluation expertise at the table to figure it out. So I think it actually only deepens the the impact of collaboration because of that abstract to concrete in such a quick turnaround time.

Ryan Koch (08:33)
I can definitely get behind the statement that it's it's fun to build things and it's fun to get to test your ideas like that. And your mention of the service we're gonna talk about here throughout this episode, Document AI, is a good segue towards our main topic. I wanna start a little bit base level and talk about kind of the challenge that this project's meant to go after. You know, we're here to talk about the challenge that comes with handling and processing documents for for programs, like those that might handle social well-being.

or other very important functions in folks' lives. Why is document processing a challenge worth solving for?

Laurence Goolsby (09:11)
Yeah, that's that's a great question. I I think when you look at, you know, whenever someone is applying for benefits or just in general, right? Processing a document, you know, there's the person uploading the particular particular document and receiving feedback in real time whether or not the document is good or not. And then there's also kind of the back end like caseworker side. So, you know, are they

looking at documents and then typing things in manually? Is there something we can do to help kind of ease that burden? And I think there's a lot that kind of factors into it, right? Like how do we ensure that good documents are getting through, that documents are of high quality and are readable. And you know, if there's a way that we can extract some information from the document and you know help pre-populate things or give someone you know, maybe they don't have to type as much, or we can

give them some sort of visual cue that, you know, yes, we extracted this, but this is something maybe you should take a look at. I think that definitely makes it a problem worth solving.

Sophia Philip (10:15)
Absolutely. And I think all of us have experienced iterations on systems that are just trying to keep track of the information that they need about you as an individual. And so we have watched forms digitize. We have watched maybe not watch, but backends integrate that would make us not have to watch so many things on our own accounts or across respective platforms. So

I think document processing is is a means to an end. And so some of the work that we are doing now is looking to the public and private sector and figuring out how can we pull in opportunities so that we are raising the floor of public the public's expectations for how benefit programs might function, how government services in general. I think

It is not just benefit programs, whether you're applying for a driver's license or a permit to build an extension on your home. You're going to have to submit a document of some kind. So it is again about raising those expectations for how public services can perform and bringing in pretty standard and expected practices that have already been tried and true in other industries so that it can be a seamless digital experience, regardless if you're engaging with a public service.

or a private entity. And I think as well, humans are complex and that's wonderful, and also can make for a lot of paperwork. And so so one of the things that we're really is very core to ensuring we're doing human centered design is that complexity does not cause a deterrent for the way the shirt service should work. It is about saying let's make it a little bit easier on you so you don't have to con constantly justify or explain

unique turns of your life every time you engage with a public service. How might might we make that experience smoother so that you can just come as you are and get the service that you intended to apply for without unnecessary jumping through hoops and also it should feel as seamless as the way you might access your online online bank account.

Ryan Koch (12:22)
Real quick, it's editing Ryan here. I just wanted to give a note of thanks to Nava Public Benefit Corporation for sponsoring this episode, where we talk a lot about document processing and about their tool document AI. So with that said, let's go ahead and get you back into that conversation.

Ryan Koch (12:38)
in the fight against administrative burden or the fight for good user experiences, sometimes we're saddled with responding to a burden that's imposed by purposeful policy change, whether we're talking state laws or federal laws. What do you think about these situations where we find ourselves needing to make a by design requirement easier to comply with?

Sophia Philip (13:01)
Thanks for that. I think I always remind our teams to keep top of mind what is our answer if we were asked in order to do what? And I think it when it comes to complex policy or ambiguous policy implementation, it's ensuring we have the internal alignment, which again, huge shout out to Nava for anticipating the need and role of policy strategists on delivery teams. This has gotten nipped this in the bud, at least for our work with Maryland.

We have immensely benefited from having SNAP policy expertise to really talk through even if I'm not a policy expert, how might I understand the way this is structured in order to do what? And so that helps us create a shared definition of the outcome we're building toward and in turn building back of okay, if this is where we're going to if this is where the policy wants to get to, how can we build this into the design?

But also where's the wiggle room, just as you said, to ensure the user is not accidentally left out of the experience that we're building toward compliance? So I think a great way that we've navigated it this is collaborating with jurisdictions like Maryland, where we understand what is the current benefit application experience. And given their AI policies around requesting consent, how can we always ensure not only an option to opt in or out, but how do we build a path?

Pathway to create reversibility for consent. Because I think it's always important to ensure that the user has autonomy as you adapt a service or introduce new capabilities. And so something like that really helps ensure that we are honoring compliance with the state's policy, but also building in steps of consent in a way that isn't surprising to the user, provides just in time plain language information of what's the service that they're consenting to.

And then the ability to revisit that should their mind change, but really just to ensure that they have confidence that they can control the experience that they're moving through that has certain policy required steps.

Ryan Koch (15:15)
Y'all have been working on an open source solution to address some of the challenges we've talked about so far called document AI. As I I think it the name has floated into our conversation already. for folks out there, maybe they haven't come across it yet, what is document AI?

Laurence Goolsby (15:32)
Yeah, if I were to summarize document AI at like the highest level possible, it is a service where you upload a document and the document or the service rather decides whether that document is clean and legible or if it's not clean and legible. So that is the simplest explanation for it, right? Is this a good document or a bad document?

But there are a whole host of other things that go into that determination, but again, at the most simplest terms. I've uploaded a document. Is it good or bad? Provide that feedback to the person uploading the document and give them the option to proceed if or I guess upload a new document if it comes back with a non-positive response.

Ryan Koch (16:26)
And why did this project kick off?

Laurence Goolsby (16:30)
Yeah, so this project started back last year, 2025, right around this time, as a rapid response to the HR1 legislation that was passed. so we started it here in Pennsylvania. we were on site early on. We had six weeks to implement this process from beginning to end. we did not have an architecture yet defined.

we worked really closely with folks from AWS and others to figure out how we were going to do this in the time allotted. certainly, you know, it was a challenging time, a lot of long hours, but we implemented it officially on October first, 2025. and again, the goal, as mentioned previously, was to

help improve the quality of documents that are coming through. So if someone's uploading a cat picture, for instance, flagging that and saying we don't believe this to be a document. We're also doing some extraction from the document. So if someone uploads a driver's license, we'll extract name, address, anything we can from the document, really. And it's kind of morphed into this open source initiative that we have where we recognize that this problem isn't specific to Pennsylvania or mainland or anyone

or even the type of document, right? It's very generic. I have a document. I need to upload it. So this is a problem I think that a lot of places are trying to solve. And if we can help accelerate that, then you know that's why we're here.

Sophia Philip (18:11)
And I think this identifying this problem, but also opportunity, speaks to what we discussed earlier, where states have been tasked with reducing their payment error rates. And I think this is a great example of looking at the data within states to understand what might be the most feasible and impactful area to start that might help reduce those payment error rates. And so the form that it took in this case.

was an intelligent document processing service at Document AI. And I think the opportunity that really presented itself is as Lawrence knows this is a shared problem across states. And yes it is in response to new policy requirements and compliance, but in the way that it was set up to be open source, to be a reusable service for document processing.

It can be reused in for different policy needs or different program objectives. So, like I mentioned earlier, you know, we may be submitting an application to for a permit or even for a public university application as a student, and we're using our mobile phone to take photos. We check these are the ones I want to upload, and accidentally in there is that photo you took of let's say a very cute dog on the street, which

Never stop yourself from appreciating the little moments while you're walking through life. But it sometimes is difficult when you're doing paperwork to make sure you're including everything necessary and also removing those things you don't mean to upload. And so something like Document AI helps reduce those unexpected uploads or unintentional uploads so that you can really get to the thing that you want through this public service, whether that is submitting your application to pursue higher education through a public institution or

I would just like to build a porch on my house and make sure it's in compliance.

Ryan Koch (20:10)
it sounds like y'all definitely didn't waste a crisis there as that a big policy change was coming through. that kind of core thing you started with, did you find that that payment error rate ended up being affected by the work?

Laurence Goolsby (20:23)
Yeah, I I think payment error rates, it's a very complex thing. So there are multiple ways that it can be addressed, right? document AI is just one of many ways that it can be addressed, right? And the hope there is that the quality of documentation that we're receiving goes up. And that in turn should have a positive impact on the error rate. There are also things that can be done upfront, right? Like doing some

you know, analysis or like pre-quality checks on the existing data to see things that may be flagged, right? So yes, I think document processing itself had an impact, positive impact on payment error rates. But again, it's it's such a complex thing that, you know, one thing in and of itself probably isn't going to move the needle significantly. But a lot of things in combination with each other can really, you know, help kind of drive that number.

Ryan Koch (21:21)
And in in in the answers before, the term open source was used a a couple of times there. what do you mean when you say open source in the context of this project and why did you all set it up that way?

Laurence Goolsby (21:32)
Open source in this regard, is that the code itself lives in GitHub. It is publicly available for anyone to take a look at. see how we've implemented it, see what services we're using. and you know, basically just take the code. They can fork the repository if they want to, they can follow along by starring the repo if they want to.

And that way as you know, we make changes and updates to it and improve the service. anyone available in the public can take a look at it, see what we're doing, open up a PR in theory if they wanted to and help contribute to make it a really better solution.

Sophia Philip (22:15)
I think as we've discussed many a times, document processing comes up in a lot of different offices for a lot of different objectives. And I think with open source, one thing we were anticipating is that given, dare I say, a very universal need for a little help to make document processing easier, why not prioritize an outcome that we are not going to create vendor lock-in with document AI?

Having it as open source enables different jurisdictions, different offices to be able to look at what is an option we could consider to improve our document processing. So it enables, much like we are anticipating for the user document AI, that autonomy, that ability to understand what's within the realm of my control. How can I make an informed decision given the outcome that I'm driving toward?

So again, thinking of the perspective of an office jurisdiction as they're evaluating what tools they might want to procure, build in-house, and in this case, fork a branch from to then bring into their own system. And I think something that's really key about open source, it's and the way that we've set up the repo specifically is it provides the clarity on how a government might provide the service.

But it does not require that government agency or office to compromise their security or privacy standards. So we've made it visible what security checks are built into document AI, but at no point does our public repo store private information, require agencies or offices to pass PII or PHI.

Rather, they can fork the service into their system, stand it up so that it remains in a secure environment. They can adapt it to their security and privacy requirements. Again, ensuring and prioritizing the safety and needs of the end user, but also again helping remove those blockers and burdens to finding the most appropriate tools for.

their mission for the problems that they are trying to work through to improve public service delivery.

Ryan Koch (24:25)
And you mentioned folks that would ultimately interact with this. who are those main personas, the the types of people that would interact with this service?

Laurence Goolsby (24:35)
Yeah, I I think there's probably like three main categories. first and foremost, it's the individual uploading a document, right? So getting that real-time feedback, you know, is this document a good document? Does it have the information that it should have, or is it missing something? Right. So for example, if I upload a driver's license and

For whatever reason, let's say the address is unreadable, you know, letting me know in real time that hey, this document isn't readable or you're you're missing something here. so you know, either take a new picture or upload a new document. So again, first persona of someone uploading a document. Second, I think is you know, whoever is receiving the document and evaluating it and looking at it and saying, okay, does this document contain the information that I need? And the hope there is that the documents that come through.

have all of the information that the individual needs right then and there without having to reach out to the individual who uploaded it and then have back and forth cycles, which can take time, right? So we want to reduce that. And I think lastly, you know, anyone who's managing document AI, and I I think Sophia mentioned this earlier, but, you know, looking at any sort of metrics. So how many documents have we processed or what types of things are we

seeing or how many documents are being rejected. I think those are the three main three main people. Again, those uploading the documents, two, number one, number two, those receiving the documents and evaluating them. And three, people who manage the service so they can, you know, quantify it and have some numbers about how it's actually performing or what it's actually doing.

Ryan Koch (26:23)
For those folks uploading who I I imagine an agency might refer to them as like a customer or something of that of that nature, what is the experience like for them as they interact?

Laurence Goolsby (26:33)
Yeah. So the places where we've implemented it so far, it's been a synchronous experience. So the document is uploaded, the user customer receives a message that says, you know, your document is being processed. And and I should say also, you know, we've given the applicant or customer the ability to either opt in or opt out of document AI. So if they don't want to to go through it, then they don't have to. but if they choose to go through it.

It is a synchronous process where document is uploaded, document AI runs 25, 30 seconds or so, and then the uploader receives a message that says your document is good, green, or we've detected an issue with your document. Again, I'm summarizing those words. but I think the important thing to note here is that we're not preventing anyone, or the goal is not to prevent anyone from uploading a document, but rather to give an option. So if the document

We detect something, there's a message that's you know provided that says, would you like to continue with this document or upload a new one? So in that case, you know, document AI isn't making any decisions about whether the document should be uploaded. That lives with the person who's uploading the

Ryan Koch (27:55)
And it's been mentioned that there's some layers of consent for a person that's interacting with this. what's the experience like for someone let's say someone doesn't want to give their consent and they're interacting with a service that uses documented AI? What kind of fallback ends up being recommended?

Laurence Goolsby (28:12)
Yeah, I think, you know, from what we've seen, or we don't process the document and it's the same as it always has been for that user experience, right? It's a no op, essentially, right? Doesn't do anything with document AI, isn't evaluated or anything. So it's the exact same process that exists before document AI was ever implemented anywhere. So

Ryan Koch (28:41)
And let's talk about for the folks that are receiving the documents, caseworkers as a term often here for that. what is it what's it like for them as they interact with it?

Laurence Goolsby (28:51)
Yeah, I think it's so I'm not a caseworker. so I can only generalize. but I know that our expectation or our goal is to increase the quality of documents that are coming through, and also extracting information from the document so that a caseworker or someone can quickly look and validate that yes, this is correct. and again, I think document AI there are multiple stages in in the pipeline. So the first one again, depending on

where someone is in the life cycle is is this document good or bad? That's it, right? That's step zero. step one, step two, step three, et cetera, is actually extracting the the document, which we're doing. And then further down the line, actually using that information to populate, you know, a form or an application, if you will. And you know, that I think is the ultimate goal to help pre-populate things so that people can do a quick check and say, yes, that's correct, or no, it's not, and edit it.

And not necessarily spend as much time again assuming that people are typing things in. or for example, if someone faxes something or scans it in, right? That's something that may need to be typed in from that upload to the actual form. So if we can simplify that, then fantastic.

Sophia Philip (30:07)
Earlier we were talking about I mentioned how humans are complex and their paperwork can maybe make it difficult to understand exactly what that complexity means for where they are today, what is the service that they need. and so, but that shouldn't necessarily mean it's an immediate blocker for them getting that service. And so something that we've considered in building Document AI that is isn't set up in the code base to support is

It can process anticipated forms or standardized forms such as W-2s, but or identification cards like a passport or a driver's license, but there might be some more unexpected documents like a handwritten invoice. So something that, especially in light of HR1, we wanted to make sure was supported is types of work that may not have

formal documentation or standardized form. So thinking of gig economies, thinking of people who might be driving Ubers and being able to support submitting those invoices from gig economy formal businesses to also, hey, this is someone who mows my lawn every week. And that handwritten invoice or documentation is acceptable because it includes the required data fields that DocumentAI is looking out for.

so I think it's something really important to note that we've again, how do we keep the human at the center of this service? is anticipating the conventional and maybe unexpected ways people will document or try to comply with policy requirements and support that effort. Additionally, we are able to process documents that are in different languages like Spanish as well. So

supporting bilingual documents is again something that we understood early on from data. We needed to be able to support. And I think shifting gears to the caseworker perspective as well, reducing the caseworker needing to put effort in going back and forth with a customer. Maybe it's a sending out a notification that a document needs to be re-uploaded.

They don't always have the control to specify which document, why does it need to be reuploaded? So that can create a lot of or not create, but cause a tot of cause a lot of time to be lost in just getting that clarity and alignment. So just as Lauren said, by providing that real time feedback up front so the customer can immediately respond to it, that enables them to go on with their lives, provide the things that are required for compliance, but also for the caseworker to have that high.

the documents that they need to check for that compliance right off the bat, not have to go do that back and forth. And if there is maybe additional attention needed, more more effort put into potentially complex cases that need their support, they're able to prioritize that work over back and forth and time loss to just saying, I don't know if you realize, but you uploaded a photo of a cat and

It really needed to be a W two and that happens and that is understandable, but we can provide support at earlier on at the right time to reduce time loss for caseworkers and customers.

Ryan Koch (33:26)
A sort of failure mode that comes to mind is thing a lot of things could be coming in, right? And I might as a caseworker enter a state where it's like, you know what, I should just a hundred percent trust what the machine's doing and just stamp the thing. is that something y'all have seen? And I guess how would you advise an agency in feeling out whether that's something to be worried about?

Laurence Goolsby (33:51)
Yeah, I think that has to do with where they are in the lifecycle. Again, if it's just validating document quality, then no issues there. But if it's looking at the extracted information, we provide confidence scores currently, along with the data that we extract. Right. So that gives us some sort of indicator. And again, it's not foolproof, but it gives us some sort of validation to say that, hey, you know, this was extracted, but maybe it's not as confident as some of the other fields.

Right. Some of the other things that we've talked about as well is potentially, you know, if we were to do this, you know, highlighting something in a specific color to give someone something to look at to say, hey, this is something that needs to be evaluated, right? Whether it's yellow or red or green or whatever the case may be. but yeah, I I think that is something that can happen, right? Where you get used to a service being in place that's

Pretty good most of the time. So you just stamp it. But I think ultimately that's not something that we would want, right? We still want the human in the loop. I think that's the phrase that everyone uses. and we want to make sure that that still happens, right? things should be validated and looked at. But if we can, you know, reduce that time, that's something that certainly we would.

Sophia Philip (35:08)
Something I've really appreciated in being able to both see Document AI used for Pennsylvania and then work with Maryland on implementing it is understanding what are the objectives that need they need to make sure they're upholding and implementing a service like Document AI and what our compliance then they need to be aware of and constraints that we need to work within. And so right off the bat, both states made it very clear.

that with document AI, it is not making a determination. So something that in the way that we have created Document AI is we have not coded in policy to the service. And that was very intentional, both as I mentioned in realizing a lot of different offices use document processing. So an intelligent assistance would be useful across many different policy needs.

But also to anticipate that risk and support states so that they can confidently use document AI and still preserve the necessary compliance and keep the control and decision making within those merit employees, those caseworkers to make the decision on assessing whether a case should be approved or denied. And so those boundaries have been extremely helpful in understanding what's within the control of document AI to make an assessment on, which is

Just document quality. It is not making a determination on the case. And that in turn preserves the expected compliance and process that the policy was designed for states to follow. So it preserves that decision making with the case manager. I think something that is always available to states is to explore an accompanying rules engine if they want something that's more policy-specific. And that's something not as explored with Oscar for Medicaid, and something that we are very intentional again in not making the product.

Because if you start hard coding in policy, it's gonna be really hard to reuse and maintain as policy can evolve and change. So again, thinking of those three different users that Lawrence mentioned, we made sure to uphold that third year's user of those managing the service and keep in mind what are the requirements that you need to comply with and how can we ensure we uphold those constraints in a way that sets you up for success with your mission without compromising our groups went into those.

those users and the caseworkers supporting them.

Ryan Koch (37:36)
A tool like this in inevitably becomes a veritable wish list of neat stuff you could do or you could build. How'd you go about developing a vision for what it should be?

Sophia Philip (37:48)
Yeah, I'll I'll jump in here. I I know I spoke to it just a little bit earlier, but I think as as Lawrence has mentioned, Document AI was set up to be a decoupled service. And again, we this is not a service to respond just to HR1 forever and always. In the way that it was designed, it was very intentional so that it could be paired with any type of office or policy to support document processing and and the relative mission that.

they needed to respond to and requiring customers to submit documents in order to the classic question or to do what and that's not for document AI to answer. That's for the program to say, okay, now that we know it is the right type of document, it isn't blurry and includes the data fields that we're expecting, then we can take it and move forward with this. And they no longer need document AI for the steps that follow.

so I think again, where this also really sets up agencies for success who are using tools like Document AI is it can be reused and scaled across an enterprise, across a state. So again, as I mentioned, different examples. even now I'm thinking of other federal agencies where you're asked to submit a document, whether it's for passports, whether it's for claims, health benefits.

So services a service like document AI can be reused ac across it. And I think I'll pass it to Lawrence to I think crack open the hood and talk a little bit more about how we chose to use API to help support that.

Laurence Goolsby (39:21)
Yeah, I mean, as as far as the vision is concerned, you know, this is a back-end API, right? And I think whenever we've demoed it, you know, we use a developer tool called Postman to perform those demos. And while those are okay for a developer, you know, most of the time people who want to see document AI don't want to see me go through Postman and here's a post request and here's the API response.

So that kind of helped drive our thought or vision. Like we need a way, a better way to demo to show what this thing actually does. So part of that helped drive some of the user interface where you can upload. We have a demo site where you can upload a document and watch it process in real time. And then when it completes, you get the image, and then you also get what we call bounding boxes over that, and you can hover to see.

Where Document AI extracted the information and what the values were. So that helps give some sort of visual to an otherwise very back-end heavy service. And that was something that, you know, wasn't part of our initial implementation, but kind of grew organically out of we need to demo this quite a bit. And Postman probably isn't the best tool for.

Sophia Philip (40:42)
thinking back also on what I mentioned earlier of Lawrence and I using AI both in collaboration and internally across the different capabilities on our team.

as well as listening to what jurisdictions expected to do with document AI, that in enabled us to see the need for an administrative or management layer. And so that was something we evaluated to determine is this within scope? Does this serve many means or many objectives? Is this flexible enough for the problems that people are trying to solve with intelligent document processing? And by evaluating prototyping, bringing it back,

to jurisdictions as well as across NAVA, we were able to realize, this is how we evolve a standalone decoupled service to a scaled enterprise level capability.

Ryan Koch (41:32)
With any software, you tend to learn things as you build it. when you learn something surprising, how do you all handle that with your workflow for this?

Laurence Goolsby (41:44)
Yeah, I think it's it's evolved. So when we first started, it was, you know, a singular document being uploaded in real time, right? And that was the use case. And as we've, you know, gone down this journey, we realized that there are significantly more use cases, right? So there's I have a stack of documents that I want to run through document AI. And by stack I just mean like a folder somewhere.

Right. So that's a different input. Or I have a zip file that needs to go through. Or someone's uploading, you know, a single page at a time and they want us to be able to combine that into a a single document. So I think, you know, it's these are things that are, you know, obvious as we look back on it. but given the, you know, initial six weeks that we had to build things, we didn't, you know, our main focus was getting this up and proving that it works. But now we've seen the need to expand it and continue to expand it based on

you know, all of the things that we've learned throughout the process and working with other people and having people use the service.

Sophia Philip (42:49)
Yeah, and taking a step back about how Nava teams might approach complex problem spaces or introducing a new service or updating one, a legacy system. We definitely I you know I jokingly say this, but I do love this idea, which is fundamentals are the building blocks of fun. So we use agile, the fundamentals of agile and scrum. I I I am a nerd for it and love it so because it

anticipates exactly that that surprises will come, things will break unexpectedly, but that shouldn't mean your team every every time that surprise or break happens, it is an all-out fire drill. And I think that's where remembering we have these methodologies that have been tried tried and tested for decades and work vel very well in exactly the space of software development and also

navigating just the unexpected. So something that as as Lawrence noted in a very tight turnaround of those just a few weeks with Maryland, we have launch we stood up this service and launched a pilot within, let's see, February to end of August, so about six months and are planning to scale it across the entire state by December. We were able to do that because we leaned on the things that work best through Agile. And so it was both having

clear clearly defined sprint objectives over two weeks, having those regular touch points both internally as well as with our partners. So we're constantly sharing out where we at in the work, explicitly asking, are there any blockers? Is there anything that's unclear? Referring back to our shared definitions of what's the problem solving for, what are we focusing on now, what's for later, having those sources of truth, and then also in those moments of surprise of the unexpected

using those sprint ceremonies and also touch points as needed to make sure do we have the right people at the table to understand what went wrong, what we can learn from it, and also does this do we need to adapt the way we're working so we can mitigate the risk or likelihood of this happening again? And that's been something it is always an honor to have time with people who are tasked with incredibly challenging policies to uphold and programs to implement.

But it makes it that much more impactful when we can have them at the table, not just to define the problem, but work with them along the journey so that they understand what we're building, why we're building it, and we can stay aligned with the what their outcomes that they're driving for. And it's enabled us to deploy an intelligent document processing service in six weeks, redo it again, but adapted for Maryland's system. And alongside all this.

bring back those lessons and actually continuously update the code base in Nava's open source repo. So it is not static in any way. We are constantly updating it and refining it so that not just, okay, can it, it can process more documents, but hey, we need to make it clear how we're doing quality assurance so that it can be picked up and integrated as a capability with any jurisdiction that might want to use document AI.

Ryan Koch (46:04)
This business of trying to pull useful information off of a variety of document formats, the images people take with their phone and the obviously very perfect lighting and then upload it. I imagine there's a lot of technical challenge to this kind of thing. So I'd like to talk a little bit about how that works. The name document I itself implies AI use in there, but AI is kind of a term that means a lot of things these days. You know, it's generative, it's machine learning, it's

All kinds of fun stuff. In this context, what do we mean by AIU and how it's used?

Laurence Goolsby (46:41)
Yeah, that's a great question. And I think one of the things that is important to flag here is that the documents that are being uploaded today are not being used to train any models anywhere whatsoever. So each document is evaluated independently and nothing is aggregated. We're not learning from it or or anything at this moment. But in terms of AI, I think one of the things that we're doing is

When we evaluate a document, it is a very difficult task to understand the content or context of a document through just pure code. Right. So if someone uploads a W2, I mean, yes, in theory, you could OCR and say, is there a W2 in here or whatever? But then you get into some really heavy, very brittle logic as we've talked about.

So one of the things that we're doing is we're using Amazon Bedrock and also Amazon Bedrock data automation to use their vision model to tell us what type of document it is. Right. So one of the first things we do is we say, is this, you know, what type of document is this? Right. and for that, again, we're using a large language model or vision model, I should say, Amazon Nova Pro, I believe.

and then once we have that, we're also using Amazon Bedrock data automation in the background to match that document against a predefined schema, a predefined set of fields that we want to extract. So for a W-2, again, we've defined 30 some odd fields that we expect to be in a W-2. Pass that over to Bedrock, which is again an AWS service. So how it works behind the scenes, I'm not privy to. But again,

there is some AI component there which takes the or extracts the information from the document, puts it into a nicely defined schema based on what we've defined, and then returns that information to us.

Ryan Koch (48:46)
You mentioned that the the model used is able to identify different sorts of documents. Is that something where an agency might find themselves like trying to find tune that for like they have some hyper specific form one zero five B six

Laurence Goolsby (49:00)
Yeah.

Ryan Koch (49:01)
Q thing?

Laurence Goolsby (49:04)
Yeah, yeah, exactly. And I think that's one of the really great things about Amazon better automation is it gives us the ability to define the types of documents that we are looking for and describe those in plain language, right? So for the document that you mentioned with all of those letters, you know, we would have a high level description of the document. So this document contains XYZ, we're just describing it. And then below that we would

also define the fields that we would want to extract along with just plain language of the document or the fields that we wanted to extract for example. So if it's a first name, we would say, you know, extract the first name. And we're not specifying where in the document it's located. We're just giving it a description of it and it's able to make the associations for us and return it to us again in that structured format. So as long as it we get the same exact structured output.

for the similar type of document. So every WT that we process comes out the exact same way. There may be data, there may not be data, but if it matches it, it gives us in that same format every single time.

Ryan Koch (50:12)
And to connect back, you mentioned earlier that

Laurence Goolsby (50:14)
Yeah.

Ryan Koch (50:14)
these things come with a confidence score. I would imagine that kind of probability of, hey, did this model get that this is actually the first name correct is kind of where that confidence calculation starts to come from.

Laurence Goolsby (50:27)
Yeah, exactly. and confidence is tricky because we've seen, you know, ninety percent confident. And yes, that's great. We've also seen like sixty percent confident. And I think it's more important to note that confidence isn't necessarily accuracy, right? So we've seen some of our sample documents where the first name was actually first name and it extracted it perfectly, but that's not a legitimate name. So the confidence was a little bit lower because that's

Probably not someone's name, right? but it was accurate. So I think it's a very important distinguish, important to distinguish that confidence does not mean accuracy. It means how confident the model is and what it extracts.

Ryan Koch (51:13)
Earlier we also talked about languages. I think in it might have been the Maryland example, or it might or the other state that you had a bilingual requirement. I'd be curious when the model hits something where it's say a language you haven't prepared it for yet, what what is the kind of fallback in a situation like that?

Laurence Goolsby (51:34)
Yeah, so I don't know the exact number offhand, but I know Amazon Bedrock Data Automation supports several languages out of the box. Spanish, certainly, English, French, I think Russian might be in the list. Portuguese

Sophia Philip (51:47)
In Portuguese.

Laurence Goolsby (51:49)
is also in the list. So if there is a document that is uploaded with a language that we don't or that Amazon does not support, that's a good question. I imagine, you know, what we will do is we will

not necessarily find something to match to, but we'll identify that there's like text in the document and that it's not an image. And one thing we don't want to do is penalize, you know, an upload. So if if we upload something and we don't have something defined for it, we don't reject the document. We say we know that this is a document. We haven't defined or we're not expecting whatever this is. We're not going to punish you for it. So, or we're not going to be punitive and say, well, we're rejecting this. No, like it it is a document. It's not.

you know, the person's uploading it the person who uploads it's fault that we haven't defined it. Right. There are many different document types that can be uploaded and we're continuing to build out that corpus of of documents as well. So we understand that we may not always match, but that is not a negative

Sophia Philip (52:51)
And something that we've set up the repo to support is exactly that. We're going we understand that document AI is or any document processing are gonna get we didn't know that was a document that one could submit to satisfy this requirement. And again, just as noted, not designing to force people to comply with what we thought at the time were the appropriate guardrails. They'll always be able to submit, but rather than it sets us up, we

Set up the repo so that people can contribute back to it, can submit PRs. But additionally, it's also encouraging for those using document AI. How might you be reviewing the documents that are coming in, setting up evaluation so that you are building in a feedback loop of, okay, we're noticing the same document keeps getting rejected, but we're learning from case managers this is actually acceptable. So

again, building in both quantitative but also qualitative feedback loops and continuing to normalize that practice, I think is pretty standard in in most approaches to software development for public services, but not losing track of those in something like Document AI, which can move at a pretty high speed of adaptation.

Ryan Koch (54:07)
I'm hearing that we get a lot out of the box from the fact that that Amazon tools are being used. the other side of that then also is then like, well, you gotta use the Amazon tools in order to get that benefit. so I can imagine there could be an agency or an organization where they might see a concern with kind of being stuck on that. how would you address someone bringing that up?

Laurence Goolsby (54:31)
Yeah, that's something that we've talked a lot about internally. and so far we've only built an AWS as mentioned. And I think that gives us a really good baseline in terms of expectations. it is our goal certainly to make this cloud service provider agnostic. So whether you're on Azure or Google or some yet to be named service, that you can pick this up and use it out of the box.

for now we've been really focused on pilots and proving out that this thing works. And then once we have that, that'll give us the ability to really validate other things. So do we want to use Azure? Do we want to use something that, you know, do we want to use an open source model, for instance? Again, I'm not saying we will or we won't, but at least we'll have a baseline to compare to to say, yes, we think this is good enough to, you know, use alongside. And I think that's kind of the other thing. So

While we are using Amazon bedrock data automation, we're also using Amazon Text Track. So we do have the ability to configure based on the type of document what we use. And again, I I think in summary, it's we have something that works and that gives us the ability to iterate on it and apply it to other CSPs as necessary. But yes, certainly we've heard that and we're we're actively talking.

Ryan Koch (55:57)
here and openness for the idea, but maybe it's like a question of where does it fit in the vision, which I think would lead to the question like, what what would be the signal for y'all to say, like, like it's time to bump that up versus other things.

Sophia Philip (56:17)
So I think in regards to when what informs our decision of what provider to use, what architecture to lean into, we are very much keeping our ear to the ground of what are states and jurisdictions currently using right now. Cause I think of course we are always trying to model, I think especially with Document I, what are tools or approaches that are setting up

public entities to be again raise the floor so that their quality of seamless service experience like that of the private sector. So we always want to encourage opportunities to improve that maybe a state hasn't considered, but we also do not want to shock the system by bringing a service that would require an incredibly infeasible investment in something or completely change the infrastructure

that they have been using in a way that wouldn't set their teams up for success. So the reason we've used AWS so far is that is actually what the states we've been working with have been using. So really it's about understanding where where are states at, what are the systems and tools they are using and meeting them there in a way that again will still ensure the flexibility that we believe makes document AI most effective. So

Something that we're doing now alongside the Maryland effort is participating in a multi-state cohort with led by the American Public Human Services Services Association as well as social finance. And that actually is a multi-state cohort. And so we're using that to understand what what service providers are you all using? What from their technical teams, how is architecture set up? And the intent is so we can anticipate

us building and iterating on Document AI in a way that is flexible to it. So even in Maryland, we made a decision to adjust our architecture approach so that it would best complement where their team is at. Because it's one thing to say, hey, we have this thing, we'll that'll help. But if you have to completely gut the house that you're going to put this new service into, that makes it really hard to convince everyone this is an effort worth investing in. So we definitely, again, are trying to keep

the human at the center of this design. In this case, the human happens to be a statewide organization or entity. So a little bit of a a larger abstraction of that analogy, but one that very much is driving our decision making.

Ryan Koch (58:39)
The services that would rely upon something like this often require documentation that has pretty sensitive stuff on it. You know, it could be a driver's license, it could be a pay slip, you know, all the stuff you would need for one to take your identity. as we think about like safeguards and kind of things like that for the security of something like this, I'd like to ask you know, some questions on that on that topic. for example, when, if at all, would data leave the bounds of an agency?

when using a service like this one.

Laurence Goolsby (59:13)
Yeah, the goal is never really. so this so far we've stood up in each agency's existing AWS infrastructure. so it is secure to the standards that they have set up. And we've worked closely, I think, to understand what that is. you know, everything is private. you cannot get to the AWS environments unless you have the proper credentials.

and also I think from an API perspective, we default to not returning data that we extracted. So for instance, if once a file is uploaded, you can get the information back about what was uploaded. And by default, we do not return it, nor do we store it anywhere besides S3 and those buckets are locked down and no one can view it. so by default, you know, i if

You just ask for the response, you will not get the data. and we also you know, have auditing and a whole bunch of other things. But yes, in summary, built within the existing infrastructure alongside of the agency's infrastructure teams. And I've worked closely with many of them to make sure that we are complete.

Ryan Koch (1:00:35)
not returning the data from a a work is work through is interesting because it occurs to me that that means that in a way you're kinda closing off at a vector for something like a prompt injection into the into the model. Is is that kind of thing what you're what you're after?

Laurence Goolsby (1:00:51)
Yeah, I think there are a lot of things. So as you mentioned, it is sensitive data. And we recognize that everyone is in a different part of the journey with document processing. So by default, if you just want to upload a document and see whether or not something was extracted, not necessarily the values, but like what was extracted, we do not return that. You have to explicitly say, please give me the results of this document.

and include the extracted data, at which point we will look over to where the extracted data lives and hydrate the response that we give back. And also keep track of a record that like, yes, this was requested, you know, et cetera.

Ryan Koch (1:01:37)
As an agency seeks to implement document AI into its stack, what kind of impact do you think it has as the agency tries to govern their software that then would seek to use it?

Laurence Goolsby (1:01:50)
Yeah, I think the ultimate goal for us is that agencies are able to take document AI and manage it and customize it and update it as they need to without us necessarily. Right. So the long-term goal is this is a starter. And if you find that there are other things that you need to change to meet your architecture or

add additional document types that folks are able to do that and manage it on their own without need for for us.

Ryan Koch (1:02:26)
As we get to the close of our conversation, if folks are listening and they're like, Man, I wanna get involved, I wanna contribute to this thing, how can they go about contributing to Document AI?

Sophia Philip (1:02:38)
So for contributing to Document AI, we encourage everyone to check out again the open open source GitHub repo. this is a publicly accessible website, so you do not need a specific login to access it. you can peruse it on whatever browser you'd like. And so as you go through it, there in the code base, you can absolutely submit a PR. If you're noticing something, you can deploy it in your own environment. We have

both guidance that includes even videos on how to deploy this on your own machine, as well as what's already going to be running once you do it, what was it, what are things that you can adjust. You are also welcome to request access to our demo sites as well. So we encourage you if you're not quite ready to actually build the whole thing in your own environment or device, you can check out our interactive demo sites for both.

what might be the customer experience with a little bit of insight into what the caseworker might see, should Document AI be integrated, as well as the management layer. So again, that that third user group we kept in mind of those managing or overseeing Document AI. additionally, you are always welcome to reach out to to Nava. we'll be sharing out the contact information for that if those want to learn more about Document AI and

Yeah, Lawrence, anything else to add for how folks can reach out?

Laurence Goolsby (1:04:07)
You nailed it.

Ryan Koch (1:04:10)
And what's one thing y'all hope folks will take from this conversation we had today?

Laurence Goolsby (1:04:16)
Yeah, I I think that, you know, we certainly continue to learn as we go through this process. And we have a solution that we believe meets the needs of document processing. And with that, we'll continue to iterate on it as we learn more. And we're going to continue to improve it as we go along. this is something that

Yes, we've delivered two pilots so far in production. but you know, that doesn't mean the work is completed, right? There are always things that we can do to improve. so if there's, you know, I I think that's probably the biggest takeaway from me is that, you know, two pilots in production, actively working on it, understand that there's more to do, and we're

Happy to take input from anyone's suggestions. that's it.

Sophia Philip (1:05:25)
Definitely something I hope everyone takes away from this conversation is no one is alone in figuring this out. Whether it's figuring out how to respond to the new policy requirements of HR one or whether it's understanding what does it take to introduce a new I new AI capability or just AI at all in a public service program. And I think those are questions that we've grappled with and iterated on, both with Pennsylvania, with Maryland, but also

across other projects and internally as well. So I think just reminding everyone, you are not in a vacuum. You are not stuck in a cavern needing to figure out how to get yourself out. You have there are tons of resources available and other individuals, organizations that are trying to solve this problem or what a way forward looks like. And so we encourage you, please reach out, please engage with

something that I appreciate about product development is it only gets better if we get feedback and if we are talking about it and questioning things and wondering what if. And so I think just remember you can ask what if with others and I think many are ready to to have that conversation.

Ryan Koch (1:06:39)
Sophia Lawrence, thank you both for coming on Civic Tech Chat and sharing about your work trying to address some challenges in document processing.

Sophia Philip (1:06:50)
Thank you so much, Ryan,

Laurence Goolsby (1:06:50)
Thank you for having us.

Sophia Philip (1:06:51)
for your time today. It's been great.

Episode Video