Explor.ai

A better-governed generative AI stack on Amazon Bedrock for stronger SaaS margins

Explor.ai is an AI consulting company that builds custom solutions for its clients. One of them is Estimai, an automated construction plan intelligence platform that pairs generative AI with Explor.ai’s own purpose-trained computer vision models to extract structured information from complex architectural drawings. As Estimai scaled, its multi-provider AI stack made it difficult to understand the true cost of processing each document or the consumption generated by each customer. Ingeno consolidated Estimai’s generative AI workloads onto Amazon Bedrock and introduced a governed access layer that measures and attributes every model call. The result is a clearer view of cost by service, model, document type and tenant, giving Explor.ai the data it needs to optimize margins, support usage-based pricing decisions and simplify the operation of the platform without compromising processing time.

Duration

3 months

Category

Online Services

Project

Generative AI cost governance on Amazon Bedrock

AI consulting behind a construction SaaS

Explor.ai is an AI consulting company that designs and builds custom artificial intelligence solutions. Estimai is one of those solutions, a SaaS product for the construction industry. From architectural drawings in PDF format, Estimai extracts structured information such as windows, electrical outlets, curtain walls, tables and legends.

Contractors use Estimai to prepare accurate bids far faster than by reading plans manually. Estimai is a live multi-tenant SaaS serving paying customers and processing complex documents at scale, with results expected in seconds.

Beyond the foundation models, Explor.ai’s own data science team develops and continuously retrains computer vision models on its own annotated construction data, for the specialized detection and segmentation work that plan reading demands. That capability is built in house at Explor.ai.

Inference is not an IT line item, it is cost of goods sold

Generative AI is at the core of Estimai. Every document uploaded by a customer triggers foundation model workloads. That makes the cost of inference something other than an IT line item. It is cost of goods sold.

As the platform evolved, those workloads were distributed across several providers: one for language models, another for document OCR, another for pipeline orchestration, in addition to AWS. Each provider exposed costs differently, making it difficult to establish a unified view of the platform’s variable operating costs.

This created two closely related business challenges.

First, Explor.ai needed to understand what it actually cost to process a document and which services, models or document types were driving that cost. Without that visibility, opportunities to reduce cost were difficult to prioritize.

Second, the company needed a better understanding of consumption by tenant. In a multi-tenant SaaS where usage can vary significantly from one customer to another, that information is essential to support pricing and billing decisions that reflect actual consumption.

Any change also had to preserve the platform’s processing performance. Estimai supports interactive bid preparation, so improving cost governance could not come at the expense of latency.

A governed access layer on Amazon Bedrock

Ingeno consolidated Estimai’s generative AI workloads onto Amazon Bedrock and implemented a governance layer designed to measure and attribute every model invocation.

Rather than rewriting each application service to communicate directly with Amazon Bedrock, Ingeno preserved the model abstraction already used by the platform and connected it to a governed proxy running on AWS Fargate. Application services refer to stable model aliases, while the proxy maps those aliases to dedicated Amazon Bedrock Application Inference Profiles.

This approach allows model usage and cost to be attributed by service and environment in AWS Cost Explorer. In parallel, the proxy records each request with the consuming tenant attached, providing Explor.ai with a near-real-time view of customer consumption alongside the AWS cost reconciliation.

With this visibility in place, Explor.ai can identify the most expensive processing paths and act on them. Individual tasks can be routed to appropriately sized models, Bedrock prompt caching can be applied to recurring instructions and overlapping capabilities from external providers can be consolidated onto AWS. Bringing more of the workload onto AWS also opened access to AWS funding opportunities that were not available across the previous fragmented stack.

The same governance layer also centralizes model version management through aliases, supports fallback behavior when a provider error occurs, shares rate limits and cooldowns across instances and prevents application services from bypassing the governed path. The workload runs in the AWS Canada (Central) Region to support data residency requirements.

Beyond generative AI, the engagement consolidated additional platform workloads onto AWS. Lighter pipeline tasks run on AWS Lambda and AWS Step Functions, heavier document-processing services run on AWS Fargate, documents are stored on Amazon S3 and OCR workloads were moved to Amazon Textract. This reduced the number of external providers involved in operating the platform and brought more of the overall cost picture into a single cloud environment.

Full cost attribution, half the compute cost, same latency

Cost attribution coverage went from effectively 0 percent to 100 percent of the generative AI requests routed through the governed layer. Every request is now attributed to a tenant, a service and a model. Before the engagement, none of that spend could be attributed at all.

Explor.ai can identify expensive processing paths and prioritize concrete optimization opportunities such as model right-sizing and prompt caching.

Consumption can be measured by customer, providing the data required to support pricing and billing decisions based on actual usage rather than blended averages.

Processing time was preserved. The platform runs up to 75 parallel foundation model calls on a single document with results expected in seconds. The cutover carried real customer traffic through the governed layer to Amazon Bedrock with no regression against the agreed latency criteria.

Compute cost was cut in half. Compute represents about a third of the platform’s variable cost.

About 90 percent of the platform’s variable cost now runs on AWS. Three external providers were removed from the operating surface.

Centralized model aliases, fallback behavior and shared rate limiting make model operations easier to manage consistently across services.

A key lesson from the engagement is that moving a model to Amazon Bedrock does not, by itself, reduce the price of inference. The business value comes from the governance and visibility created around those workloads: understanding where costs originate, selecting the right model for each task, caching repeated instructions, consolidating overlapping services and taking advantage of available AWS programs.

Amazon Bedrock, AWS Fargate, AWS Lambda, AWS Step Functions, Amazon Textract, Amazon S3, Amazon RDS, Amazon ElastiCache, Amazon Cognito and AWS Secrets Manager.

AWS

Have a project in mind?