Back to Insights

Case study

We Replaced a Specialist AI Vendor with Custom Computer Vision on AWS

August 15, 2026

5 min to read

Read summarized version with

About project

Working time:

2025 – 2026

Industry:

Consumer Technology, Collectibles

The service:

Computer Vision, Generative AI, Managed Services

Overview

The client’s most important feature was the one thing they didn’t own. Here is how we fixed that: 61 percent lower cost per scan, sub-second recognition, and full ownership of the stack.

A North American consumer technology company running one of the largest digital platforms for collectibles. More than 10 million users. Over 800 million tracked products. More than 13 million daily active sessions. A team that grew from under ten people to around fifty in a single year.

The product’s defining feature is the camera. A user points a phone at a card. The app identifies it, reads its grade, and prices it into a live portfolio.

The problem

That feature ran on a specialist third-party recognition vendor. Four pressures made that untenable.

Cost. Pricing scaled with every scan, on a platform processing 2 to 3 million scans a day. A vendor price increase was already scheduled for the following year. Growth was making the product more expensive instead of more profitable.

Speed. Recognition averaged roughly five seconds. The vendor hosted in Europe. Every scan from a North American user crossed the Atlantic twice. That latency was designed in by someone else’s data center placement.

Reliability. The client’s own monitoring logged 24 vendor outage incidents in a single 20-day window. Outages they could only report and wait on.

The bar. The incumbent was no generalist. Collectibles recognition was its entire business, running at 97 percent accuracy. Any replacement had to match that on a catalog of more than 500,000 products, from real phone photos shot through plastic sleeves under kitchen lighting. Not clean reference images.

What we built

Prove it first. We ran an AWS-funded proof of concept and validated parity with the incumbent on a representative slice of the catalog before touching production. The client committed with evidence, not hope.

The architecture. Custom computer vision models locate the card in the frame, separate it from the background, and classify the finish type: holofoil, reverse holofoil, or flat print. That distinction matters because valuation turns on it. The isolated image becomes a vector embedding, searched against a Qdrant vector database on Amazon EC2, which returns the closest match from the full catalog. Structured data lives in Amazon RDS, the image archive in Amazon S3. For graded cards, a GPU-accelerated OCR stage reads the slab label and extracts the grading company, grade score, and certificate serial number across PSA, BGS, CGC, TAG, and ACE. Everything except the load balancer sits in a private subnet, with Amazon GuardDuty, AWS KMS, and AWS IAM handling threat detection, encryption, and access.

Train on the task, not the catalog. Models were trained on Amazon SageMaker and deployed for inference on Amazon EC2 GPU instances via NVIDIA Triton Inference Server, tuned so a single GPU serves multiple models without idle spend. The decisive choice: we trained the models on the general task of locating, isolating, and classifying cards, not on any fixed card set. A new release never touches the model. The client’s own team loads new images into the vector database and the cards are searchable in minutes.

Ship it like infrastructure. The inference tier runs as an Auto Scaling Group launched from an immutable, Packer-built golden AMI, scaled on Amazon CloudWatch application metrics and NVIDIA DCGM GPU telemetry, and shipped through CI/CD. Rollout was staged: a user cohort first, then a rising share of live traffic validated head-to-head against the incumbent. Full production cutover came roughly four months after first assessment.

Run it. The service now operates inside our 24/7 managed operations practice, monitored by the same team that runs the rest of the client’s platform.

The results

  • 61 percent lower cost per scan than the previous vendor, measured in a like-for-like production comparison at live traffic volume
  • 0.73-second average end-to-end recognition, down from roughly five seconds
  • 2 to 3 million scans processed per day
  • Full catalog of more than 500,000 products, across English, Japanese, Spanish, and Chinese
  • New card sets live in 5 to 15 minutes with zero model retraining, replacing a multi-week vendor cycle
  • Full intellectual property assigned to the client

The last line is the one we care about most. We build the system, prove it in production, and leave the customer owning it outright. The feature that defines their product is no longer rented.

Why this pattern matters

Plenty of platforms outgrow the vendor that got them to market. A black-box AI service was the right first move. Then the numbers turn: costs scale the wrong way, the roadmap belongs to someone else, and the product’s core capability sits on infrastructure you can’t see into.

Closing that gap is what we do. As an AWS Premier Tier Services Partner holding both the Generative AI and Agentic AI Competencies, we pair the deepest tier of AWS engineering credentials with AI delivery that holds up under real production load.

If your most important feature runs on someone else’s platform, let’s talk.

Contact our experts!


    By submitting this form, you agree with our Terms & Conditions and Privacy Policy.

    File download has started.

    We’ve got your email! We’ll get back to you soon.

    Oops! There was an issue sending your request. Please double-check your email or try again later.

    Oops! Please, provide your business email.