fbpx SEO + AEO Package: Get Found in AI and Google Top at a Special Price
Table of Contents

Stat.vin is a service for checking and buying vehicles from Copart and IAAI auctions (USA and Canada). For each lot, the site generates a separate page with sales history, price, damage, mileage, and other VIN-based data. The site runs in multiple language versions and has millions of URLs.

Initial Situation

At the start of the project:

  • about 15 million pages had dropped out of Google’s index;
  • organic traffic had fallen by roughly 40%;
  • the site was under a spam link attack;
  • competitors and third-party sites were creating duplicates of vehicle pages;
  • the site was undergoing major changes to its structure and page generation logic.

The task was to identify the cause of the indexation loss, restore Googlebot crawling, and return the maximum possible number of pages to the index. In addition, the site had to be prepared for a new direction: pages for buying and shipping vehicles from the USA.

Challenges

  • The site was submitting millions of URLs to Google, far from all of which had equal SEO value.
  • The indexation loss coincided with a spam attack, external duplicates, and structural changes.
  • There was no full server logging at the start, so bot data collection had to be set up separately.
  • Auction data contained thousands of spelling variants of makes and models, which potentially created weak and duplicate pages.

Approach

Instead of mass-submitting URLs for reindexing, we worked through the entire chain:

Discovery → Crawling → Rendering → Indexing → Internal Linking → Site Structure → Monitoring

1. Googlebot Behavior Analysis

The first step was to check whether Google actually receives the URLs Stat.vin submits for indexing. Since there was no server logging, the client set up bot data collection via Cloudflare.

Based on crawl logs, we analyzed:

  • which pages and URL types Googlebot crawls;
  • the frequency of repeat visits;
  • which sitemaps are crawled more actively;
  • Google’s response to URLs submitted through the forced indexing service.

This made it possible to replace assumptions with analysis of Googlebot’s actual behavior.

2. Testing the Mass Indexing Service

One of the first hypotheses was that the third-party forced indexing service was not working correctly. To test it, we ran an experiment: submitted 100 test URLs (orphan pages) to the service and checked the Cloudflare crawl logs.

Result: Googlebot visited 83 of the 100 URLs (~83%).

The mechanism for delivering URLs to Googlebot was working. The cause had to be found elsewhere: in crawl budget, site structure, URL prioritization, and the signals the site was sending to Google.

3. Technical SEO Audit

  1. Rendering. We found no critical issues that would prevent Google from accessing the main content of pages.
  2. Duplicate content. A large How to Buy a Vehicle block was repeated across many vehicle pages. We recommended either making it unique with template/spin content or excluding it from indexing. Later, we prepared a specification for generating unique content based on vehicle data: make, model, year, VIN, mileage, auction, damage, sale status, and more.
  3. Duplicate URLs. Multiple URLs could be generated for a single vehicle/lot. Some of them had a canonical tag but still:
  • consumed crawl budget;
  • created extra URLs for Googlebot;
  • periodically got into the index;
  • aggravated the duplicate problem.

We recommended restricting crawling of such links.

  1. Duplicate meta tags. Title and Description tags were repeated. We prepared meta tag templates for each page type and language version.

4. Last-Modified and HTTP Recrawl Mechanics

We checked how Last-Modified, If-Modified-Since, and HTTP 304 Not Modified worked.

On the main part of Stat.vin, the mechanism worked correctly. In the Auto.RIA section, the Last-Modified headers and the 304 Not Modified server response did not work correctly.

The goal was to help Google determine which pages had actually changed and needed to be re-downloaded, and which ones were not worth spending crawl budget on.

5. New XML Sitemap Strategy

Bot logs led to a hypothesis: Google actively crawls sitemaps of current vehicles and much less often those of sold vehicles. At the same time, the site was submitting a huge number of URLs with varying SEO value.

We changed the principle from “submit as many pages as possible” to “prioritize URLs with the highest value and likelihood of indexing.”

For sold-vehicle sitemaps, we proposed keeping predominantly:

  • vehicles no older than 10–11 years;
  • vehicles with a last sale price of $2,000 or more;
  • premium brands, without these restrictions.

In addition:

  • we recommended regenerating sitemaps regularly instead of continuously accumulating old ones; the chosen frequency was once a week;
  • for sold vehicles, it was decided to reduce the number of language sitemaps and use the English version, leaving Google the ability to discover other language versions via hreflang.

6. Priority for New Vehicles

Vehicles that had just received a full VIN and a complete Stat.vin page were given separate priority.

The logic: the earlier Google receives a new Stat.vin VIN page compared to other sites, the more likely it is to be treated as the original source rather than another data duplicate.

This logic was built into the third-party indexer: vehicles fully added to the database the previous day, with a VIN and an accessible page, are automatically submitted for indexing.

7. Internal Linking

We analyzed the internal linking and prepared recommendations. The goal was to give Googlebot additional paths for discovering pages and redistribute internal link equity in favor of important sections and vehicle pages.

8. Make and Model Normalization

Auction data stored the same make in different variants: Mercedes Benz, Mercedes, Mercedes-Benz. For models, there were even more variants. Thousands of make and model variants had accumulated in the database, which potentially created a large number of weak and duplicate pages.

The client developed a system for daily normalization of makes and models. After the release:

  • makes and models were normalized;
  • old URLs were redirected to new ones;
  • unnecessary vehicle types were excluded;
  • sitemaps were fully regenerated;
  • the site’s main menu was reworked.

9. Server Speed and Crawl Budget

For a site with millions of URLs, even a slight degradation in server response time can significantly affect the number of pages Googlebot is able to crawl. That is why server response speed was monitored throughout the project.

In September, Google Search Console showed: increased server response time → fewer crawl requests.

We passed the issue on to the client. The development team restored lazy loading and made changes to speed up the site. After implementation, the client noted an improvement in traffic.

10. Protection Against Link Spam

In parallel with indexation recovery, we worked on the link profile:

  • analysis of the existing Disavow file;
  • link profile audit;
  • search for new low-quality and spam domains;
  • Disavow update.

Analysis and cleanup of new spam from the profile were performed regularly.

11. Post-Implementation Monitoring

After the technical changes, we regularly tracked:

  • the number of indexed pages;
  • crawl requests and Googlebot behavior;
  • server response time;
  • clicks, impressions, CTR, positions;
  • dynamics of priority VIN pages.

In July, there was a period of high Google volatility in the automotive niche, during which crawling and organic traffic temporarily declined. For this reason, changes were evaluated after the search results stabilized, not within a few days.

Results

Indexation

Number of pages in Google’s index:

  • July: 7.9M → 8.4M;
  • August: 9M → 12.9M;
  • last recorded stage: 12.9M → 13.6M.

Over the course of the work, the number of indexed pages increased by 5.7M (+72%). The July growth occurred despite a temporary significant reduction in crawl budget. At the last recorded stage, the number of indexed pages continued to grow.

Stat.vin: Recovering Indexation After Losing 15 Million Pages - 1

Organic Traffic

Indexation recovery did not produce an immediate linear increase in overall organic traffic: the metrics were affected by changes in search results, fluctuations in demand, and periods of volatility. After a decline in June–July, the trend turned positive in August.

Organic clicks according to Google Search Console:

  • June: 436,503;
  • July: 436,917;
  • August: 454,631.

Impressions for the same period (July → August): 7.22M → 6.02M.

In August, clicks grew by 4.1% compared to July, and CTR rose from 6.05% to 7.55%. The site received more clicks with fewer impressions.

On vehicle VIN pages, one of the key page types for the project, both clicks and CTR increased according to internal reports.

Stat.vin: Recovering Indexation After Losing 15 Million Pages - 2

New SEO Landing Pages

After the technical side stabilized, we began expanding the site’s semantic structure. We researched demand for buying and shipping vehicles from the USA and prepared:

  • the structure of a new section and hub pages;
  • pages for shipping to Ukraine;
  • city pages;
  • make and model pages;
  • Title/Description and on-page specifications;
  • spin content templates;
  • internal linking.

In the first period after launch, the new shipping pages began receiving organic clicks and appearing close to the TOP for target queries.

Summary

Starting point: ≈15M pages dropped out of the index, organic traffic fell by roughly 40%.

Results:

  • pages in Google’s index: 7.9M → 8.4M → 9M → 12.9M → 13.6M (+72%);
  • organic clicks (July → August): 436,917 → 454,631 (+4.1%);
  • CTR: 6.05% → 7.55%;
  • VIN pages: growth in clicks and CTR;
  • new direction: shipping pages are getting their first clicks and rankings.

The project moved from recovering lost indexation to expanding organic visibility in commercial areas.

Is Your Large Website Losing Indexation?

For sites with millions of pages, mass-submitting URLs for reindexing is not enough. Luxeo analyzes Googlebot’s actual behavior, identifies priority URLs, and configures structure, sitemaps, and technical signals with crawl budget in mind.

If a large number of your pages have dropped out of the index, submit a request, and we will analyze your site’s indexation and propose a recovery plan.

How useful was this post?

Click on a star to rate it!

Average rating / 5. Vote count:

No votes so far! Be the first to rate this post.

Author
Dmytro Kovshun

Dmytro Kovshun is the founder of Luxeo Team – an SEO Outsourcing Company. As a leading specialist in the industry, he is recognized as an expert in SEO promotion of websites. With years of experience and a deep understanding of the field, Dmytro continues to drive success and innovation in SEO strategies, helping businesses achieve their online goals.

DO YOU HAVE ANY QUESTIONS? WE ARE READY TO ANSWER THEM!

LuxeoPartners

+351960165177

Contact us

    No file selected
    Thanks for your application!

    Thanks for your application!

    Our specialists will contact you within 24 hours

    To up