What data is actually collected through price web scraping?

Livre-Blanc-Comment-choisir-solution-supplychain-optimix-solutions

Check out our white paper

This guide offers you a clear method and concrete benchmarks for identifying the Supply Chain solution best suited to your needs, in the face of growing complexity and ever higher expectations.

donnees-web-scraping-prix-optimix-solutions

When we talk about web price scraping, the first instinct is often to think of the price listed on a competitor’s product page. However, a price alone provides only a very limited view of an offer’s true competitiveness.

A product listed at €99 may be in stock immediately at one retailer, out of stock at another, or sold by a third-party seller with high shipping fees on a marketplace. Similarly, a 15% price drop may reflect a new pricing strategy or simply a promotion valid for just a few hours.

In both cases, the price list is correct. But its interpretation changes completely depending on the context.

That is precisely why web scraping for prices is generally not limited to simply retrieving a price. Depending on the needs, it can collect the listed price and its various components, promotions, availability, shipping costs and delivery times, the seller’s identity, product specifications, and the date and time of the data collection. Other information, such as changes in the product lineup, ratings, or the number of reviews, can also supplement the analysis.

However, not all of this data is equally useful. The appropriate scope of data collection depends primarily on the pricing decisions the company wishes to make.

What data can price scraping collect?

As part of a pricing monitoring process, the key data collected can be grouped into several categories:

– pricing information: regular price, original price, or crossed-out price;

-promotions: discounts, coupons, bundles, or conditional offers;

-availability: in stock, out of stock, or estimated delivery time;

-Delivery terms: shipping costs, delivery times, and estimated delivery date;

-information about the seller, particularly on online marketplaces;

-the identifiers and attributes required for product matching;

-the date and time of each collection to maintain a record;

– Strikethrough prices

– certain additional indicators, such as ratings, the number of reviews, or rankings.

Web scraping automates the retrieval of this information from e-commerce websites and marketplaces. However, it should not be confused with web crawling: crawling primarily involves traversing a site, following its links, and identifying the pages or URLs to explore, whereas scraping aims to extract specific data from within those pages.

Depending on the site’s structure, the information you are looking for may be found directly in the page’s visible content or in structured elements of its HTML or source code.

However, data extraction produces raw data. Before it can be used in a pricing analysis, it generally needs to be cleaned, standardized, put into historical context, and matched to the correct product references.

To learn more, learn about the difference between web scraping and web crawling

Pricing data: much more than just a single price

Price remains, of course, the key factor in competitive intelligence. It allows you to compare multiple retailers, measure price differences, and track your positioning relative to the market.

However, a product sheet may list several different values.

This could be a list price, a current price, a previous price crossed out, or a price reserved for members of a loyalty program. Some offers also charge different prices depending on the quantity purchased.

Reducing all this information to a single value results in the loss of a significant amount of context.

Let’s say, for example, that a competitor lists a product at €50, with a previous price of €65. While €50 is indeed the price observed at the time of data collection, knowing that this is a discounted price allows us to interpret it differently from a permanent price that has been set at €50 for several months.

A pricing survey should therefore distinguish between the following, when the information is available:

  • -the current price;
  • -the initial price or reference price;
  • -the crossed-out price;
  • -the promotional price;
  • -possibly the unit price or quantity-based terms.

This distinction is the first step in avoiding comparisons between business situations that are not truly equivalent.

Promotions and Their Terms and Conditions

A price cut by a competitor should not necessarily be interpreted as a lasting change in pricing strategy.

Web scraping can also retrieve information used to describe a promotion: old price, new price, discount percentage, coupon, bundle offer, bulk purchase, or limited-time offer.

These data are important because a permanent 15% discount and a 15% discount valid for three days may appear the same if only the final price is recorded.

However, they do not necessarily elicit the same response.

When a retailer regularly offers discounts on certain product categories, the history of price observations may also reveal a recurring promotional pattern. In such cases, the observed price drops do not necessarily reflect a change in the underlying price positioning.

Distinguishing between the regular price and the promotional price helps prevent reacting too quickly to a temporary price drop and, in some cases, unnecessarily sacrificing profit margins.

Availability and Inventory Information

Comparing prices isn’t always enough to gauge the true competitiveness of an offer. But the product still has to actually be available for purchase.

Availability therefore provides essential information for correctly interpreting competitors’ prices. When the source site provides it, scraping can collect information such as:

  • in stock;

  • unavailable;

  • limited stock;

  • delayed shipment;

  • available starting on a certain date;

  • Estimated time before shipment.

The goal is not necessarily to know the exact number of units held by the competitor, but to understand the level of availability presented to the consumer.

This data helps provide context for interpreting the price. An offer listed at a lower price but temporarily unavailable does not exert the same competitive pressure as a product that is available immediately.

The intersection of price and availability also provides insight into certain market trends. For example, a simultaneous increase in prices accompanied by a decrease in availability among several players may indicate supply constraints. Conversely, an isolated price increase by one competitor while others remain stable may reflect the retailer’s own business strategy.

Combining price and availability thus provides a more accurate picture of the competitive landscape.

Shipping Costs and Delivery Times

The price listed on the product page does not always match the final cost paid by the consumer.

For example, two competitors might offer the same product for €48 and €50. The first offer seems more competitive. But if it adds a €5 shipping fee, while shipping is included in the second offer, the comparison changes immediately.

When this information is available from the monitored source, the data collection may therefore include:

  • -shipping costs;
  • -the conditions for qualifying for free shipping;
  • -the announced deadlines;
  • – the estimated delivery date.

This data becomes particularly important in categories where fast delivery is a key factor in the value of the offering.

At the same price, next-day delivery and delivery within eight days do not present the same commercial proposition.

The cost and terms of delivery thus make it possible to compare not only the listed price, but also the actual offer presented to the customer.

On a marketplace, the seller’s identity is essential

Collecting data becomes more complex on marketplaces, since the same product may be offered simultaneously by multiple sellers.

The platform itself may sell the product, while several third-party merchants offer the same item at different prices, with their own inventory levels, shipping fees, and delivery times.

In this context, the seller’s identity becomes essential information.

Without this information, a price reduction by a third-party seller may be mistakenly interpreted as a price change by the marketplace itself or by the competitor that the company actually intends to follow.

The number of sellers offering the same product can also provide valuable insight. The arrival of new sellers can increase competitive pressure, while a decrease in the number of listings may indicate more limited availability.

Rather than simply noting that a product is priced at “€79,” the data becomes much more useful when it also includes:

  • -the product in question;
  • -the seller;
  • -the price;
  • -availability;
  • -delivery terms;
  • – the statement date.

Price is no longer a standalone metric. It has become a contextualized observation.

Product attributes make it possible to compare the right items

Even if the price was collected correctly, the analysis is still incorrect if two different products are compared.

Web scraping therefore also retrieves the information needed for product matching.

Depending on the category, these may include:

  • -the brand;
  • -the part number or model;
  • -the EAN, if available;
  • -size or format;
  • -capacity;
  • -color;
  • -dimensions;
  • -capacity;
  • -certain technical specifications;
  • -included accessories;
  • -the warranty;
  • -the variant;
  • -the composition of a bundle.

These attributes make it possible to determine whether two sales records actually correspond to the same offer.

For example, two devices with nearly identical names may differ in terms of capacity, a specific feature, or the included accessories. If the comparison is based solely on their names, the observed difference may be interpreted as a price difference, when in fact it partly reflects a difference in the product itself.

The challenge increases when each retailer uses its own product nomenclature, when standard identifiers are not available, or when certain SKUs are sold in specific package formats.

Once the data has been collected, automated matching systems can use multiple attributes to identify the most likely matches. Artificial intelligence, in particular, can help handle cases where a unique identifier is not sufficient.

However, this step depends on the quality of the available information. If the key characteristics needed to distinguish between two items are missing or incorrect, the match will remain unreliable.

A matching error can thus make a competitor appear to be cheaper when the product being compared is simply not equivalent.

The date and time of the reading allow you to build a history

Pricing information has different value depending on when it was observed.

A price of €79.99 becomes truly usable over time when it is linked to a date and, if necessary, a pickup time.

By repeating the readings and recording each observation, it becomes possible to build a price history.

This history then makes it possible to examine:

  • – the frequency of changes;
  • -the range of variation;
  • -the duration of an award;
  • -reversals to a previous rate;
  • -repeated promotions;
  • -the response time of the various competitors.

Here, it is important to distinguish between the collected data and the indicator derived from the analysis.

For example, the scraper might record that a product cost €80 on Monday at 9 a.m., €72 on Monday at 4 p.m., and then €80 again on Tuesday morning. It is the analysis of these successive observations that will then lead to the conclusion that a short-term promotion took place.

Scraping, therefore, collects time-stamped observations. Historical analysis then transforms this sequence of observations into actionable insights about pricing behavior.

How often should prices be collected?

There is no single ideal frequency that applies to all categories and all products.

A daily reading may be sufficient when rates change only slightly. In a particularly volatile category, however, this frequency may mask a large portion of the fluctuations.

For example, two price checks taken 24 hours apart may show exactly the same price, even though the competitor changed it several times in the meantime.

Conversely, systematically increasing the number of repetitions does not guarantee a better analysis.

Very frequent inventory counts across several hundred thousand SKUs rapidly increase the volume of inventory that must be stored, monitored, and processed.

The frequency should therefore be tailored to the value of the information being sought. Bestsellers, items that are frequently compared, loss leaders, or items that are strategic for price positioning may warrant more frequent monitoring. For very stable items, a high monitoring frequency often provides little additional information.

Scraping can also track changes in the product assortment

Competitive intelligence isn’t limited to products that are already known.

By regularly reviewing the catalogs, it is also possible to spot certain changes in the product lineup:

  • -addition of new items;
  • -missing products;
  • -an increase or decrease in the number of variants;
  • -Changes in the depth of a category;
  • – the introduction of new brands or product lines.

This information becomes particularly interesting when cross-referenced with pricing data.

For example, a competitor might offer very aggressive prices on a few prominent products while maintaining a much less aggressive pricing strategy for the rest of its product line. An analysis based on too small a sample would therefore provide an incomplete picture of its strategy.

Conversely, the introduction of many new products in a segment may indicate that a company is strengthening its presence in that category.

Monitoring the product assortment thus helps to contextualize price movements within the context of actual available supply.

Ratings, reviews, and rankings can provide additional context for the award

Ratings, the number of reviews, or certain marketplace rankings are not always the main focus of a pricing survey.

Nevertheless, they can provide additional context for the analysis.

A competitor that is 5 percent cheaper but whose reviews are getting worse and whose delivery times are getting longer does not necessarily pose the same threat as a cheaper competitor that is immediately available and has excellent reviews.

Conversely, a slightly higher price may still be acceptable if the product has a better reputation or offers a shopping experience that is perceived as superior.

However, this data should remain merely supplementary information. Its purpose is not to directly determine the price to be charged, but to prevent a bid from being deemed less competitive based solely on a comparison of the listed price.

From a Web Page to Actionable Pricing Data

Scraping retrieves information from monitored sources. It does not automatically generate a database that is ready to be used for pricing decisions.

Between the web page and the pricing analysis, several processing steps may be required:

web page → extraction → raw data → cleaning → standardization → deduplication → product matching → historical data → pricing analysis

One of the main purposes of data cleaning is to correct or remove incorrect values and to handle missing data.

Standardization makes it possible to compare information that may be expressed differently across websites.

Matching links competing observations to the correct internal references.

Historical tracking preserves successive changes to enable analysis over time.

Data quality is largely determined at this stage.

A price that has been correctly extracted but associated with the wrong product remains unusable. A promotion recorded as a permanent price distorts the history. An undetected out-of-stock situation can cause an unavailable offer to appear as the best deal on the market. On a marketplace, linking a third-party seller’s price to the wrong listing leads to the same type of error.

Once prepared, this information can be used to power pricing solutions, analytics tools, business intelligence systems, or specific analyses.

Matching this data with internal data then adds another layer of insight.

By linking competitive data to your own prices, costs, margins, sales volumes, or product segments, you can move beyond simply observing the market to obtaining information that can be directly applied to decision-making.

For example, a 5% price difference will not necessarily be interpreted the same way for a high-margin item, an essential product where price is a key factor, or an item that is not very sensitive to changes in the competitive landscape.

Does all the data need to be collected?

No.

Web scraping technologies make it possible to retrieve very large amounts of information, but collecting more data does not automatically lead to better analysis.

Each additional attribute must then be verified, stored, logged, and made usable.

The appropriate scope therefore depends on the objective.

For certain purposes, three pieces of information can cover a large part of the analysis:

  • -the price;
  • -promotion;
  • -availability.

In other situations, it will be necessary to include shipping costs, the seller, detailed technical specifications, or a much more frequent update history.

The same logic applies to the depth of monitoring.

Tracking an overall price index across several thousand SKUs and detecting a price change in 300 strategic products within a few hours are two different challenges. They do not require the same data collection frequency or the same level of detail.

The scope of scraping must therefore be defined based on the desired outcome, not simply on all the information that is technically possible to retrieve.

What data should you actually collect for pricing?

There is no single list that applies to all companies.

Nevertheless, many price monitoring systems share a common foundation:

The price, to measure competitive position.

The promotional nature of the offer, to distinguish a temporary price reduction from a new price level.

Availability, to avoid using an offer as a reference that cannot actually be purchased.

The date and time of the reading, to place each observation in context and build a history.

Information needed to correctly identify the product, in order to avoid matching errors.

Other data may be essential depending on the context.

On a marketplace, the seller’s identity is essential. When shipping costs significantly affect the final price, they must be factored in. When products come in many variations, technical specifications become more important. In highly volatile categories, a more detailed history may be necessary.

The level of monitoring may also vary within the same catalog.

A product that is frequently compared to others or plays a significant role in price perception often warrants more detailed data collection than a product that faces little competition.

The logic, therefore, is to prioritize data based on its actual usefulness for decision-making.

Web scraping of prices primarily collects the context surrounding the price

Web scraping for prices isn’t just about collecting the prices listed by competitors.

For a price to be truly useful in an analysis, it is often necessary to know which product it corresponds to, who offers it, when it was observed, whether it is part of a promotion, whether it is available, and under what conditions it can be purchased.

Product attributes make matching more reliable. Availability information prevents offers that cannot actually be purchased from being used as references. Promotional data makes it possible to distinguish between a temporary fluctuation and a change in positioning. Finally, timestamping and historical data make it possible to analyze variations over time.

In addition to this essential information, the following may be included as needed: shipping costs and delivery times, the seller’s identity, changes to the product lineup, notes, or other supplementary information.

However, collecting more data does not automatically lead to better analysis.

A large volume of data does not make up for poor matching, outdated information, a misidentified promotion, or a price taken out of context.

So the right question isn’t: “What data can we scrape?”

It’s more like: “What data do we need to properly understand a competitor’s price and decide whether to respond to it?”

Based on this decision, the sources to be monitored, the information to be collected, the frequency of data collection, and the required level of detail must be determined.

Subscribe to our Newsletters :

Our Last Articles :

Trade news

Immerse yourself in the latest Pricing and Supply Chain news!

Découvrez nos actualités liées au Pricing et à la Supply Chain