All resources

Dataset Python

Products Dataset

A compact synthetic product catalog for pandas joins, margin calculations, category summaries and reference-data validation.

Type
Dataset
Level
Beginner
Updated

In short

Thirty fictional products across Electronics, Office, Accessories, Home, and Software with deterministic costs and prices.

Who it is for

  • Readers learning reference-data enrichment
  • Analysts practicing category and margin calculations

What it helps you do

  • Validate schema and numeric columns
  • Enrich orders through a many-to-one join
  • Aggregate sales and margin by product category

Dataset shape

The file contains 30 data rows and five columns.

Column Meaning
product_id Product key used by the orders dataset
product_name Fictional product name
category Electronics, Office, Accessories, Home, or Software
unit_cost Synthetic per-unit internal cost
unit_price Standard catalog price

What to practice

Use this small catalog as the reference side of a many-to-one merge with orders.csv. It supports revenue, cost, and gross-margin calculations without overwhelming the explanation with hundreds of product definitions. Category values are intentionally clean so they can contrast with the inconsistent country values in the customer dataset.

The order dataset mostly references these 30 identifiers, but it contains one deliberately invalid product key. A left merge with an indicator column exposes that bad relationship. Readers can then decide whether to quarantine, reject, or report the unmatched transaction instead of silently losing it through an inner join.

All products and prices are fictional and deterministic. The catalog spans physical goods and software so category aggregation produces varied results. It is suitable for inspecting cardinality, checking join assumptions, calculating markup, and comparing transaction prices with standard prices.