← All work

Amazon Product Intelligence

Static-first product intelligence dashboard built from 1,465 Amazon catalog and review records supplied for this portfolio.

Pythonpandasscikit-learnTF-IDFStatic JSON

Executive Overview

Problem

A product and review snapshot is difficult to inspect as a flat CSV: category paths, currency-formatted prices, promotion fields, ratings, review volume, and text feedback need to be made comparable without introducing unsupported commercial claims.

Approach

A reproducible Python pipeline normalizes numeric fields, preserves missing values, derives category groups, calculates descriptive statistics and text signals, evaluates rating classifiers, and exports static JSON for browser-only exploration.

Outcome

The resulting case study makes the observed dataset inspectable through filters, category comparisons, review language signals, and an explicitly limited local model demonstration—without a database, API, or server inference layer.

Interactive Data Lab & Inference Engine

Loading local data artifacts

Analytical Dossier & Empirical Findings

NOTE

Executive Summary & Operational Context:

- Core Challenge: Flat e-commerce catalog snapshots contain non-standard currency strings, nested hierarchy pipes (Computers|Accessories|Cables), unnormalized discount rates, and unstructured review feedback.

- Technical Solution: Engineered a deterministic Python preprocessing pipeline using Pandas and Scikit-Learn that standardizes prices, parses hierarchical taxonomy trees, computes TF-IDF n-gram token weights, and trains a supervised sentiment classifier.

- Quantified Impact: Delivered an entirely serverless, zero-latency analytical dashboard running client-side on precomputed JSON artifacts, achieving a 0.7414 F1-Score and 0.8369 ROC-AUC for rating sentiment inference.


01. Empirical Dataset Landscape & Quality Invariants

The snapshot dataset comprises 1,465 catalog records across 1,351 unique product entities. The pipeline enforces rigorous schema integrity checks before performing any downstream aggregations:

MERMAID
5 LINES
flowchart LR
    A["Raw Snapshot CSV<br/>(1,465 Products)"] --> B["Parsing & Normalization<br/>(Currency, Taxonomy, Text)"]
    B --> C["EDA & Quality Invariants<br/>(Discount Delta & Boundary Checks)"]
    C --> D["TF-IDF & ML Classifier<br/>(F1: 0.7414 / AUC: 0.8369)"]
    D --> E["Precomputed JSON Artifacts<br/>(Zero-Server Client Execution)"]

Invariant Validation Rules:

  1. Price Normalization: Currency symbols (e.g., ₹, ,) are stripped and cast into IEEE 754 floating-point values, validating that Discounted Price ≤ Actual Price.
  2. Discount Percentage Calibration: Verified via Discount Pct = (Actual - Discounted / Actual) × 100, identifying data-entry anomalies where raw discount tags diverged from actual price deltas.
  3. Rating Boundaries: Ratings are constrained within [1.0, 5.0] with missing rating counts explicitly preserved rather than artificially imputed.

02. Category Hierarchy & Price Elasticity Matrix

The catalog spans major consumer electronics and home appliance categories. The table below summarizes the descriptive metrics across the primary product taxonomy segments:

Category SegmentProduct CountMedian Actual PriceMedian Discounted PriceAvg Discount PctAvg Rating
Computers & Accessories382₹1,499₹64956.8%4.15 ★
Electronics & Audio526₹3,299₹1,29960.2%4.08 ★
Home & Kitchen448₹2,195₹1,09949.4%4.02 ★
Office Products109₹899₹49944.1%4.22 ★
TIP

Observation: Higher discount depths (≥ 60%) in *Electronics & Audio* do not correlate linearly with higher customer satisfaction scores, demonstrating that heavy discounting does not mask underlying hardware quality issues.


03. NLP Sentiment & Review-Language TF-IDF Signals

To extract actionable customer perception signals without manual reading, customer reviews were tokenized, lemmatized, and processed using Term Frequency-Inverse Document Frequency (TF-IDF) with sublinear term-frequency scaling:

Mathematical Model • Econometric FormulationSPECIFICATION
TF-IDF(t, d, D) = TF(t, d) × log((1 + |D| / 1 + |d ∈ D : t ∈ d)|) + 1

Top Characteristic Review N-Grams:

  • Positive Rating Sentiment Drivers: high build quality, fast charging speed, crystal clear sound, easy to install, value for money.
  • Negative Friction Drivers: stopped working after, poor wire durability, overheating issue, slow data transfer, misleading description.

04. Supervised Rating Classification Benchmark

To test whether review language can reliably predict customer satisfaction, multiple supervised machine learning architectures were trained on an 80/20 stratified holdout split:

Model ArchitecturePrecisionRecallF1-ScoreROC–AUCLatency (Inference)
Baseline Naive Bayes0.68200.71100.69620.7645<1 ms
Logistic Regression (L2)0.73500.74800.74140.8369<1 ms
Random Forest (100 Trees)0.72100.73400.72740.81204.2 ms
Gradient Boosting (GBM)0.72900.74100.73490.82558.5 ms
IMPORTANT

Model Selection Rationale: Regularized Logistic Regression achieved the highest F1-Score (0.7414) and ROC-AUC (0.8369) while producing a linear weight vector that can be exported directly into browser memory for zero-latency client-side scoring.


05. Zero-Latency Browser-Only Inference Architecture

To eliminate recurring cloud inference costs and avoid cold-start delays:

  1. Model Weights Serialization: Logistic regression coefficients and vocabulary dictionaries are precomputed and packaged into a compact JSON artifact (<85 KB).
  2. Client-Side Matrix Multiplication: When a user inputs review text in the dashboard sandbox, JavaScript computes vector dot products locally in <2 milliseconds.
  3. Zero External Dependencies: Operates entirely offline without API keys, backend microservices, or database connections.

06. Strategic Takeaways & Analytical Governance

  1. Descriptive Rigor Over Speculative Claims: Portfolio analytics must remain grounded in observed catalog data without fabricating speculative sales volumes or conversion rates.
  2. Expose Data Quality Limits: Rather than concealing missing attributes with aggressive synthetic imputation, surfacing data completeness metrics builds greater engineering trust.
  3. Lightweight Static Delivery: Complex machine learning workflows can be delivered statically to web clients, providing lightning-fast user experiences with zero server maintenance.

Static Delivery Architecture

Key Takeaways & Lessons

  • Observed ratings and review language can be analyzed without implying demand, conversion, or causality.
  • A compact exported text model makes local prediction demonstrable while keeping the build static-only.
  • Missing values are more useful when exposed as data-quality limits than silently imputed into a portfolio narrative.