Amazon RecSys is a package-owned multi-stage recommender built on Amazon review data. It trains hybrid retrieval plus ranking pipelines, exports versioned runtime bundles, serves the active bundle through FastAPI and Jinja, and records MLflow and monitoring artifacts around that flow.
src/amazon_recsys/ is the source of truth. notebooks/amazon_recsys_pipeline.py and notebooks/RecSys.ipynb remain as notebook-facing compatibility layers over the same package-owned ML core.
- trains a hybrid recommender with popularity, cooccurrence, latent CF, content-based retrieval, and an optional two-tower retriever
- ranks candidate sets with XGBoost by default, with
dlrmkept as an experimental backend - exports and activates versioned ONNX-backed serving bundles for online inference
- serves recommendation, history, model, evaluation, and monitoring endpoints through FastAPI plus a Jinja UI
- logs training and monitoring runs to MLflow when enabled
- computes batch feature drift and concept drift from served inference logs and delayed outcomes
This project’s biggest lesson is that recommender quality is usually won or lost before the ranker ever sees a row. The low offline hit rate from the quality bundle made it clear that XGBoost tuning was not the next high-leverage move; candidate recovery, source balance, category-aware backfill, and retrieval diagnostics mattered more. The slow and low-recall neural retriever experiment also sharpened the operating rule for this repo: neural retrieval is useful only when it independently improves candidate recall under the available compute, so it stays behind explicit experiment profiles instead of becoming the default path.
The second major lesson is operational: a recommender becomes trustworthy only when training artifacts, serving latency, and monitoring semantics are treated as one system. The project moved from notebook-driven modeling into versioned bundles, serving indexes, MLflow lineage, drift dashboards, synthetic-outcome safeguards, and diagnostics that separate feature drift from true concept degradation. That work exposed practical production concerns: avoid serving-time scans over full interaction tables, make readiness checks manifest-based, interpret small monitoring windows cautiously, and build enough observability to know whether a quality issue is caused by retrieval, ranking, traffic mix, or delayed outcomes.
pip install -e .[dev]
Copy-Item .env.example .env -Force
python -m amazon_recsys.cli.main export-bundle --run-name debug-local --run-profile debug --activate
python -m amazon_recsys.cli.main serveOpen:
http://127.0.0.1:8000/http://127.0.0.1:8000/docs
The training and export entry point is:
python -m amazon_recsys.cli.main export-bundle --run-name debug-local --run-profile debug --activateThis project uses the Amazon Reviews 2023 dataset. Download and documentation are available from the official dataset page:
To download the default categories expected by this project into amazon_review_data, run:
python scripts\download_amazon_reviews_2023.pyPlease cite the dataset paper when using this project or derived experiments:
@article{hou2024bridging,
title={Bridging Language and Items for Retrieval and Recommendation},
author={Hou, Yupeng and Li, Jiacheng and He, Zhankui and Yan, An and Chen, Xiusi and McAuley, Julian},
journal={arXiv preprint arXiv:2403.03952},
year={2024}
}Training is configured from .env, temporary AMAZON_RECSYS_* environment variables, and a few CLI flags. The CLI flags are intentionally small:
--run-name: names the artifact folder underartifacts/amazon_recsys/<run-name>/--run-profile: choosesdebug,quality,quality-neural, orfull--force-rebuild: rebuilds cached corpus artifacts when you need a clean preprocessing run--version: sets the exported bundle version forexport-bundle--activate: makes the exported bundle the serving bundle
The most important data-volume controls are:
AMAZON_RECSYS_RUN_PROFILE: broad preset.debugreads up to 100k review rows per category and keeps small training caps;qualityuses the full available category files with moderate caps;quality-neuralis the same quality-sized experiment with the neural retriever enabled;fulluses the full files, enables the neural retriever, and removes the retriever training cap.AMAZON_RECSYS_CATEGORIES: JSON list of categories to train on, for example["All_Beauty"]for a smaller single-category run or["All_Beauty","Automotive","Industrial_and_Scientific"]for the default three-category run.AMAZON_RECSYS_DEV_MODEandAMAZON_RECSYS_DEV_FRACTION: deterministic sampling before preprocessing. This is useful for quick experiments against a fraction of each configured category.AMAZON_RECSYS_K_CORE: minimum positive interactions required per user and item. Lower values keep more sparse data; higher values produce a denser but smaller training graph.AMAZON_RECSYS_RANKER_TRAIN_EXAMPLE_CAP,AMAZON_RECSYS_RANKER_VAL_EXAMPLE_CAP, andAMAZON_RECSYS_EVAL_USER_CAP: cap ranker and evaluation workload without changing the raw data scan.AMAZON_RECSYS_ENABLE_NEURAL_RETRIEVER: controls whether the TensorFlow two-tower retriever is trained. Keep this off for fast local and CPU runs; enable it for production-quality experiments when compute allows.
After a run, inspect candidate recovery without retraining. Add --persist when you want the active bundle's candidate recall summary saved into monitoring history and shown in the dashboard:
python -m amazon_recsys.cli.main diagnose-candidates --bundle-version active --split test --sample-size 500 --persistFor a quick small run:
$env:AMAZON_RECSYS_CATEGORIES='["All_Beauty"]'
$env:AMAZON_RECSYS_DEV_MODE="true"
$env:AMAZON_RECSYS_DEV_FRACTION="0.1"
python -m amazon_recsys.cli.main export-bundle --run-name beauty-dev --run-profile debug --activateFor a production-style training run:
$env:AMAZON_RECSYS_ENVIRONMENT="production"
$env:AMAZON_RECSYS_USE_MOCK_BUNDLE_IF_MISSING="false"
$env:AMAZON_RECSYS_MLFLOW_ENABLED="true"
$env:AMAZON_RECSYS_MLFLOW_EXPERIMENT_NAME="amazon-recsys-prod"
$env:AMAZON_RECSYS_MLFLOW_LOG_FULL_BUNDLE="false"
$env:AMAZON_RECSYS_CATEGORIES='["All_Beauty","Automotive","Industrial_and_Scientific"]'
python -m amazon_recsys.cli.main export-bundle --run-name prod-2026-04-28 --run-profile quality --version prod-2026-04-28Review the metrics and exported artifacts before activation:
python -m amazon_recsys.cli.main activate-bundle prod-2026-04-28Use docs/running-the-app.md for the detailed training configuration guide, including how profiles change the amount of data and compute used.
In practice, this system would usually sit behind a web or mobile product as an online recommendation service. A common Azure shape is:
For the expanded Azure multistage architecture, see Multistage Recsys Design.
- train and export a bundle from CI, Azure ML, or a scheduled batch job
- store the exported ONNX bundle under durable storage and activate the current version
- deploy the FastAPI service to Azure App Service, Azure Container Apps, or AKS
- expose
/recommendor a product-specific endpoint behind Azure Front Door or API Management - let the website, app, or backend call that endpoint whenever a user lands on a page, opens the app, views a product, or refreshes a personalized feed
The request flow is typically:
- A user opens the site or app.
- The frontend or backend sends the known
user_id, or a short interaction/history payload for a cold-start session, to the recommender service. - The service loads the active bundle, generates candidates, ranks them, and returns the top items for that placement.
- The client renders those recommendations on the homepage, product detail page, cart page, email widget, or in-app feed.
- The serving layer logs the inference event so monitoring can later compare live traffic against the reference bundle.
- When outcomes arrive later, such as clicks, purchases, ratings, or add-to-cart events, those can be ingested into the monitoring flow to measure drift and online performance.
On Azure, the clean separation is usually:
- website or mobile app: owns page rendering and user session context
- application backend: decides when to request recommendations and what business rules apply
- recommender service: owns ranking logic and active model bundle selection
- storage and MLflow: keep bundle artifacts, metrics, and lineage
- monitoring jobs: compute drift and online quality from served traffic plus delayed outcomes
That means this repo can be used either as:
- a direct recommendation microservice called synchronously during page load
- an internal backend service called by another API before the final page payload is assembled
- a batch or near-real-time generator for precomputed recommendation slots that are later cached at the edge
The simplest production version is usually a backend call on page load: when a user lands on the homepage or opens the app, the product backend calls this service, gets the top recommendations for that user, injects them into the response payload, and returns the final page or screen data to the client.
Detailed guides live under docs/.
- Docs Hub
- Architecture Guide
- Multistage Recsys Design
- Running The App
- Release Process
- MLflow Guide
- Drift Monitoring Guide
src/amazon_recsys/ml/core.py: recommender training, retrieval, ranking, and bundle-facing artifactssrc/amazon_recsys/application/services.py: active-bundle recommendation service and serving fallback behaviorsrc/amazon_recsys/api/: FastAPI routers for health, models, recommendations, and monitoringsrc/amazon_recsys/monitoring/: reference profiles, inference and outcome logging, and drift computationsrc/amazon_recsys/cli/main.py: train, evaluate, export, activate, serve, and monitor commandsnotebooks/amazon_recsys_pipeline.py: notebook compatibility import surface
Primary runtime:
- package-owned recommender logic
- bundle export and activation
- FastAPI plus Jinja serving
- MLflow integration
- batch monitoring based on served traffic and delayed outcomes
Secondary support:
- notebook compatibility for existing exploration workflows
template.pyfor bootstrap/template generationinfra/azure/and workflow files for deployment-oriented support
Use src/amazon_recsys/ as the source of truth and treat notebooks and deployment/template assets as consumers or support layers around that package.