Scrape and normalize tender data from India's Central Public Procurement Portal (CPPP).
- Scrapes ePublishing (latest tenders) and eProcurement (active tenders) endpoints
- Normalizes data to canonical procurement schema
- Validates output against JSON schema
- Command-line interface for easy data collection
- Python package for programmatic use
- Handles dates, categories, and contract forms intelligently
pip install -e .Or from git:
pip install git+https://github.com/pardarsh/cppp-collector.gitCollect 5 tenders from each source:
cppp-collectCollect 20 tenders from ePublishing only:
cppp-collect --epublish-only --limit 20Save to file:
cppp-collect --output tenders.jsonValidate against schema:
cppp-collect --validate --schema schema.json --output tenders.jsonfrom cppp_collector import CPPPCollector
# Create collector
collector = CPPPCollector()
# Collect tenders
tenders = collector.collect_all(epublish_limit=10, eprocure_limit=10)
# Use tenders
for tender in tenders:
print(f"{tender['identity']['tender_id']}: {tender['procurement']['title']}")
print(f" Closes: {tender['timeline']['closing_at']}")
# Save to file
collector.save_tenders(tenders, "cppp_tenders.json")Tenders are output in normalized JSON format:
{
"identity": {
"tender_id": "BHU/UED/WO/2026-27/90",
"source": "cppp",
"source_url": "https://eprocure.gov.in/..."
},
"procuring_entity": {
"organisation": "Government of India (via CPPP)",
"department": "...",
"contact": { "email": "...", "phone": "..." }
},
"procurement": {
"title": "SITC of Spare Parts and Comprehensive Servicing...",
"category": "Supplies",
"contract_form": "works",
"estimated_value": { "amount": 850000000, "currency": "INR" }
},
"timeline": {
"closing_at": "2026-10-06T14:00:00Z",
"opening_at": "2026-10-06T14:30:00Z"
},
"metadata": {
"extracted_at": "2026-10-03T10:07:25Z",
"extracted_by": "CPPPCollector",
"extraction_method": "scraper",
"confidence": 0.85
}
}- URL: https://eprocure.gov.in/epublish/app
- Data: Latest 10 tenders, basic info (title, dates)
- Table:
id="activeTenders"
- URL: https://www.eprocure.gov.in/eprocure/app
- Data: Active tenders with advanced search, richer metadata
- Status: May have CAPTCHA protection
Tenders conform to schema_procurement_normalized.json:
- Required: identity, procuring_entity, procurement, timeline, metadata
- Optional: documents, eligibility, awards, reference_number, etc.
- All dates: ISO 8601 format (UTC) with Z suffix
# Install with dev dependencies
pip install -e ".[dev]"
# Run tests
pytest tests/
# Format code
black cppp_collector/
# Lint
flake8 cppp_collector/
# Type checking
mypy cppp_collector/- ePublishing: Only retrieves 10 latest tenders per run
- eProcurement: CAPTCHA-protected; requires session handling
- Updates: Re-running collects duplicates; deduplication needed upstream
- Organization names: Generic "Government of India" for ePublishing
- Handle eProcurement CAPTCHA with Selenium
- Implement intelligent deduplication
- Add support for state procurement portals
- Add data.gov.in integration
- Incremental collection (track last collected date)
- Webhook notifications for new tenders
- Database backend option
MIT
Contributions welcome! See CONTRIBUTING.md for guidelines on:
- Setting up development environment
- Running tests and linting
- Submitting pull requests
- Code of conduct
For issues, questions, or suggestions, open an issue on GitHub.