Colleve

Data Analyst

📍 India (Bangalore · Hyderabad · Chennai · Pune · Mumbai · Delhi NCR); remote-friendly💼 Contractor

Hiring for Data Analyst role:

  • Role type : Contractor

  • Engagement window: Full-time during the base-brain build (Aug-Dec) → fractional ongoing (Jan-Apr)

  • Location: Tier-1 India (Bangalore · Hyderabad · Chennai · Pune · Mumbai · Delhi NCR); remote-friendly

  • Reports to : CTO

  • Team : Likely a lead Data Analyst + AI, scaling to a small pool for the global expansion

About the role

The Data Analyst is the primary resource who marshals the documentation that builds the platform's base Brain. The Brain produces a credible screening PCF by ontological equivalence — estimating a supplier's missing data from comparable suppliers (same material × process × country). That only works once a reference base (~1,500 suppliers) populates the equivalence space. Building that base is what makes meaningful output possible for any customer — so this role is foundational and on the critical path.

The platform serves automotive OEMs and their Tier-1 suppliers, whose supplier base is global — so the base you build is global and multi-language, not India-only. You will not do this alone or by hand: you work AI-first (Claude to extract, translate and structure across languages) on the pipeline the AI/ML Platform Engineer builds, and the base is built coverage-first — EU + India (the high-yield, high-relevance regions) first for general audience, then China / SE Asia / Americas as expansion. You run the ingestion; you do not build it.

Key responsibilities

Source & marshal the documentation (global)

  • Global sources — EU company registers + CSRD-cascade disclosures, national registries, OEM / Tier-1 supplier portals, India (BRSR / MCA), Japan / Korea filings, industry associations and global databases — per the approved-source registry.

  • Multi-language, AI-first — use Claude to extract, translate and structure non-English sources at scale; you curate and QA, the AI does the volume.

  • IP discipline — ingest derived attributes only; never redistribute source documents; every field carries a source and confidence tier (A/B/C).

Build the base — coverage-first

  • EU + India first — build coverage of the equivalence neighbourhoods (material × process × country) for OEM/Tier-1 suppliers in the high-yield regions first, so the platform can ride on a covered sub-footprint.

  • Default-anchor the hard regions — where public data is sparse (China, SE Asia, N. Africa), anchor with regional / sector defaults + a few representative suppliers rather than fully sourcing each; improve as primary data arrives.

  • Structure & resolve — normalize into the L2 schema; resolve entity-resolution and source-format exceptions; one canonical profile per supplier.

Quality, provenance & drift

  • QA against the golden set — check profiles within tolerance; triage and flag gap-fills; keep provenance + tier on every field.

  • Monitor drift — watch source-format changes across regions; keep the source registry current.

What success looks like — coverage-driven phasing

  • Foundation (10) — taxonomy, ingest contract, entity-ID and gates frozen; per-region automated yield measured; golden set seeded.

  • 50 — cohort model proven; EU + India yield confirmed on real cohorts.

  • ~150 — EU + India core equivalence neighbourhoods started; first meaningful screening outputs.

  • ~500 (GA-ready) — EU + India covered → credible screening PCFs for EU + India suppliers.

  • GA-usefulness threshold.

  • 800 → 1,500 (post-GA) — add China / SE Asia / Americas — multi-language, default-anchored where data is sparse.

Required qualifications

  • Data literacy — spreadsheets and SQL; some Python a plus; can structure messy source data into a defined schema.

  • Research & sourcing skill across jurisdictions — finding and extracting the right data from reports, filings and portals; strong attention to detail.

  • AI-first working — fluent using Claude and similar to extract, translate and structure at scale, with human judgement on exceptions and QA.

  • Comfort with multi-language sources (via AI translation); rigour on provenance (source · date · confidence tier).

  • Willingness to learn the carbon / LCA and automotive supply-chain domain; 2–4 years in a data / research / analyst / operations role; volume data collection & QA a strong plus.

Boundaries

  • Owns — sourcing, structuring and QA of the global base-Brain data.

  • Not this role — building the ingestion pipeline or L4 (AI/ML Platform Engineer); taxonomy / methodology governance (CTO); the programme (PM); the customer-facing product (Product Engineer).

  • Seam — operates the pipeline the AI/ML Platform Engineer builds; escalates taxonomy questions to the CTO

Apply

Takes ~2 minutes. No account required.

Open to *Tick all that apply.
Currency

Already applied? Update your CV.