Environmental, social, and governance (ESG) data now serves very different jobs: portfolio screening for asset managers, supplier due diligence for corporates, ESG reporting for compliance teams. Most ESG data providers describe themselves in similar terms: broad coverage, real-time signals, regulatory alignment. The meaningful differences show up in daily use, when an analyst looks for a supplier controversy but can’t find it in the feed, or when a compliance team reconstructs the evidence behind a screening decision from eighteen months ago.
A demo shows you the ESG data software. It does not show you how the ESG data collection underneath actually works: where the data comes from, how often it refreshes, or which languages the models read. That is what the seven questions below cover. The first five establish what the data is, and the last two establish whether your teams can work with it. Run them in a live session, with the ESG data platform open and your own entities loaded.
First, define what you are buying
The ESG software market bundles several products under one label, and a scoping error here wastes the rest of the evaluation.
Carbon accounting platforms measure your own emissions. Carbon accounting tools calculate a carbon footprint across scopes and draw most of their input from systems you already run, such as purchasing and energy records. No ESG risk dataset replaces carbon accounting, and no carbon accounting platform substitutes for outside-in evidence.
ESG reporting software assembles disclosures, mapping them to GRI, SASB and the other reporting standards a filing requires. That work also starts from your own records rather than from evidence about anyone else.
ESG risk data points outward instead. It draws on sources outside the companies it covers rather than on what those companies choose to file, which is what makes it useful to risk management teams carrying exposure through holdings, suppliers and counterparties they do not control. Whether a vendor delivers it inside a broad ESG solution or as point solutions is a packaging question, not a substantive one.
ESG ratings agencies assess companies rather than supply raw evidence. A rating draws on outside-in sources and on what the company provides directly, including questionnaire responses and details that were never published, which is the depth an external feed cannot reach. The cost is breadth and frequency: rated universes are smaller, and scores are refreshed on a review cycle. Some ESG ratings agencies license third-party risk data to widen coverage and fill the gaps between reviews.
The seven questions in this buyer's guide concern the data underneath either kind of ESG solution, which is what separates an ESG data provider decision from a software decision.
1. How many companies are covered, public and private?
Coverage of listed companies is well served across the market. Roughly 50,000 companies trade on public exchanges worldwide, they publish disclosures, and every established ESG vendor ingests them. The differences start beyond that perimeter.
Private equity portfolio companies, tier-two and tier-three suppliers, and acquisition targets rarely publish sustainability reports, so an ESG rating built on corporate disclosure returns an empty record for them. That empty record then reads as low risk, even though it reflects limited visibility. Disclosure-dependent methods carry a second weakness even where filings exist, since the company is describing its own ESG performance and its own ESG initiatives, which is where greenwashing risk enters and where disclosure-derived ESG scores diverge most from observed behavior.
What matters is the private company count as a number, alongside an explanation of how it is produced. An ESG data provider that derives signals from external sources rather than self-reported filings can cover a private supplier in Vietnam on the same basis as a listed multinational. How that coverage is counted also matters, since operational risk often sits within a subsidiary or a plant trading under a different name.
2. How frequently is the data updated?
Coverage tells you whether an entity appears in the dataset. Update frequency tells you whether what it says is still true.
Controversies develop faster than reporting cycles. A labor investigation or a product recall can escalate within days, while an ESG rating, refreshed annually or even quarterly, describes the company as it stood some time ago. What settles it is the typical lag between a story's publishing and the signal reaching your feed. Measured in minutes and hours, the ESG risk data supports active monitoring. Measured in weeks, it supports periodic ESG reporting, a valid use case, and a narrower one.
Depth of history matters for the same reason in reverse. A decade or more of continuous archive lets you judge whether an incident is isolated or part of a pattern, and lets you backtest before committing budget.
3. Can you see the source behind every signal?
Current data counts for little if the ESG score behind it cannot be traced to a source. Transparent, source-backed evidence determines how far the data travels within your organization.
Follow an ESG score down to its evidence, and keep going until you reach documents: publication name, date, original language, and a working link. If the trail ends at a category label such as "governance concern, moderate," then the underlying governance data cannot be verified, and it will be hard to place the signal in an investment committee paper or a regulatory filing. Investment decisions presented to a client or a regulator need verified sources, not the ESG score alone.
That same trail settles internal disagreement, since a challenged ESG score can be resolved on the document level rather than through proprietary weighting.
Grouping matters just as much. When 40 outlets cover a single strike at a factory, a well-built system presents a single case supported by 40 documents, not 40 separate alerts. Without that consolidation, alert fatigue follows.
4. Which languages does the data cover?
Following a signal to its source assumes the relevant document is in the dataset to begin with, and that depends on the language.
Incidents surface first in local reporting. A pollution complaint against a plant in Guangdong appears in the Chinese regional press long before the environmental impact reaches an English-language wire. A land rights dispute in Brazil appears in Portuguese, often through an NGO bulletin. On the flip side, if coverage is limited to English, it's more likely to arrive late.
There are two important things to look at: the full language list and whether the models analyze source text natively or rely on machine translation, which tends to lose entity names, local aliases, and the sentiment carried by idioms. Language depth carries particular weight for any ESG due diligence provider, since supplier and third-party risk is documented almost entirely in non-English sources.
5. Analyst-based, AI-based, or hybrid methodology?
AI now sits somewhere in almost every provider's pipeline. The question is not whether AI is involved, but where the balance sits between machine and human.
At the analyst-heavy end, people make most of the judgment calls, and AI mainly narrows what they read, which buys depth and a consistent house view at the cost of reduced coverage and a slower refresh. At the AI-heavy end, models carry the work end-to-end, delivering the scalability manual review cannot reach, though at the risk of poor data quality and too much noise. Most ESG vendors describe their approach as hybrid, somewhere in the middle, so the useful follow-up is what each side actually does. Defining the risk taxonomy, validating model output against labeled samples, and adjudicating edge cases are substantive human roles, worth distinguishing from review of finished output. The other number to get on the record is the false positive rate, and how it is measured, since that is where the balance shows up in practice.
6. How does the data fit into your existing workflow?
The first five questions establish what an ESG data provider is worth. The last two decide whether you get that value back.
ESG data management succeeds or fails on the match between the delivery route and the people using it. Many teams work entirely in a dashboard, where single sign-on matters more than any integration, since access runs through credentials the organization already manages. Quantitative and credit risk teams feeding internal models want an API or scheduled files. Others rarely open the platform and work from alerts instead. A growing number of users query the data through an AI assistant connected over MCP, which allows them to create a custom report without ever using an ESG reporting software. None of these routes is necessarily better than the others; it just depends on what you need.
Entity resolution comes next, since the provider has to map its coverage onto the identifiers you already use, whether ISIN, LEI, or internal supplier codes. Manual mapping becomes a recurring cost and a recurring source of error.
Data integration is where costs are concentrated, so be sure to test integration capabilities in concrete terms rather than on a checklist. Total cost of ownership is driven less by license price than by the engineering time each route consumes, and usability belongs on the same axis, since a dashboard your team avoids has the same practical value as a feed nobody wired up. Data security review is worth starting early, because procurement and InfoSec sign-off often run longer than the technical work.
Alerting logic is the piece that teams most often underestimate, whichever route they take. Thresholds should be configurable by severity, risk category, and entity group, so an asset manager covering 300 holdings or a compliance officer covering 400 suppliers sees only the cases that matter this week. Without that control, an ESG data platform gets filtered out rather than used.
7. Is the methodology audit-ready for regulators (SFDR, CSDDD)?
Workflow fit decides whether the data gets used week to week. Documentation decides whether it can be cited.
SFDR principal adverse impact reporting, CSDDD risk prioritization, Corporate Sustainability Reporting Directive (CSRD) disclosure and EU taxonomy alignment all want evidence behind an ESG metric rather than an ESG score copied into a template. TCFD-aligned climate risk reporting works the same way. Omnibus I, adopted in February 2026 with CSDDD transposition due by July 2028, narrowed who falls in scope of both CSDDD and CSRD without softening those regulatory requirements. At procurement, this comes down to one property: the provider's documentation has to be usable by someone other than the provider.
Three points establish whether it is. Is the ESG data provider's methodology written in a form your auditor can read? Is it versioned, so the provider can say which version produced a given signal and when the logic changed? Methodologies shift with regulatory changes, and regulatory changes have arrived steadily since Omnibus I, so a provider that cannot date its own revisions cannot help you date yours. Can the provider reproduce point-in-time data rather than overwriting history with the current view?
Reproducibility is where providers separate most sharply, and it is the easiest to test. Name a company in your book and a date twelve months back, then see whether the provider can produce the record as it stood on that date. A methodology that only ever shows the current view leaves no audit trail behind its own signals, and regulatory compliance rests on the trail as much as on the score.
Running the evaluation
Environmental, social, and governance data is bought once and lived with for years, so the evaluation is worth running properly. Score every ESG data provider against the same seven questions, using your own entity list rather than a curated demo universe, then weight the criteria by use case. Portfolio monitoring and sustainable investment screening prioritize update frequency and alerting. Pre-deal due diligence prioritizes private company coverage and traceability. ESG reporting prioritizes documentation and audit trails.
One pattern runs through all seven. Coverage, freshness, traceability, and language depth are properties of the data itself, shaped by architecture decisions made years earlier. They sit deeper than the ESG data software and the ESG data management tooling layered on top, which is why they deserve more weight in a scorecard than roadmap commitments do. An ESG strategy is only as strong as the evidence under it, and so is every ESG score built on it.
SESAMm was built against these constraints. TextReveal® analyzes over 30 billion articles from 4 million public and premium sources in over 100 languages, covering 5 million public and private companies, with 10 million new documents added daily and an archive reaching back to 2008. Every signal resolves to its original document, which is what lets asset managers, banks, and corporates see the ESG risks and opportunities across a portfolio or supply chain rather than only those a company chooses to disclose.














.png)