Skip to content
Updated: 18 min read

AI in Service of ESG: Delivering Environmental, Social and Governance Goals

How artificial intelligence is used as an instrument for delivering environmental, social and governance goals — what it measures, what it improves, what it can prove to an auditor, where the social pillar is least served by tooling, and how an AI-assisted sustainability claim fails.

Anna Polak Author: Anna Polak

Delivering on environmental, social and governance commitments means setting targets, gathering evidence that they are being met, and proving that evidence to someone outside the company. Artificial intelligence enters as an instrument in that work: it measures what used to be estimated, optimises what used to be left alone, and tests claims that used to be simply asserted.

Quick Overview

What you’ll learn from this article:

  • Where AI genuinely changes an ESG programme, and where it only changes the presentation
  • What AI does for environmental targets in the operating business, not in the IT estate
  • Why the social pillar has the least tooling and the highest exposure, and what can be built there
  • How AI supports the governance pillar and the assurance obligations that come with it
  • Why data quality decides the whole exercise, and how an AI-assisted sustainability claim fails

Who this article is for: sustainability and ESG leads deciding where technology helps, data and analytics teams asked to support reporting, and executives accountable for claims made in a sustainability statement.

Reading time: 15 minutes

Where AI Actually Fits in an ESG Programme

An ESG programme does distinct kinds of work, and confusing them is the fastest way to buy the wrong technology. There is measuring — establishing the organisation’s actual impact. There is improving — changing operating decisions so the impact moves. And there is proving — producing evidence that survives an auditor’s questions. These call for different systems, and a tool bought for one rarely serves another.

Artificial intelligence is a strong fit for the first two, because both reduce to problems it handles well: inference from incomplete signals, pattern detection across sources too numerous for manual review, and optimisation under constraints that shift faster than a person can re-plan. It is a much weaker fit for the third on its own, because proof requires traceability from a reported figure back to a primary record, and a model that estimates without an auditable lineage has made the assurance problem harder rather than easier.

That asymmetry should determine sequencing. Organisations that begin with the reporting pillar automate the assembly of numbers whose provenance nobody established, then discover during assurance that the automation cannot show its working. Those that begin with measurement build the lineage first and find that reporting largely follows from it.

One boundary is worth setting immediately, because it is where most confusion lives. This article is about AI used in service of environmental, social and governance objectives across the business — not about reducing the environmental impact of the technology function itself, and not about governing AI systems as a risk class. Both are separate subjects, and the closing section says where each sits.

Environmental Goals: Changing Operating Decisions, Not Just Recording Them

The environmental pillar is where AI has the longest operational track record, because the underlying problems — forecasting, control and anomaly detection — were being solved with machine learning long before anyone framed them as sustainability work.

Energy and resource optimisation is the clearest case. A production line, a building management system or a distribution network generates a continuous stream of sensor readings, and the setpoints governing them are usually fixed by a rule written years ago for average conditions. A model trained on that stream adjusts setpoints against live demand, weather, occupancy and tariff signals, holding output constant while consumption falls. The same logic applies to water use in process industries and to scrap rates in manufacturing, where better scheduling and tighter process control reduce material waste directly.

Emissions measurement matters disproportionately because of what reporting now demands. Direct emissions from owned sources and purchased energy are comparatively easy to meter. Value-chain emissions are not: the Corporate Value Chain (Scope 3) Accounting and Reporting Standard organises them into categories spanning purchased goods, transport, use of sold products and end-of-life treatment, most of which sit with counterparties the reporting company does not control. Machine learning converts what is available — spend records, shipping documents, supplier disclosures, product specifications — into category-level estimates, and flags the entries where the estimate is weakest and primary collection would pay off most.

Remote sensing extends measurement beyond the perimeter. Models trained on satellite imagery detect land-cover change and encroachment near operating sites; sensor networks and image recognition identify leaks and fugitive emissions that periodic inspection misses; acoustic and camera-trap analysis supports biodiversity monitoring where a company has committed to a habitat outcome and needs better evidence than an annual site visit.

Circular-economy work points the same capabilities at materials. Vision systems improve the yield and purity of sorted recyclate; models trained on lifecycle inventories identify which design decisions dominate a product’s footprint, so redesign effort goes where it changes the answer rather than where it is easiest.

The caveat across all of this is that optimisation is only as good as the objective it is given. A model asked to minimise energy cost will not minimise emissions where the two diverge — a cheap hour on the grid is not always a clean one — and the objective function is a policy decision belonging to the sustainability function, not to the engineering team implementing it.

Social Goals: The Pillar With the Least Tooling and the Most Exposure

The social pillar is where the gap between ambition and instrumentation is widest. Environmental performance has meters; social performance mostly has surveys and incident reports. That gap is precisely why AI is interesting here, and also why it is risky here.

Occupational health and safety is the most mature application. Incident histories, near-miss reports, maintenance records and environmental readings from a work site can be modelled to identify the conditions preceding injuries, so intervention happens before the incident rather than in the investigation afterwards. The value is in prediction at the level of a task or a shift pattern, not an individual worker, and the distinction is not cosmetic — a system that scores people rather than conditions converts a safety programme into surveillance and will be treated as such by the workforce.

Pay and progression equity is a genuinely analytical problem. Establishing whether a pay difference reflects role, tenure and performance or reflects something else requires controlling for many variables at once across a population, which is what regression-based analysis does. The same techniques surface progression gaps — where candidates from a given group consistently stall at one grade — that headline representation figures conceal entirely.

The counterpart risk has to be stated with equal force. A model trained on historical hiring or promotion decisions learns the pattern in those decisions, including the part that should not be repeated. The Artificial Intelligence Risk Management Framework (AI RMF 1.0) published by the National Institute of Standards and Technology treats this as a design and measurement problem rather than an aspiration: harmful bias is to be characterised, measured against defined criteria and monitored over time, not declared absent. Systems used in employment decisions also fall within the highest-scrutiny categories of Regulation (EU) 2024/1689 (Artificial Intelligence Act), which is the practical reason such deployments need documented evaluation before they go live rather than after.

Accessibility is the social application with the clearest external benchmark. Speech recognition, text-to-speech, automatic captioning and image description have moved from assistive niche to general infrastructure, making products usable by people they previously excluded. The benchmark is published: the Web Content Accessibility Guidelines (WCAG) 2.2 define success criteria a digital product either meets or does not. Generated captions and alternative text help only where their accuracy is verified — an inaccurate caption is a compliance artefact, not an accommodation.

Value-chain conditions carry the most explicit regulatory pull. Directive (EU) 2024/1760 (Corporate Sustainability Due Diligence Directive) frames human rights and environmental due diligence as an ongoing obligation across a company’s chain of activities, structured around identifying adverse impacts and acting on them. Across a supplier base numbering in the thousands, identification is a screening problem: models combining audit findings, sectoral and geographic risk indicators, media and civil-society reporting can rank where scarce audit capacity should go. Screening prioritises investigation; it does not substitute for it, and treating a low model score as evidence of a clean supplier inverts the purpose of the exercise.

Governance: Evidence That Survives Assurance

The governance pillar is where AI meets the reporting obligation, and where the requirement changes from useful to defensible.

Directive (EU) 2022/2464 (Corporate Sustainability Reporting Directive) brought a substantially wider population of companies into sustainability reporting and, critically, subjected the resulting information to assurance. Commission Delegated Regulation (EU) 2023/2772 (European Sustainability Reporting Standards) sets out what has to be disclosed and in what structure. The practical consequence for anyone deploying technology against this is blunt: a sustainability figure now needs the same lineage a financial figure needs.

That reframes what AI should be asked to do. Language models are effective at locating relevant content inside contracts, policies, permits and supplier questionnaires, and at mapping scattered internal metrics onto the disclosure structure the standards require. They are effective at consistency checking — reconciling a narrative figure against the underlying dataset, catching a unit change between reporting periods, flagging where a prior-year comparative has silently moved. They do not replace the control environment: an extraction step that cannot show which document and which passage produced a number will fail assurance regardless of its aggregate accuracy.

Risk and compliance monitoring is the most continuous governance application. Regulatory obligations, supplier events, litigation, media coverage and internal transaction data change daily; models watching those streams and alerting on material change turn a periodic review into an ongoing one. The output is a prioritised queue for people, not a verdict.

Traceability is the third. Where a claim depends on provenance — a certified input, a chain-of-custody assertion, a recycled-content share — the recording technology and the analytical layer do different jobs. A tamper-evident record establishes that a document was not altered after the fact; validating that it was accurate when written is an analytical problem, and conflating the two produces confident traceability over unverified inputs.

Deciding which of these belongs in-house, who signs off on model outputs feeding a published statement, and what documentation an assurance provider will ask for is an organisational design question rather than a technical one. Working through a governance framework end to end — ownership, documentation set, review gates — is the ground covered by the AI corporate governance in practice programme, and it is worth resolving before the first model output reaches a disclosure.

The Data Problem Underneath All of It

Every application above rests on data the organisation may not have in usable form, and this is where AI-for-ESG programmes most often stall.

Sustainability data is scattered by construction. Energy sits in facilities systems, workforce data in human resources, supplier information in procurement, product data in engineering — each designed for its own purpose, with its own identifiers and update cadence, none designed to be joined to the others. Building a common entity model — one supplier, one site, one product, consistently identified across systems — is unglamorous and is the actual precondition for everything else.

Coverage gaps then force a choice with reporting consequences. Where primary data is missing, the alternatives are to disclose the gap, to estimate with a documented method, or to quietly use an industry average. Disclosure and documented estimation are both defensible; the silent average becomes indefensible the moment an assurance provider asks how the figure was derived. Where a model supplies the estimate, its method, inputs and uncertainty are part of the disclosure, not an implementation detail behind it.

Bias in social data is not a data-cleaning problem. Historical records reflect historical decisions, and a model trained on them reproduces those decisions with greater consistency and less visibility than the humans who made them. Detecting that requires deliberate testing against defined criteria, on an ongoing basis, with a route by which an affected person can contest an outcome.

How an AI-Assisted Sustainability Claim Fails

Greenwashing has moved from reputational risk to legal exposure, and analytical sophistication makes the failure mode worse rather than better.

Directive (EU) 2024/825 (Empowering Consumers for the Green Transition) restricts generic environmental claims not supported by recognised excellent environmental performance, and constrains claims about future environmental performance that lack clear, objective, publicly available commitments and an independently verified implementation plan. A well-rendered dashboard is not evidence. A claim generated from a model whose method is undisclosed is weaker than a plain statement backed by a documented calculation, because the analytical layer adds an opaque step between the primary record and the assertion.

The failure has a recognisable shape. A metric is selected because it moves favourably rather than because it is material. A boundary is drawn to exclude the part of the value chain where performance is poor. An estimate replaces a measurement without the substitution being visible. A model output is presented as a finding rather than an inference with a confidence interval. None of these requires bad faith, and all are found by the same question: what would have to be true for this figure to be wrong, and would we know?

The defences are procedural. Fix the methodology before the number is generated. Disclose the boundary, the estimation method and the uncertainty alongside the result. Have someone outside the producing team reproduce it from primary records. Report the metrics that moved unfavourably with the same prominence as those that moved favourably — a uniformly good set of results is a signal about selection, not about performance.

The Cost Side, Stated Plainly

An article arguing that AI serves sustainability goals has to account for what the technology itself consumes. Training and serving large models draws significant electricity, and Electricity 2024 from the International Energy Agency identifies data centres, together with cryptocurrencies and AI workloads, as a fast-growing component of demand. The research literature has quantified the training side for individual models — Energy and Policy Considerations for Deep Learning in NLP brought the question into the mainstream, and Carbon Emissions and Large Neural Network Training showed how sharply the answer depends on model architecture, hardware generation, data-centre efficiency and the carbon intensity of the grid supplying it.

The practical reading is that the footprint is a variable under management rather than a fixed property of the technology. Choosing a smaller model that meets the requirement, running training in a region and at a time with cleaner supply, and reusing an existing model rather than training from scratch all have measurable effect. What follows for the IT estate specifically — measurement method, workload placement, right-sizing — is a separate discipline covered separately here; the point is only that a sustainability programme cannot claim a net benefit it has not netted.

Sequencing a Programme That Survives Its First Year

The pattern that works starts narrow and builds lineage before it builds scope.

  • Anchor to a target that already exists. An AI initiative attached to a committed, externally visible objective survives budget review; one attached to a general interest in innovation does not.
  • Choose the first use case by data readiness, not ambition. The best candidate is a repeated decision where data is already collected and the effect is measurable within a reporting cycle — an energy optimisation on one site, a supplier screening on one category.
  • Establish lineage before automation. Whatever the pilot produces, be able to trace it back to primary records. Retrofitting that later costs more than building it first.
  • Assemble the team across functions from the start. Sustainability, data, operations, procurement, human resources, legal and internal audit each hold part of the problem, and a project staffed only from technology rediscovers this in month four.
  • Define success as a decision that changed. A model producing a more accurate estimate while leaving every operating decision in place has improved reporting, not performance. Both are legitimate goals; conflating them makes evaluation impossible.
  • Publish the method alongside the result. Internal transparency about estimation and uncertainty is what makes external transparency possible.

What This Article Deliberately Leaves Out

This article is about artificial intelligence used as an instrument for environmental, social and governance objectives across a business. Several adjacent subjects have their own treatment and are not duplicated here.

The first is the environmental footprint of the technology function itself — sustainable practices for engineering teams, workload placement, right-sizing, efficient code, and measuring an organisation’s own infrastructure emissions. Its direction of travel is the reverse of this article’s: it asks how to reduce what technology costs the environment, where this article asks what technology can do for environmental, social and governance goals in the wider business. The two are complementary and should not be run as one programme.

The second is the mechanics of sustainability reporting itself — double materiality assessment, the internal structure of the reporting standards, and the preparation sequence before a first report. That is covered on its own terms elsewhere in this knowledge base and appears here only where artificial intelligence changes how the work is done.

The third is governance of artificial intelligence: who owns an AI system, what documentation a formal audit requires, how an ethical audit is conducted, and what the consequences of non-compliance are. That is set out in the AI governance and ethics practical guide. The direction of the relationship is the distinction worth keeping sharp: that article governs AI, this one puts AI to work on governance and sustainability objectives.

The fourth is the normative frame — the principles of responsible AI, how values become design commitments, and how the regulatory landscape is navigated as a whole — which is the subject of ethics and responsibility in AI. Regulatory references above are deliberately narrow: they appear only where a specific provision constrains what a company may do or claim, and every such requirement is cited from the text of the act rather than from commentary about it.

Also outside the scope: the mechanics of explaining a model’s individual decisions, which has its own treatment. Knowing that stakeholders must understand an AI-derived sustainability figure does not tell you which technique delivers that understanding.

Frequently Asked Questions

Does AI actually improve sustainability performance, or only sustainability reporting?

Both, through different mechanisms and on different timescales. Reporting improves quickly, because assembling and reconciling data is what the technology is good at. Performance improves only where a model output changes an operating decision — a setpoint, a route, a supplier selection, a design choice. The test is simple: name the decision now made differently. If none changed, the programme improved measurement, which is worth having but is not impact.

Where should a company start if its ESG data is incomplete?

Start with the incompleteness rather than around it. Map what is collected today against what has to be disclosed, and classify each gap as missing, scattered or unreliable. Missing data needs a collection process; scattered data needs a common entity model; unreliable data needs a control. Deploying analytics over the third category produces confident output from inputs nobody trusts — worse than the original gap, because it is harder to see.

How do you keep an AI system used for social metrics from doing harm?

Model conditions rather than individuals wherever the problem allows it — task-level and shift-level safety prediction is legitimate in a way individual scoring is not. Test for disparate outcomes against defined criteria before deployment and on a schedule afterwards, not once at launch. Keep a human decision-maker accountable for any outcome affecting a person’s employment, and give that person a route to contest it. Treat the evaluation record as part of the system, because under European rules for employment-related systems it effectively is.

What does an assurance provider ask about an AI-generated figure?

Where the number came from, how it was derived, what was estimated rather than measured, and whether the result reproduces from primary records. A model that cannot answer on provenance and reproducibility makes assurance harder, however accurate it is. Building lineage before automation is not a compliance formality — it is the difference between a figure that can be published and one that cannot.

Is the energy cost of AI large enough to cancel out the benefit?

It depends on what is being compared, and the comparison has to be made rather than assumed. Training a large model from scratch differs materially from using an existing one; a workload on efficient hardware in a region with a clean grid differs materially from the same workload elsewhere. The defensible position is to account for the AI deployment’s footprint explicitly within the programme it belongs to, so any claimed net benefit is a net figure rather than a gross one presented as net.

Anna Polak
Anna Polak Opiekun szkolenia

Request a quote

Develop Your Competencies

Check out our training and workshop offerings.

Request Training
Call us +48 22 487 84 90