Browse resources
Synthesis Methodology
prepared by Xavier Lerin and Albert Carter
For a combination of risk management, regulatory, and reputational reasons, most of the world’s largest banks release policies publicly stating what they will or will not fund. Common commitments include, for example, forbidding the funding of slavery, terrorism, and some environmentally destructive activities.
It is little wonder that campaigning organisations, investors, and stakeholders, care about bank policies and what they say.
Watchtower contains a “synthesis” module to help people understand such policies, starting with environmental policies that restrict financing to fossil fuels.
Synthesis is how we translate bank documents into a structured list of commitments to restrict, condition, or phase out financing for coal, oil, and gas. These structured commitments can then be used to understand, compare and benchmark policies, and to monitor compliance.
On some parts of Watchtower, logged-in users from whitelisted domains can click the synthesis button to see what we mean by structure.

Here is a guide to understanding our synthesis methodology, what we mean by bank “commitments,” what aspects of commitments we capture, and the limitations of this process.
----
Commitments
The central idea behind synthesis is to capture each distinct commitment a bank makes and, for each one, the applicability rules that can shift its breadth and strength well beyond what the headline suggests.
By commitment, we mean any stated policy position that limits how a bank will finance fossil fuels. We do not consider statements that describe ambition, context, or general climate strategy to be commitments.
For example, in this 2025 document, HSBC commits to phasing out financing for thermal coal-fired power and thermal coal mining by 2030 in EU/OECD (§ 4) countries … except that it can still finance them if they abate (filter) some of their emissions (§ 5) and even if not, HSBC can arbitrarily choose to waive that commitment for almost any reason (§ 2). We categorize this as a phase-out commitment with a number of loopholes.
Distinctness
We consider a commitment to be distinct when it differs from other commitments in at least one of the following key dimensions: effective date, theme (coal vs oil & gas), value chain segment, or the type of financing (project vs corporate finance).
When a single passage in a policy contains several distinct rules, we separate rather than bundle them. For example, if a bank states: "We will not provide financing for upstream oil & gas expansion and pipelines to be used for such expansion," we record this as two commitments: one for upstream oil & gas, and one for midstream oil & gas.
Scope
Presently, we limit the scope of commitments we catalog to those related to coal, oil, and gas financing. A bank's asset management, wealth management, and insurance activities are out of scope, as are commitments on other themes (e.g., human rights, biodiversity), unless they explicitly relate to coal, oil, or gas.
We cover three types of commitments:
- Restriction: outright exclusion or conditioning financing based on specific criteria (e.g. revenue thresholds)
- Phase-out: setting a timeline to phase out financing for a specific sector or category of clients
- Enhanced screening: performing additional checks (e.g., enhanced due diligence) to inform risk appetite and governance, without committing to a financing outcome
We distinguish phase-out commitments from portfolio reduction targets: a commitment to exit any type of fossil fuel financing, including commitments to exit with residual exposure, is in scope, but a target to cut financing or financed emissions without an exit is not. This is true even if such targets represent interim milestones within an exit strategy.
Commitment fields
For each commitment, we extract a defined set of fields that determine its scope and strength. This standardizes the explicit and implicit applicability rules that most commonly determine a policy's breadth and effectiveness.
| Structured fields | What they capture |
|---|---|
| Sector |
|
| Geography |
|
| Counterparty |
|
| Product |
|
| Timing |
|
Triggers
Commitments to not fund fossil fuels are often based on threshold values defining what constitutes a fossil fuel company or project. To capture this, we record a number of composite fields:
- quantitative triggers: numeric inequalities specifying a threshold (>25% of revenue from coal)
- qualitative triggers: written descriptions of the threshold (“‘significant’ revenue from coal”)
- timeframe: when a measurement is specified to be taken (“revenues based on the most recent fiscal year”)
Loopholes
Carve-outs to commitments often prove too varied to standardise. As a result, we also capture a less-standardized “loophole” field with two subfields
- Exceptions: a list of written exceptions to the policy
- Weaknesses: descriptions of how loopholes might undermine a policy’s effectiveness.
Keeping loopholes separate from the structured dimensions reflects a deliberate choice: the dimensions standardise the limits that recur across banks, while loopholes capture bespoke conditions that otherwise would not fit a fixed field.
Synthesis & verification process
On a high level, the process of cataloging commitments is very similar to how linguists diagram sentence structure, but applied to a text as a whole. While there is ambiguity in commitments, as there often is in language, there is still useful information to be extracted and synthesised.
Policy synthesis is derived using the above fields and large language models (LLMs, often known as chatbots) in a multiple-pass processing pipeline.
We verify the accuracy of our processing pipeline by comparing its results to a ground-truth set of syntheses derived by human analysts. This ground-truth dataset covers a range of bank fossil fuel exclusion policies, ranging from simple to extremely complicated.
Read how Watchtower evaluates synthesis accuracy.
In our testing, chatbot-derived syntheses exhibit 90% recall (the number of commitments found) with roughly 85% accuracy (the accuracy of each commitment). We consider a commitment accurate only if it correctly captures all of the human analyst-specified fields. Chatbots have a habit of artificially combining commitments and fields, artificially lowering their accuracy and recall scores. With this in mind, their real-world performance is closer to 95% recall and 90% accuracy.
Work on the synthesis data structure and technique was first pioneered by Albert Carter with significant structure technique, benchmarking, and ground-truth dataset work by Xavier Lerin. Aaron Hoffer made further contributions to the technical extraction pipeline. Quentin Aubineau contributed to one of the ground-truth analyses.
Current limitations & future work
This technique, while powerful, has a number of limitations.
Availability of public data
Synthesis is necessarily limited to the documents, webpages, and press releases that institutions publish. (In our case, we currently track the world’s 100 largest banks.) In cases of ambiguity, we are unable to contact banks to ask for clarification.
Furthermore, while Watchtower is continuously scanning the web for bank policies, there is no guarantee that we have been able to find the most relevant documents or that the published policies we have found remain in force. This is a significant technical limitation, as banks sometimes hide these documents from search engines, block automated systems from downloading them, or take other obfuscatory measures against automated analysis.
Instead, we take the most recent versions of documents that Watchtower was able to find, filter for documents containing fossil fuel exclusionary language, and synthesise them.
Adaptability to other fields and institution types
The current syntheses and their related fields are tuned exclusively to fossil fuel exclusion policies. While it is technically feasible to use this technique for other types of policies (human rights, sustainable financing targets, corruption, gender equality, biodiversity, etc) and other sorts of institutions (insurance, asset management, etc.), this would involve additional work and effort that is outside the scope of the current project.
Non-English language
Our synthesis benchmarks contain only English-language documents. While synthesis is possible for non-English-language policies, we have not benchmarked performance in this area and expect that this would lead to somewhat degraded performance. This could be improved upon but is outside of our current scope.
Cross-referencing
When the text of a commitment, definitions of terms used for that commitment, and exceptions for that commitment are scattered throughout one or multiple documents, chatbots often struggle to piece these together to generate an accurate synthesis. Instead, we tend to see parts of this information inconsistently captured. This accounts for 5 to 10 points of our 100 point accuracy scores and is an active area of refinement.
Ambiguity and scoring
What qualifies as a “commitment,” a “loophole,” or most other fields can often be a matter of opinion, even for human analysts, and even after significant clarification and definition work. (Even when benchmarking experienced human analysts against one another, it is practically impossible to achieve above 95% accuracy or recall.)
Chatbots have a habit of combining multiple commitments into one, and combining other fields into one. This limits the accuracy of our accuracy scores, themselves.