Generative AI and The Future of Finance: A Deep Dive with Sylvain Forté
July 19, 2023
•
5 mins read
Hello, and welcome to our ongoing series on Generative AI in Finance. I’m Sylvain Forté, CEO and co-founder of SESAMm, and in our first article of the series, we’ll explore how generative AI is reshaping the financial industry. At SESAMm, we’ve been fortunate to be at the forefront of this revolution, witnessing the transformational power of large language models like ChatGPT and its iterations.
Unprecedented Evolution in Generative AI
Let's begin with an overview of the current generative AI landscape. Over the last few years, we've seen an explosion in the capabilities of generative AI, particularly in text processing. From BERT to GPT4, large language models (LLMs) have demonstrated increasingly impressive capabilities. These models, performing at a human level for many tasks, are rapidly evolving, making the past six months feel like an exponential leap in the AI domain. Generative AI is no longer a speculative idea but rather an early adoption phase of a powerful technology. At SESAMm, we've leveraged our partnerships with OpenAI and other organizations to gain high-level access to these AI models, allowing us to harness this potential and democratize access to intelligence. It’s an exciting shift that, while reshuffling business models and job roles, promises enormous productivity gains and increased overall value. The key, of course, is ensuring that these benefits extend to society as a whole.
The Disruption in Financial Sector
The finance sector, with its vast array of text-based tasks, stands to gain enormously from generative AI. Any repetitive yet intelligence-heavy tasks — think verification, document generation, or client communication — are ripe for automation. Finance, despite being a highly intelligent sector, often sees that intelligence is misspent on routine tasks. Generative AI can realign this balance, reducing costs, enhancing service quality, and building trust. Whether private equity, asset management, or commercial banking, AI can streamline processes, delivering an efficiency boost that significantly enhances customer satisfaction. In private equity, the automation possibilities could reshape the sector, bringing it closer to the public markets. In asset management and banking, cost reduction and service enhancement could lead to a dramatic rise in customer satisfaction.
Concrete Use Cases of Generative AI in Finance
So, how does this look in practice? Generative AI can automate numerous finance tasks, including creating reports, verifying information, summarizing news or earnings calls, and even making internal data searchable. Imagine a system that can help asset managers match various types of datasets based on a user query in natural language, thereby making data access and interpretation simpler. This could vastly improve the user experience with business software, reducing effort and time spent. While some applications, like a fully automated financial advisor or AI-led trading and hedging, might present more significant challenges, their potential benefits could revolutionize these sectors.
Overcoming Roadblocks
Of course, every transformation comes with challenges. The key is discerning which use cases are suited for full automation and which require human oversight. Data privacy concerns will also influence decisions about whether to use proprietary or open-source models. There will inevitably be resistance to change within organizations, but the 'Google test' can help navigate data privacy issues: if an employee would conduct that search or share that data on Google, it's likely safe to share with a proprietary Generative AI solution.
Generative AI and Risk Mitigation
Risk mitigation strategies can greatly benefit from generative AI. From detecting and preventing fraud to managing market risks, generative AI can verify identities, cross-reference databases, and analyze vast amounts of data. For instance, at SESAMm, we're developing an ESG controversy detection solution that can be an invaluable tool for risk mitigation.
Improving Investment Decision-Making
Generative AI’s ability to process and analyze massive amounts of data accurately and quickly makes it a formidable tool for investment decision-making. By identifying patterns, trends, and correlations that humans might miss, generative AI can provide a more comprehensive, data-driven perspective, aiding portfolio optimization, asset allocation, and investment risk management.
Streamlining Operations for Efficiency
Generative AI's efficiency and accuracy promise to transform financial institutions. By automating time-consuming tasks like report generation and client communication, AI can free employees to focus on strategic tasks that require critical thinking. Imagine being able to ask complex questions to your banking app in natural language and getting immediate, accurate responses. Such high-quality service was unthinkable a few years ago, but with generative AI, it's within our reach. The future of the financial sector is undeniably tied to the successful implementation of generative AI solutions. The potential is vast, the challenges are surmountable, and the rewards are great. In my view, generative AI is the key to a more efficient, cost-effective, and customer-centric financial sector.
To learn more about SESAMm’s innovative solutions and how we’re pushing the boundaries with generative AI, read the second part of this series here.
Reach out to SESAMm
TextReveal's web data analysis of over five million public and private companies is essential for keeping tabs on ESG investment risks. To learn more about how you can analyze web data or request a demo, contact one of our representatives.
In private equity, as in most industries, decision-making counts on accessing accurate and valuable information. However, these firms often encounter significant challenges when sourcing reliable data, especially when dealing with small, private companies. This article dives into the complexities of identifying high-quality information on smaller companies and underscores its value in investment decisions, operational efficiency, and risk management. It also explores how advanced artificial intelligence (AI) technologies are revolutionizing the identification of these risks, leading to higher rewards and more secure investments, thus providing a competitive edge.
The challenge of identifying valuable information for Smaller Firms
Lack of valuable data
Sturgeon's Law, which states that "Ninety percent of everything is crap (or noise)," becomes particularly relevant in the context of data sourcing. For private equity and investment firms focused on small companies, finding the golden nuggets of information amid the overwhelming amount of digital noise can be daunting. The data available on these companies is often sparse, fragmented, and difficult to uncover using conventional methods. This scarcity of reliable information makes it challenging for private equity firms to make informed decisions, heightening the risk of overlooking critical issues that could impact their investment process.
The difficulties extend beyond just locating information. Many small companies operate without a significant online presence or may not be required to disclose as much information as publicly traded firms. This lack of transparency can further blur critical data points. Furthermore, the data that is available is often unstructured, residing in various forms such as social media posts, obscure local news articles, or industry-specific reports. Extracting meaningful insights from these disparate sources requires sophisticated data processing capabilities, which traditional methods often lack. As a result, private equity firms are left with a significant challenge: how to separate valuable data from the noise without missing critical risk indicators, thereby optimizing their deal sourcing and investment strategies.
Diverse language and terminology
Smaller firms frequently face existential risks, and the potential rewards for identifying these risks early on can be significant for the private equity firms that invest in them. However, mainstream methods of risk identification often fall short, as these companies may not use standardized language to describe materiality. Instead, risks are discussed in varied and context-specific ways, complicating the task of recognizing relevant information. Therefore, it is essential to adopt a specialized approach that analyzes and decodes these firms' unique terminologies and business idiosyncrasies, ultimately translating them into a standardized language that can be effectively used in risk assessment.
The diversity in language is not just a barrier to risk identification but also to the communication of these risks within and between private equity firms. When a small firm uses industry-specific jargon or localized expressions to describe potential threats, it can lead to misunderstandings or underestimations of the actual risk. For instance, a manufacturing startup in a developing country might describe supply chain disruptions in terms that do not translate easily to a global investor’s risk framework. Additionally, cultural differences in how risk is perceived and reported can lead to further complications. This linguistic diversity necessitates the use of advanced natural language processing tools that can interpret data through a common lens while considering industry-specific contexts. For an insurance company, understanding financial models, insurance principles, and regulatory frameworks is crucial. Conversely, assessing risks for a beauty company requires a focus on product safety, consumer preferences, and market trends. By appreciating the specific contexts of each industry, private equity firms can better identify and evaluate potential risks, enhancing decision-making processes, risk and portfolio management strategies, and operational efficiency.
The dynamic nature of the industries themselves further complicates the challenge. For example, the tech industry evolves rapidly, with new risks emerging as technologies develop and consumer expectations shift. What might be considered a negligible risk today could become a significant issue tomorrow as regulatory landscapes, market conditions, and technological advancements alter the playing field. In contrast, industries like agriculture or real estate might have more stable risk profiles but are subject to sudden changes due to environmental factors or policy shifts. This variability across industries means that a one-size-fits-all approach to risk assessment is inadequate. Private equity firms must adopt flexible, industry-specific risk models that can adapt to the unique characteristics and evolving landscapes of the sectors they invest in, thus optimizing their AI capabilities.
The Power of AI in Enhancing Risk Management in Small Firms
AI technologies, particularly natural language processing (NLP) and machine learning algorithms, are important tools for private equity firms aiming to monitor and manage risks in small firms. These technologies can sift through vast amounts of data, extracting the valuable 10% and identifying patterns, trends, and subtle nuances in the language used to describe risks. By detecting these patterns, AI can reveal potential risks that might not be immediately apparent through traditional methods. This proactive approach to risk identification allows firms to address issues before they escalate, providing a more comprehensive and nuanced understanding of the risks facing small firms.
AI's ability to process unstructured data is particularly valuable in this context. Many of the risks that small firms face are discussed informally in places like social media, niche blogs, or local news outlets. Traditional risk management tools might overlook these sources, but AI-powered tools can analyze them in real-time, detecting emerging threats as they develop. Moreover, AI can cross-reference these insights with structured data from financial reports, regulatory filings, and other formal documents to create a holistic risk profile. This multidimensional analysis helps private equity firms not only identify risks but also understand their potential impact, enabling more informed, data-driven decision-making that enhances operational efficiency and competitive edge.
Beyond risk identification, AI also enhances risk mitigation strategies. By continuously monitoring data and learning from new information, AI systems can adapt to changing conditions, offering updated risk assessments that reflect the latest developments. This dynamic approach allows private equity firms to stay ahead of potential issues, making it possible to implement preventative measures rather than reacting to crises after they occur. In this way, AI capabilities contribute significantly to the optimization of risk management processes.
How SESAMm’s Advanced Technology Enhances Risk Assessment
SESAMm’s TextReveal® is at the forefront of this technological revolution, enabling private equity firms to efficiently navigate the vast digital landscape and extract the crucial information needed for informed decision-making. Through our proprietary data lake amounting to over 25 billion online articles with 15 years of historical data and our AI algorithms, TextReveal® can quickly identify and retrieve valuable insights, even when the information is deeply buried or highly specific. The tool's ability to analyze and understand the diverse language and terminology used in discussions about risks on the web empowers private equity firms to objectively assess the materiality of certain risks or identify emerging threats that have yet to be formally recognized.
TextReveal® goes beyond merely identifying risks—it categorizes them, providing context that helps private equity firms understand the severity and relevance of each risk. For example, if a small biotech firm is mentioned in discussions about regulatory hurdles, TextReveal® can determine whether these mentions are isolated incidents or part of a broader trend. It can also assess whether the language used suggests an imminent threat or a longer-term concern, enabling firms to prioritize their responses accordingly. Additionally, TextReveal® integrates sentiment analysis, which can gauge the overall tone of discussions surrounding a company, offering further actionable insights into potential reputational risks.
SESAMm has developed a proprietary metric – the Intensity Score, which calculates an event's relevance based on its news coverage and sentiment. It uses negative sentiment, article dispersion, and empirical ESG risk measures to determine how likely an article is to represent a high-risk controversy. The Intensity Score gives TextReveal users a clear understanding of which events require their attention.
Users can also opt to receive email alerts for the more severe controversies, ensuring they’re always aware of significant risks. In addition to the severity, controversies are also categorized by risk and sub–risk type, making it easy to analyze specific areas of concern.
Moreover, SESAMm's platform is designed to be intuitive and user-friendly, making it accessible to investment professionals who may not have a technical background. This ease of use ensures private equity firms can quickly incorporate AI-driven insights into their risk management processes without a steep learning curve. By streamlining the data analysis process, TextReveal® allows firms to focus on strategic decision-making, confident they have a comprehensive understanding of the risks and opportunities associated with their investments and portfolio companies. This level of operational efficiency and optimization is key to maintaining a competitive edge in the fast-paced world of private equity.
TextReveal’s Risk Assessment module enables deep company and thematic research in multiple languages through on-the-fly keyword searches. Users have full access to articles, sentiment analysis, and trending topics to get a complete understanding of the risks. We’ve even developed an AI Text Summary feature that provides a quick summary of a selected article, saving time and enabling a faster analysis.
In summary, the integration of AI tools and natural language processing technologies is transforming risk management in private equity, particularly for firms dealing with small, private companies. By leveraging these advanced tools, private equity firms can enhance their due diligence processes, better monitor risks and controversies, and ultimately make more informed investment decisions that lead to higher rewards and operational efficiency.
Reach out to SESAMm
TextReveal's web data analysis of over five million public and private companies is essential for keeping tabs on ESG investment risks. To learn more about how you can analyze web data or request a demo, contact one of our representatives.
Financial and ESG insights begin with big data coupled with data science.
At SESAMm, our artificial intelligence (AI) and natural language processing (NLP) platform analyzes text in billions of web-based articles and messages. It generates investment insights and ESG analysis used in systematic trading, fundamental research, risk management, and sustainability analysis.
This technology enables a more quantitative approach to leveraging the value of web data that is less prone to human bias. It addresses a growing need in public and private investment sectors for robust, timely, and granular sentiment and environment, social, and governance (ESG) data. This article will outline how the data is derived and illustrate its effectiveness and predictive value.
Content coverage and ESG data collection
The genesis of SESAMm’s process is the high-quality content that comprises its data lake, the source from which it draws its insights. SESAMm scans over four million data sources rigorously selected and curated to maximize coverage of both public and private companies. Three guiding criteria—quality, quantity, and frequency—ensure a consistently high input value.
Every day the system adds millions of articles to the 16 billion already in the data lake, going back to 2008. The coverage is global, with 40% of the sources in English (the U.S. and international) and 60% in multiple languages. The data lake, expanding every month, comprises over 4 million sources, including professional news sites, blogs, social media, and discussion forums.
The following tables illustrate SESAMm’s data lake distribution (Q1 2022):
Respect for personal privacy figures highly in the data gathering process. We don’t capture personal data, like personally identifiable information (PII), and respect all website terms of service and global data handling and privacy laws. SESAMm’s data also doesn’t contain any material non-public information (MNPI).
Deriving financial signals and ESG performance indicators
SESAMm’s new TextReveal® Streams platform applies NLP and AI expertise to process the premium quality content gathered in its data lake. This complex process involves named entity recognition (NER) and disambiguation (NED)—the process of identifying entities and distinguishing like-named entities using contextual analysis—and mapping the complex interrelationships between tens of thousands of public and private entities, connecting companies, products, and brands by supply chain, location, or competitive relationship.
Process representation for NER and NED
Using SESAMm’s TextReveal Streams, this wealth of information is filtered to focus on four crucial contexts for systematic data processing, risk management, and alpha discovery:
Sentiment covering major global indices: world equities (and Small Caps, Emerging), U.S. 3000, Europe 600, KOSPI 50, Japan 500, Japan 225
Sentiment covering all assets and derivatives traded on the Euronext exchange
Private company sentiment on more than 25,000 private companies
ESG risks covering 90 major environmental, social, and governance risk categories for the entire company universe, which includes more than 10,000 public and more than 25,000 private companies with worldwide coverage
TextReveal Streams data sets and assessments are used by financial institutions, rating agencies, and the financial services sector, such as hedge funds (quantitative and fundamental) and asset managers, to optimize trade timing and identify new sustainable investment opportunities. Private equity deal and credit teams also use the data for deal sourcing and due diligence. Private equity ESG teams use it to manage initiatives like portfolio company environmental, social, and governance risk and reporting.
Methodology and technology for processing unstructured data
NLP workflow, from data extraction to granular insight aggregation
Data is continually extracted from an expanding universe of over four million sources daily. As it enters the system, it is time-stamped, tagged, indexed, and stored in our data lake to update a point-in-time history extending from 2008 to the present. The source material is then transformed from raw, unstructured text data into conformed, interconnected, machine-readable data with a precise topic.
NLP workflow for TextReveal Streams
Mapping relationships between entities with the Knowledge Graph
At the heart of the text analytics process is SESAMm’s proprietary Knowledge Graph, a vast map connecting and integrating over 70 million related entities and their keywords. It’s essentially a cross-referenced dictionary of keywords, relating each organization to its brands, products, associated executives, names, nicknames, and their exchange identifiers in the case of public companies.
Entities within the Knowledge Graph are updated weekly and tagged to ensure changes are correctly tracked. The CEO of a company today, for example, may not be the CEO tomorrow, and brands may be bought and sold, changing the parent company with each sale. Weekly updates within the Knowledge Graph ensure the system is aware of these changes.
Named entity disambiguation (named entity recognition plus entity linking) is one of the NLP techniques used to identify named entities in text sources using the entities mapped within the Knowledge Graph universe.
At SESAMm, NED identifies named entities based on their context and usage. Text referencing “Elon,” for example, could refer indirectly to Tesla through its CEO or to a university in North Carolina. Only the context allows us to differentiate, and NED considers that context when classifying entities. This method is superior to simple pattern matching, limiting the number of possible matches, requiring frequent manual adjustments, and cannot distinguish homophones.
SESAMm uses three other NLP tools to identify entities and create actionable insights. These are lemmatization, embeddings, and similarity. Each is explained in more detail below.
Analyzing the morphology of words with lemmatization
News articles, blog posts, and social media discussions reference organizations and associated entities in various forms and functions. Lemmatization seeks to standardize these references so the system knows they mean the same thing.
For example, “Tesla,” “his firm,” “the company,” and “it” are all noun phrases that can appear in a single article and refer to a single entity. Even where the reference is apparent, it can take different forms. For example, “Tesla” and “Teslas” both refer to the same entity but have slightly different meanings (semantics) and shapes (morphology).
The lemmatization process standardizes reference shape (morphology) to facilitate identification and aggregation. Lemmatization is a more sophisticated process than stemming, which truncates words to their stem and sometimes deletes information.
Encoding context and meaning with word embedding
In NLP, embedding is a numerical representation of a word that enables its manifold contextual meanings to be calculated relationally. Embeddings are typically real-valued vectors with hundreds of dimensions that encode the contexts in which words appear and, thus, also encode their meanings. Because they are vectors in a predefined vector space, they can be compared, scaled, added, and subtracted. An example of how this works is that the vector representations of king and queen bear the same relation to each other as the representations of man and woman once you subtract the vector that represents royal.
Vectorized representation of embeddings
Using embedding is key to analyzing how words change meaning depending on context and understanding the subtle differences between words that refer to the same concept: synonyms. For example, the words business, company, enterprise, and firm can all refer to the same thing if the context is “organizations.” But they represent different things and even different parts of speech if the context changes.
In the phrase, “[Tesla] will be by far the largest firm by market value ever to join the S&P,” for example, one could replace the word firm with company or enterprise without affecting the meaning significantly. Contrast that with “a firm handshake,” where a similar substitution would render the phrase meaningless.
Also, words referring to the same concept can emphasize slightly different aspects of the concept or imply specific qualities. For example, an enterprise might be assumed to be larger or to have more components than a firm. Embeddings enable machines to make these subtle distinctions.
One advantage of using embedding is that it’s practical because it’s empirically testable. In other words, we can look at actual usage to determine what a word means.
Another advantage is that embeddings are computationally tractable. This understanding of a word’s definition allows us to transform words into computation objects to programmatically examine the contexts in which they appear and, thus, derive their meaning.
As lemmatization is an improvement on stemming, embeddings improve techniques such as one-hot encoding, which is close to the common conception of a definition as a single entry in a dictionary.
SESAMm uses the global vectors for word representation (GloVe) algorithm to generate embeddings. It’s an unsupervised learning algorithm that begins by examining how frequently each word in a text corpus co-occurs with other words in the same corpus. The result is an embedding that encapsulates the word and its context together, allowing SESAMm to identify specific words in a list and different forms of the listed words and unlisted synonyms.
GloVe is an extension of recent approaches to vector representation, combining the global statistics of matrix factorization techniques like latent semantic analysis (LSA) with the local context-based learning of word2vec. The result is an unsupervised algorithm that performs well at capturing meaning and demonstrating it on tasks like calculating analogies and identifying synonyms.
BERT is another algorithm used by SESAMm to generate embeddings. BERT produces word representations that are dynamically informed by the words around them. Google developed the technique, and it’s what’s known as a transformer-based machine learning technique, which means it doesn’t process an input sequence token by token but instead takes the entire sequence as input in one go. This technique is a significant improvement over sequential recurrent neural network (RNN) based models because it can be accelerated by graphics processing units (GPUs).
SESAMm uses BERT for multilingual NLP of its extensive foreign language text because it has been retained using an extensive library of unlabeled data extracted from Wikipedia in over 102 languages. BERT model was trained to predict words from context and next sentence prediction where it was trained to predict if a chosen following sentence was probable or not given the first sentence. As a result of this training process, BERT learned contextual embeddings for words. Due to this comprehensive pre-training, BERT can be finetuned with fewer resources on smaller datasets to optimize its performance on specific tasks.
Linking words, sentences, and topics with cosine similarity
Cosine similarity with centered means it’s identical to the correlation coefficient, which highlights another element of the computational tractability of the embeddings approach. It makes it easy to compare words and contexts for similarity.
Converting words to vector representations means we can quickly and easily compare word similarity by comparing the angle between two vectors. This angle is a function of the projection of one vector onto another. It can identify similar, opposite, or wholly unrelated vectors, which allows us to compute the similarity of the underlying word that the vector represents.
Two vectors aligned in the same orientation will have a similarity measurement of 1, while two orthogonal vectors have a similarity of 0. If two vectors are diametrically opposed, the similarity measurement is -1. In practice, negative similarities are rare, so we clip negative values to 0.
Vectorized representation of cosine similarities
Cosine similarity measures whether two words, sentences, or corpora are close to one another in vector space or “about” the same thing in semantic space. To answer the question, “Is this sentence referencing company X?” we embed the sentence using the process described above and compute the cosine similarity between the sentence and the embedded company profile. Analogously, we compute similarities between sentences and the ESG topics SESAMm monitors by taking the maximum similarity between a sentence and each embedded keyword associated with an ESG topic.
These similarities allow us to identify whether a sentence references fraud, tax avoidance, pollution, or any other ESG risk topic among the more than 90 that SESAMm tracks across the web.
Similarities within ESG topics combine with word counts to resolve the recall and precision problem. Word counts are precise because if a word is identified within a context, then that context, by construction, references the topic.
The virtue of using these NLP techniques is that even if a given keyword list does not include every possible combination of words that a person might use to discuss a topic, relevant entities missed by the word-count process will be identified through vector similarity.
This is the power of SESAMm’s NLP expertise. We can scan many lifetimes’ worth of data in seconds to find the concepts you explicitly ask for and the concepts relevant to your search but that you did not think of yourself.
Sentiment analysis with deep learning and neural networks
Once we’ve identified the concepts and contexts of interest in all the forms they appear, we analyze the context to determine the speakers’ attitudes.
We use sentiment classification models to score a sentence with three possible outcomes: negative, neutral, or positive. The current classification models are based on deep learning AI technologies. Specifically, we stack convolutional neural networks with word embeddings and bayesian optimized hyperparameters—parameters not learned during training. This architecture improves the accuracy and enables fast shipping of production-ready models for a given language. We also produce state-of-the-art frameworks with architecture variations enabling multilingual capabilities, such as transformers and universal sentence encoders.
Condensing information and extracting insights with daily aggregation
Similarities, embedded word counts, and sentiment are state-of-the-art tools for processing unstructured text data. The same tools are effective cross-linguistically.
Once the information has been extracted from millions of data points, it’s aggregated and condensed into actionable insights.
All entities are referenced directly or indirectly within an article. Then, sentence-level references are aggregated to obtain an article-level perspective, and finally, all relevant articles are aggregated to gain an entity-level view of that day.
In this way, reams of data are compressed into several metrics to provide a daily aggregate view for each entity, highlighting trends at a sentence, article, and entity-level comparable over a multi-year history.
ESG analysis use cases
SESAMm’s TextReveal Streams is used in various investment domains, from asset selection to alpha generation and risk management. Systematic hedge funds track retail interest in real time to identify investment opportunities and protect their existing positions. In the Private Equity industry, equity and credit-deal teams use the data in various ways, from monitoring consumer perspectives via forums and customer reviews for evaluating deal prospects to estimating due diligence risks, all to help make investment decisions. Dedicated teams use our data for monitoring portfolio companies for ESG red flags that conventional ESG reporting might miss.
Below are two examples of how aggregated TextReveal Streams data can be used to help identify investment risk and opportunity.
LFIS CapitalL: ESG signals for equity trading
ESG controversies can significantly impact asset prices in the short term, and it’s now estimated that intangible assets, including a company’s ESG rating, account for 90% of its market value.
Working in partnership with LFIS Capital (LFIS), a quantitative asset manager and structured investment solutions provider, SESAMm developed machine learning and NLP algorithms that could analyze ESG keywords in articles, blogs, and social media, to generate a daily ESG score specific to each stock, which is part of the TextReveal Streams’ platform’s core functionality.
The results were promising when these scores were incorporated into a simulated strategy for trading stocks in the Stoxx600 ESG-X index.
A simulated long-only strategy running between 2015 and 2020, using the signals, delivered a 7.9% annualized return, 2.9% higher than the benchmark for similar annualized volatility (17.3% vs. 17.1%). The information ratio of the strategy was greater than 1, with a tracking error of 2.8%. Results for the previous three years were compelling, reflecting the growing interest and news flow around ESG themes.
Researchers also backtested a hypothetical long-short strategy for all stocks in the Stoxx600 ESG-X index with a market cap of over $7.5bn. This investment strategy delivered a Sharpe ratio of approximately 1 with annualized returns and volatility of 6.1% and 5.9%, respectively, between 2015 and 2020. Like the long-only strategy, returns were particularly robust over the three years up to 2020: +6.0% in 2018, +7.3% in 2019, and +11.3% in 2020.
Finally, a simulated “130/30” ESG strategy that combined 100% of the long-only ESG strategy and 30% of the long-short ESG strategy delivered a 10.8% annualized return, 5.8% higher than that of the Stoxx600 ESG-X index. Annualized volatility was similar at 16.9% vs. 17.1%. The strategy experienced a tracking error of 3.8% and an information ratio of over 1.5, with a consistent outperformance each year.
Disclaimer: Past performance is not an indicator of future results. Theoretical calculations are provided for illustrative purposes only. The investment theme illustrations presented herein do not represent transactions currently implemented in any fund or product managed by LFIS.
Wirecard: ESG sentiment and volume as predictive indicators
The Wirecard scandal broke on June 21, 2020, when newswires carried the story that the major German payment processor had filed for bankruptcy after admitting that €1.9 billion ($2.3 billion) of purported escrow deposits did not exist.
Could SESAMm’s TextReveal Streams platform have provided investors with an early warning that the scandal was about to break?
The following chart derived from the platform shows how key ESG metrics, including ESG scores (volumes) and ESG scores (sentiment), reacted to the news.
An analysis of the charts pinpoints a shallow rise in the ESG scores (volumes) time series in the early part of June before the eruption on June 21.
The ESG scores (sentiment) metric also shows a steady increase in negative sentiment for governance, the most relevant of the three ESG factors regarding the scandal.
How key ESG metrics, including ESG scores (volumes) and ESG scores (sentiment), reacted to the Wirecard scandal news.
Additionally, before the crash, governance was the most negative of the three ESG factors most of the time. This was especially the case from late March to early April, and then before the scandal in early June, negative governance sentiment diverged higher from the other two.
The rate-of-change of negative governance sentiment as it rose and peaked in early June before the scandal broke was also extremely high, perhaps providing the basis for an early warning signal.
Portfolio managers who had been keeping an eye on the reputational slide in Governance for Wirecard may have decided the company was at high risk of a negative controversy emerging, giving them cause to drop the stock before the event.
In this way, it can be seen how while not providing a hard and fast early warning signal, SESAMm’s ESG scores can, nevertheless, be used as the basis for developing a data-driven, rules-based portfolio management approach that can help investors avoid high-risk candidates like Wirecard.
SESAMm takes on ESG data challenges
SESAMm’s NLP and AI tools analyze over four million data sources daily to identify thousands of public and private companies and their related products, brands, identifiers, and nicknames, turning reams of unstructured text into structured and actionable data.
SESAMm’s TextReveal Streams platform can be used in many quantitative, quantamental, and ESG investment use cases. TextReveal is a solution that allows you to fully leverage NLP-driven insights and receive high-quality results through data streams, modular API and dashboard visualization, and signals and alerts.
Learn how SESAMm can support you in your investment decision-making and request a demo today.
To request a demo or for access to the full SESAMm Wirecard or LFIS reports, contact us here:
On June 17, SESAMm joined WeeFin and EthiFinance to co-host ESG Version Française, a morning of discussion held at Caisse des Dépôts in Paris. The event brought together investors, asset managers, and ESG leaders from across the French and European financial sector to address a question with real consequences for how ESG risk gets measured: how do you support European sovereignty in ESG evaluation, ratings, and analysis?
The starting point for the morning was a striking figure: roughly 60% of the global ESG data market is controlled by US-based providers. For European investors, that concentration creates a structural dependency that shapes methodologies, tooling, and ultimately investment decisions. ESG Version Française set out to move that conversation from diagnosis to action.
Opening: Why ESG Sovereignty Is a Competitiveness Issue
The morning opened with a conversation between Grégoire Hug, co-founder and CEO of WeeFin, and Josselin Kalifa, Director of the Investment Management Department at Caisse des Dépôts. Their discussion framed ESG sovereignty not as an abstract policy goal but as a competitiveness issue for European financial institutions, one that touches data access, methodological independence, and the ability to set standards rather than simply follow them.
Methodology: French ESG Expertise as an Asset
The first panel, moderated by Carol Sirou, CEO of EthiFinance, examined what makes French ESG expertise distinctive and how those strengths can serve European sustainable finance more broadly. Panelists discussed the specific methodological rigor that French providers and analysts bring to ESG assessment, and why that expertise is a genuine asset rather than a niche approach, at a moment when European institutions are being asked to justify their ESG frameworks with more scrutiny than ever.
Private Assets and Technology: New Frontiers for ESG Analysis and Risk Management
The second panel, titled "Actifs privés et technologie : quelles nouvelles frontières pour l'analyse ESG et la gestion des risques ?" (Private Assets and Technology: What New Frontiers for ESG Analysis and Risk Management?), was moderated by Sylvain Forté, CEO of SESAMm. The discussion focused on private assets specifically, the segment of the market where ESG data has traditionally been hardest to access, and on how AI driven technology is opening new frontiers for covering, analyzing, and monitoring risk on private companies at scale. Panelists discussed what these tools make possible today, from broader coverage of private assets to faster, more transparent risk analysis, and why technological capability is now central to the sovereignty conversation, not separate from it.
The morning closed with a final round of questions over coffee, giving attendees a chance to continue theconversation informally with speakers and fellow investors.
Speakers
Josselin Kalifa
Director of the Investment Management Department Caisse des Dépôts
Pascale Forde Maurice
Head of Corporate Europe, Sustainable Banking Crédit Agricole CIB
Marie-Pierre Peillon
Director of Research and ESG Strategy Groupama Asset Management
Maud Colin Livet
Head of SRI Fonds de Garantie des Victimes
Cécile Goubet
Managing Director Institut de la Finance Durable
Gaëlle Achdjian
Principal ESG & Sustainability Access Capital Partners
Abigail Arellano
Sustainability Project Manager & Data Specialist Natixis Investment Managers
Grégoire Etienne
CEO Greenscope
Grégoire Hug
Co-founder and CEO WeeFin
Sylvain Forté
CEO SESAMm
Carol Sirou
CEO EthiFinance
Stay ahead with the latest in ESG and AI intelligence
Join our mailing list to receive new reports, event invites, and updates from SESAMm directly to your inbox.