Return to Blog

July 6, 2023

Are we there yet? “Hackathon mode” and the pursuit of reliability with LLMs

Human comprehensibility leads to human ingenuity.  

I don’t want to understate the scale of the technical breakthrough that OpenAI has achieved with GPT-3 and GPT-4. But it’s easy to miss how much work the “chat” part of the phrase ChatGPT is doing.  

It’s an overlooked insight that the flowering of amazement associated with the current generations of large language models is directly linked to the user-experience breakthrough of allowing anyone to chat directly with the model.

Recursively, OpenAI used its own model to make its model legible to the world.  

This is a very FUN breakthrough.  

Three weeks ago, I asked ChatGPT to write an AML policy for a fintech (generic but passable) and then I asked it to translate the policy into Japanese. Interestingly, it had to be coaxed into completing the task (continue generating), but the contrast to my own painstaking efforts to translate much less complex material when I studied Japanese in college was a joy in and of itself.  

Last week, my seventh grader wanted more fill-in-the-blank practice problems related to complex verb tenses – do you remember what the past perfect progressive tense is? GPT does, and was much quicker to generate practice problems than I would have been. (GPT is both creative and a fast typer!) Somewhat hilariously, despite three separate prompts, GPT kept on putting the correct answers right after the sentence – almost like it couldn’t resist that urge to demonstrate its own competence much like a certain 7th grader!  

In a recent survey, I found that most people (56%) are using GPT this way – really just playing around.  

But in startup land, the real value is what I would call “hackathon mode.”  In hackathon mode, a single, usually junior developer builds a novel feature that they don’t have relevant experience with. In a day or two, the build is done whereas before that build would have taken a small team weeks or even months. Nearly 37% of respondents to my survey say that they are using ChatGPT this way, including many who otherwise describe themselves as non-technical.

This is a very EXCITING breakthrough.  

Largely because of “hackathon” mode, at QED we believe that the same companies who are most likely to benefit from the LLM/generative AI moment are those that are most threatened by it.  

Usually that threat is framed based on the potential that new “open” models are going to be so powerful that they will obviate the need for specialized models and will erode the power of proprietary data advantages. The power of GPT-4 in particular suggests a kind of reasoning ability, though the underlying model does not include any explicit logical engine. Some are projecting that this common sense is going to wipe out narrowly specialized AI and other vertical SaaS.  But my sense is that we are so awed by the passing of the Turing test that we’re overlooking how difficult non-language tasks can be.

Predictions can be hard, especially about the future, but most of the data in the world is actually not public and not accessible to LLMs. This is especially true of financial data, so I think we are still years away from LLMs or other publicly trained models displacing the value of proprietary data, training informed by domain expertise, and managed feedback loops. Finance may be a particularly hard nut to crack here, because  will still be necessary to achieve reliable performance in any domain where the downsides to wrong decisions are high.  LLMs must be combined with specialized models at least for this next stage of implementations.  

But there is another pernicious dynamic, namely that the excitement around “hackathon mode” causes people to overlook obvious drivers of value and business thresholds that must be passed in order for software to be effective. Overrating their own ability to “build it quickly” and underrating the challenges of perfecting and maintaining a feature.

In financial services, the point of any system is to have a reliable representation of the world in the form of data.  In underwriting or fraud, a mistake costs real money – the quantifiable impact of a mistake changes, quite literally, by an order of magnitude each time it moves one digit to the left or the right. Automation and straight through processing requires accuracy. I once had a founder pitch me that their solution was 100% accurate 80% of the time!

Two of my companies, Ocrolus and Ntropy announced a partnership that is not only a triumph over “hackathon mode” for each of them, but also allows their customers to get the benefits of proprietary data and fit for purpose models.  

Both Ocrolus and Ntropy have dedicated data science and machine learning teams – AI is what both of them DO – and they also recognize that GPT-4 would enable them each to create rudimentary, “slide-ware” versions of each others’ products in a few weeks. But that having “slide-ware” is not good enough for lending and fraud fighting. These functions require Ocrolus and Ntropy’s customers to make predictions, so mistakes will happen. But mistakes at the level of data input must be squeezed out. .  And since both companies build accuracy testing into their product by default, they recognize that excellence is worth partnering for.  

Ntropy's core is ingesting transaction data feeds and enriching this data to provide lenders a deeper view on their customers’ behavior. While lenders can get this data from aggregators, many borrowers continue to prefer providing PDF documents directly, and unfortunately aggregation services break more often than many realize – many lenders find that 20-80% of their pipeline requires PDF documents.  

Using Ntropy and Ocrolus together allows lenders not only to use a unified pipeline for digital and non-digital applicants, it also allows them to fully take advantage of their marketing and sales – no longer wasting spend on applicants stopped by an unnecessary friction in the underwriting process.  

For Ntropy, the decision to partner with Ocrolus should have been easy (I’m on the board of both companies!), but they went through exhaustive testing on latency and accuracy and Ocrolus was still the right choice. Their human-in-the-loop approach provides exactly the guard rails that companies need to embed AI into a core business process.

So, while I’m as excited about our eventual destination as anyone, like any dad in the front seat of this summer season, my answer to the inevitable question “are we there yet?” is “not yet”.  Value still matters. Proprietary data, models that are fit for purpose, and teams that thoughtfully include LLMs to improve margin without sacrificing quality are the recipe for success.

Human comprehensibility leads to human ingenuity.  

I don’t want to understate the scale of the technical breakthrough that OpenAI has achieved with GPT-3 and GPT-4. But it’s easy to miss how much work the “chat” part of the phrase ChatGPT is doing.  

It’s an overlooked insight that the flowering of amazement associated with the current generations of large language models is directly linked to the user-experience breakthrough of allowing anyone to chat directly with the model.

Recursively, OpenAI used its own model to make its model legible to the world.  

This is a very FUN breakthrough.  

Three weeks ago, I asked ChatGPT to write an AML policy for a fintech (generic but passable) and then I asked it to translate the policy into Japanese. Interestingly, it had to be coaxed into completing the task (continue generating), but the contrast to my own painstaking efforts to translate much less complex material when I studied Japanese in college was a joy in and of itself.  

Last week, my seventh grader wanted more fill-in-the-blank practice problems related to complex verb tenses – do you remember what the past perfect progressive tense is? GPT does, and was much quicker to generate practice problems than I would have been. (GPT is both creative and a fast typer!) Somewhat hilariously, despite three separate prompts, GPT kept on putting the correct answers right after the sentence – almost like it couldn’t resist that urge to demonstrate its own competence much like a certain 7th grader!  

In a recent survey, I found that most people (56%) are using GPT this way – really just playing around.  

But in startup land, the real value is what I would call “hackathon mode.”  In hackathon mode, a single, usually junior developer builds a novel feature that they don’t have relevant experience with. In a day or two, the build is done whereas before that build would have taken a small team weeks or even months. Nearly 37% of respondents to my survey say that they are using ChatGPT this way, including many who otherwise describe themselves as non-technical.

This is a very EXCITING breakthrough.  

Largely because of “hackathon” mode, at QED we believe that the same companies who are most likely to benefit from the LLM/generative AI moment are those that are most threatened by it.  

Usually that threat is framed based on the potential that new “open” models are going to be so powerful that they will obviate the need for specialized models and will erode the power of proprietary data advantages. The power of GPT-4 in particular suggests a kind of reasoning ability, though the underlying model does not include any explicit logical engine. Some are projecting that this common sense is going to wipe out narrowly specialized AI and other vertical SaaS.  But my sense is that we are so awed by the passing of the Turing test that we’re overlooking how difficult non-language tasks can be.

Predictions can be hard, especially about the future, but most of the data in the world is actually not public and not accessible to LLMs. This is especially true of financial data, so I think we are still years away from LLMs or other publicly trained models displacing the value of proprietary data, training informed by domain expertise, and managed feedback loops. Finance may be a particularly hard nut to crack here, because  will still be necessary to achieve reliable performance in any domain where the downsides to wrong decisions are high.  LLMs must be combined with specialized models at least for this next stage of implementations.  

But there is another pernicious dynamic, namely that the excitement around “hackathon mode” causes people to overlook obvious drivers of value and business thresholds that must be passed in order for software to be effective. Overrating their own ability to “build it quickly” and underrating the challenges of perfecting and maintaining a feature.

In financial services, the point of any system is to have a reliable representation of the world in the form of data.  In underwriting or fraud, a mistake costs real money – the quantifiable impact of a mistake changes, quite literally, by an order of magnitude each time it moves one digit to the left or the right. Automation and straight through processing requires accuracy. I once had a founder pitch me that their solution was 100% accurate 80% of the time!

Two of my companies, Ocrolus and Ntropy announced a partnership that is not only a triumph over “hackathon mode” for each of them, but also allows their customers to get the benefits of proprietary data and fit for purpose models.  

Both Ocrolus and Ntropy have dedicated data science and machine learning teams – AI is what both of them DO – and they also recognize that GPT-4 would enable them each to create rudimentary, “slide-ware” versions of each others’ products in a few weeks. But that having “slide-ware” is not good enough for lending and fraud fighting. These functions require Ocrolus and Ntropy’s customers to make predictions, so mistakes will happen. But mistakes at the level of data input must be squeezed out. .  And since both companies build accuracy testing into their product by default, they recognize that excellence is worth partnering for.  

Ntropy's core is ingesting transaction data feeds and enriching this data to provide lenders a deeper view on their customers’ behavior. While lenders can get this data from aggregators, many borrowers continue to prefer providing PDF documents directly, and unfortunately aggregation services break more often than many realize – many lenders find that 20-80% of their pipeline requires PDF documents.  

Using Ntropy and Ocrolus together allows lenders not only to use a unified pipeline for digital and non-digital applicants, it also allows them to fully take advantage of their marketing and sales – no longer wasting spend on applicants stopped by an unnecessary friction in the underwriting process.  

For Ntropy, the decision to partner with Ocrolus should have been easy (I’m on the board of both companies!), but they went through exhaustive testing on latency and accuracy and Ocrolus was still the right choice. Their human-in-the-loop approach provides exactly the guard rails that companies need to embed AI into a core business process.

So, while I’m as excited about our eventual destination as anyone, like any dad in the front seat of this summer season, my answer to the inevitable question “are we there yet?” is “not yet”.  Value still matters. Proprietary data, models that are fit for purpose, and teams that thoughtfully include LLMs to improve margin without sacrificing quality are the recipe for success.

No items found.

Test column heading 1

Test column heading 2

Test column heading 3

Test column heading 4

This is a longer lorem ipsum text 1
This is a longer lorem ipsum text 2
This is a longer lorem ipsum text 3
This is a longer lorem ipsum text 4
This is a longer lorem ipsum text 5
This is a longer lorem ipsum text 6
TEst row
TEst row
TEst row
TEst row
TEst row
TEst row
No items found.
No items found.
No items found.
No items found.

01

Settlement Collapse

Value transfer moves from days - corresponding banking, T+1 securities - to seconds. Working capital tied up in float is released.

02

Cost Collapse

Marginal transaction cost approaches zero: fractions of a cent, versus 1-6% on card and corresponding rails.

03

Programmability

Money becomes an object that carries logic - escrow, splits, rebates, compliance - executed by code, not back offices.

Pure infrastructure with no revenue accrual

Layer-1 chains and general-purpose middleware — outside our circle of competence and typically outside our stage.

Speculative asset creation

NFT platforms, memecoins, prediction markets styled as products — mapping to none of the five functions; structurally uninvestable for us.

Three structural truths cut across all five functions.

(a)

Regulated-first wins

The 2020/21 cycle proved permissionless purity does not survive contact with real financial regulation. GENIUS, MiCA, CLARITY and the UK/Singapore regimes are producing founders who start from “how do we get licensed” and build backwards — precisely the founder profile QED has always preferred.
(b)

Incumbents upgraded, not disintermediated

JPMorgan, Citi, Bank of America and Wells Fargo are jointly building a tokenized-deposit network; Visa launched a stablecoin platform in July 2026; 140+ businesses signed an open stablecoin standard. Banks migrate — and new-generation infrastructure companies own the picks and shovels of that migration.
(c)

Emerging markets feel it first

Every function improves most where the fiat experience is worst: cross-border payments, dollar access, investment product availability, working-capital finance. Those are exactly the geographies where QED has fintech ventures’ deepest footprint. Our geographic distribution is not incidental to the tokenization thesis — it is the thesis.

Trade & working-capital finance

Finkargo( LatAm import finance) and OatFi(B2B working-capital infrastructure) sit directly on flows whose logical settlement layer is stablecoin.

Collateralized digital-asset lending

Tokenized Treasuries, equities and stablecoin holdings as instant, programmable collateral.

On-chain private credt

Maple, Centrifuge and emerging institutional protocols - credit funds migrating to programmable rails.

Why QED is advanced

Credit is QED's craft: distinguishing lending businesses from fintechs pretending to be one, charge-offs earned from charge-offs deferred. On-chain credit is a straight-line extension, not a stretch.