Skip to content
Insights
Case study · Proprietary sourcing

Own the market definition.

Sourcing advantage comes from defining and classifying the market around the mandate, not searching the same generic datasets with the same categories and filters as every other buyer.

  • 9 min read
  • August 2026

Every sourcing route works—until the market gets crowded or the pace has to increase.

Relationships, brokers, events, marketplaces, databases, and manual research all contribute something useful. The problem is that none gives an investor reach, context, control, and repeatability at the same time.

Company databases appear to solve the volume problem, but every buyer starts with the same categories and sees the same obvious companies. For vertical software, the gap was especially clear: a generic software label did not say which market a product served, what work it performed, or whether software was actually the product.

What the market sellsBroad company access

Useful coverage organized by someone else's taxonomy.

What the investor needsA thesis-shaped universe

A market defined in the language of the mandate.

The right method depends on which trade-off matters most.

Relationships can be high-conviction but narrow. Data products can be broad but undifferentiated. The useful question is not which method wins—it is which trade-off the mandate can afford.

Relative comparison of sourcing methods across pipeline, effort, and advantage.
FavorableMixedLimiting
  1. Personal connections

    Trusted context and access. Limited by network reach and senior time.

  2. Broker relationships

    Prepared materials and a structured process. The same opportunity is visible to competing buyers.

  3. Events and conferences

    Live context and serendipitous discovery. Low repeatable volume for the time invested.

  4. Social media

    Inexpensive access to public signals. Weak intent and a high-noise search surface.

  5. Deal marketplaces

    More opportunities that are ready to transact. High buyer competition and generic discovery filters.

  6. Company data products

    Broad, queryable coverage and a fast start. Vendor taxonomy and the same target pool as competitors.

  7. Manual search

    Mandate-specific human interpretation. Analyst time limits coverage and repeatability.

Owned approach

Custom sourcing engine

A private taxonomy and ranking shaped around the mandate. Requires maintained sources, rules, and review.

The only route here that combines broad volume, low competition, and high scalability.

Why data-driven sourcing fails.

The math looks great.

A $20K data product plus $10K–15K per offshore resource makes scale look inexpensive. The model assumes $15K per resource, 10K leads each year, and a 0.01% lead-to-close rate.

Modeled sourcing cost and output at one and five resources
ModelAnnual spendOutputCost / deal
1 resource$35K10K leads1 deal$35K
5 resources$95K50K leads5 deals$19K
The model scales activity. It does not create an advantage.

The underlying yield never changes: 99.99% remains unconverted. More throughput only becomes an advantage when the inputs, market definition, or process are different.

Everyone can run the same play.

Market convergence

Everyone contacts the same companies; the probability of a brokered process increases.

Process degradation

Cheap throughput fills the database with junk, while CRM sync exposes proprietary finds.

Why existing data products fall short.

6–7%of companies labeled Software or IT

Poor classification

Most data products do not differentiate between vertical software and regular software. Of companies classified as ‘Software’ or ‘IT,’ only 6–7% can be considered vertical software.

NA → ?coverage weakens across borders

Poor global coverage

Coverage is especially poor for companies based outside North America, which we attributed to the language barrier.

A ∩ B ∩ Cbasic filters miss the mandate

Poor filtering algorithms

Industry, geography, and product filters rarely work together with enough precision. You cannot reliably pull ‘Healthcare Software in Canada’ without cross-referencing broad categories and cleaning the result by hand.

VIEW ≠ USElimited exports keep data inside the product

Restricted data access

Many products expose far more data in the interface than they allow you to export. Teams are forced to work in batches, repeat queries, and stitch partial results together outside the platform.

Proprietary classification at scale.

The foundation for every company-intelligence product on the market tends to be built from some combination of three major data sources: LinkedIn company profiles, government company registries, and company website data.

The classification work sits on top of that layer rather than in it: multiple sourcing methodologies and a set of proprietary algorithms that resolve, verify and classify the relevant companies out of a universe of tens of millions.

Common source material becomes a mandate-shaped universe.

The complexity lies in reconciling incomplete signals at scale while keeping the resulting market aligned to the mandate.

Mandate-shaped universe

A configurable, verified research surface that remains private to the investor.

  1. Multi-dimensional research

    A configurable research surface reconciles many incomplete and sometimes contradictory signals at once. The complexity stays behind a simple mandate-specific experience.

  2. Proprietary classification

    A mandate-specific classification layer distinguishes companies beyond the broad categories available in conventional data products.

  3. Generative AI scoring

    AI supports mandate-specific review across the factors an investor defines, without replacing the underlying evidence or analyst judgment.

  4. Analyst validation

    Analysts reviewed the classified set before release, and review continues as coverage expands.

From Kiosk to a general-purpose sourcing engine.

The methodology was first proven with Kiosk—a vertical-market-software data platform built for investors. Notoriously tedious to identify and often misclassified by traditional data products, vertical-software companies had eluded all but the most resourceful investors.

Kiosk extracted more than 85,000 vertical-software companies from a universe of 13 million companies using proprietary classification. By 2025, the platform had expanded to cover more than 200,000 verified software companies and had become a core part of the sourcing process for dozens of international investors headquartered in seven countries.

The algorithms and methodologies that powered Kiosk became the foundation for a general-purpose sourcing engine. The same classification pipeline now processes more than 30 million companies and can be configured for any industry, any mandate, and any geography.

Kiosk narrowed a 13 million company source universe to more than 85,000 classified vertical-software companies. The proof later expanded to more than 200,000 verified software companies across 73 industries, with 24 average functional keywords per company and investor users headquartered in seven countries.
13MInitial source universe
Vertical-software definition
At launch85K+classified vertical-software companies
The classified set extended well beyond North America.
  1. North America42K
  2. Europe27.6K
  3. Latin America2.3K
  4. Oceania2.6K
  5. Rest of world14.9K

Your own sourcing engine.

The difference is not simply more records. It is control over the taxonomy, coverage, classification, supporting data, and confidentiality of the market itself.

The general sourcing engine processes more than 30 million companies through a market definition configured by industry, mandate, and geography to produce a mandate-specific company universe.
30M+Companies processed
Configured by industry, mandate and geography
A proprietary, prioritized company universe
Traditional data-driven sourcingYour custom engine
Market
The same data products as competitors.
A fully customizable industry taxonomy drawn from a universe of 30M+ companies.
Coverage
A focus on the easiest 50% of companies.
Coverage of whatever geography or vertical the mandate needs.
Control
Data sharing exposes proprietary findings.
Insight into—and the ability to customize—the proprietary classification algorithms.
Data quality
Low-quality leads and database pollution.
Fresh contact data and financial data where available.
Confidentiality
High competition for the same targets.
Yours alone: no other buyer sees the same universe.