Dossier · Private startup · 3 independent sources

definity

Cloud & Developer Infrastructure Priority Signal Founded 2023

Last updated: Aug 31, 2026

definity is an Israeli-founded data-infrastructure company building an agentic runtime layer for operating Apache Spark and lakehouse data platforms. Its platform combines in-motion observability, data-quality and execution context, cost optimization, and guarded AI-driven remediation so enterprise data teams can prevent pipeline incidents instead of investigating them after the fact.

Visit Website

Company Overview

**Product and the concrete problem it solves.** definity addresses the operational gap between modern data platforms and the teams expected to keep them reliable. Enterprise analytics and AI workloads increasingly depend on Apache Spark jobs, lakehouse tables, and cloud data services, yet the operating experience for data engineering remains fragmented and reactive. When a pipeline fails, slows down, produces questionable data, or consumes more compute than expected, engineers commonly correlate logs, dashboards, job history, configuration, and data behavior manually. The result is delayed detection, missed service-level agreements, wasted engineering capacity, and cloud bills that rise before anyone understands why. definity's initial product was a data-application observability and remediation platform purpose-built for Spark-heavy environments; its April 2026 launch reframed that product as Agentic Data Engineering, a runtime layer that can understand and operate pipelines while they execute. The company is not merely presenting another dashboard. Its stated aim is to combine visibility, diagnosis, optimization, validation, and controlled action in one operating surface for production data systems.

**Core technology and how it works.** The technical distinction is the company's in-motion, application-focused architecture. definity says lightweight agents can be embedded directly in Spark workloads without code changes, in on-premises, hybrid, or cloud environments. Those agents collect unified context across pipeline execution, transformation behavior, data characteristics, configuration, lineage, and infrastructure performance rather than examining only the final output or one isolated service metric. That context supports anomaly detection for data quality, job execution, and performance while a job is running; transformation-level root-cause analysis; workload profiling; and job-level cost and resource recommendations. The platform also exposes code-change validation in CI, allowing a proposed change to be tested on real data before it creates runaway spend, failures, data-integrity problems, or a missed SLA. In the newer agentic model, AI systems can act on this runtime context instead of being limited to a post hoc explanation. The important diligence question is how much of the action loop is autonomous in production and what approval, rollback, and policy guardrails customers can configure. Public materials establish the architecture and product direction, but do not disclose model choices, proprietary training data, intervention accuracy, or the full mechanics of safe remediation.

**Market, customers, and go-to-market.** definity sells into data engineering, platform engineering, and data-platform leadership teams at organizations operating large Spark or lakehouse estates. Its supported ecosystem includes Databricks, AWS EMR, and Google Cloud Dataproc, which gives the company a wedge into existing infrastructure rather than requiring a customer to replace its data stack. The commercial motion appears to be enterprise software sold through direct technical evaluation, production pilots, and ecosystem relationships. A disclosed partnership with Databricks positions definity alongside the Databricks Data Intelligence Platform, where the company adds full-stack observability, workload optimization, proactive validation, and transformation-aware execution context. definity also publishes customer stories, including an anonymized top-three U.S. television and streaming provider using AWS EMR and Spark Streaming. That deployment illustrates the buyer's pain: time-sensitive pipelines fed customer-facing recommendations and advertiser analytics, so reliability and runtime were business issues rather than internal developer conveniences. Public sources do not identify the customer, contract size, pricing, renewal rate, or total customer count. That disclosure limit matters because a technically credible platform still has to convert infrastructure complexity into repeatable enterprise adoption.

**Traction, funding, and third-party validation.** definity emerged from stealth in August 2024 with general availability for its Spark-first data-application observability and remediation product and a $4.5 million seed round led by StageOne Ventures, with Hyde Park Venture Partners and individual founders participating. On April 29, 2026, the company announced a $12 million Series A led by GreatPoint Ventures, joined by Dynatrace, StageOne, and Hyde Park, bringing reported total funding to $16.5 million. CTech independently reported the round, named the investors, and described the platform's focus on Apache Spark and lakehouse architectures. Geektime's Israeli coverage adds a useful technical description of the in-motion architecture and reports the company's claims that AI agents can reduce platform costs by more than 30% and resolve complex problems ten times faster; those are company-reported performance claims, not independent benchmarks. The strongest public customer evidence is definity's case study for the top-three TV provider, which reports a 74% reduction in heavy-pipeline runtime, a critical pipeline shrinking from 17 hours to 4.5 hours, and 58% lower platform resource consumption. The figures are material but come from the vendor's own case study and an unnamed customer, so customer references and reproducible measurement methodology remain important diligence requests.

**Founders and team background.** The founding team is Roy Daniel, CEO; Ohad Raviv, CTO; and Tom Bar-Yacov, VP of R&D. The official company profile describes them as enterprise data, platform, and product leaders, while the company's original launch post grounds the problem in their prior experience leading data engineering and product work at PayPal and FIS/Worldpay. That background is relevant because definity is selling an operating model to teams that own production infrastructure, not a research demo to model enthusiasts. Daniel is the public voice of the company in the Series A announcement and Israeli funding coverage, while Raviv and Bar-Yacov are identified as the technical and R&D counterparts. Beyond these roles and the enterprise provenance, the public record does not provide enough detail to verify specific degrees, military units, patents, headcount composition, or prior exits; those facts should not be inferred from the founders' seniority. Startup Nation Finder and Dealroom classify the business as founded in 2023 with an estimated 11-50 employees, but those are database estimates rather than a company-disclosed census. The team signal is therefore solid domain familiarity with a deliberately modest confidence level on scale and technical depth.

**Competitive dynamics.** definity sits in a crowded but still unsettled layer of the data stack. Datadog, Dynatrace, and New Relic bring broad infrastructure and application observability, but a general-purpose APM view may not capture Spark transformation semantics, data lineage, or the relationship between data quality and compute behavior. Monte Carlo, Bigeye, and Soda compete around data observability and quality monitoring, while Unravel Data and Acceldata address performance, cost, and operations for complex data platforms. Databricks itself is both an ecosystem partner and a strategic threat because it controls an important data-platform surface and can add more native monitoring or optimization over time. definity's claimed edge is the combination of application execution, data behavior, and infrastructure performance in a single real-time context, plus action and CI validation rather than alerting alone. Its 2025 Databricks partnership can reduce distribution friction, but it also increases platform dependency. The defensibility test is not whether the company can produce recommendations; it is whether its runtime instrumentation, transformation-level context, closed-loop remediation history, and customer-specific baselines improve outcomes enough to survive bundled features and incumbent contracts.

**Defense, security, and resilience dual-use relevance.** definity is strategically relevant to resilience and AI infrastructure, but the public record supports adjacency rather than a fielded defense capability. Data pipelines increasingly supply AI systems, intelligence workflows, logistics planning, fraud detection, industrial control analytics, and emergency-response decisions. A platform that detects data degradation, runaway compute, and pipeline failure before downstream consumers are affected could improve continuity for a critical-infrastructure operator or defense contractor. Its support for cloud and on-premises environments also maps to organizations that cannot place all mission data in one public-cloud operating model. The same technology could eventually help a security organization maintain trustworthy data products across disconnected or intermittently connected environments, provided the product gains suitable deployment controls and assurance evidence. However, definity publicly markets to enterprise data teams and names no military, intelligence, government, or critical-infrastructure customer. There is no public evidence of defense accreditation, classified-environment deployment, secure-edge packaging, disconnected operation, or adversarial testing. I therefore set dual_use to false: the core product is a credible commercial reliability layer with useful resilience potential, but the defense connection is not yet demonstrated strongly enough to call it true dual use under this database's standard.

**Growth stage, trajectory, and key diligence risks.** definity is best classified as early-to-mid transition: it has general availability, a named technology partnership, a funded Series A, and public production evidence, but it is still building category recognition and has not disclosed the commercial metrics normally associated with a mature infrastructure vendor. The trajectory is attractive if agentic operation turns data engineering from a monitoring cost center into a measurable control plane for AI readiness. The main risks are: (1) platform concentration, because Databricks, AWS, and Google can absorb adjacent functionality; (2) crowded observability competition from vendors already embedded in enterprise budgets; (3) autonomy and trust risk, because an incorrect automated intervention can corrupt data or interrupt a revenue-critical job; (4) evidence risk, because the best quantitative outcomes are vendor-reported and customer-anonymous; (5) services and deployment friction across Spark versions, cloud providers, and on-premises estates; (6) model and telemetry economics as AI agents inspect high-volume pipelines; and (7) Israeli-company identity and operating-location ambiguity, since public ecosystem sources identify an Israeli founding location while third-party databases list Chicago as headquarters. The next proof points should be named reference customers, retention and expansion data, independent validation of savings and incident prevention, disclosed guardrail behavior, stronger edge or sovereign deployment evidence, and sustained growth without becoming a feature inside a larger platform.

Strategic Fit Assessment

Research priority signal

Priority signal means this entry may be worth researching within the Claw & Talon thesis. It does not mean investable, suitable, endorsed, available, or likely to produce returns.

definity is a credible infrastructure-company priority signal, not an investment recommendation. (1) The product targets a real operating bottleneck: Spark and lakehouse systems are central to AI and analytics, but data teams still lack a unified control layer spanning data behavior, execution, and infrastructure cost. (2) The company has progressed from general availability to a $12M Series A in under two years, with GreatPoint Ventures, Dynatrace, StageOne Ventures, and Hyde Park Venture Partners validating the category and team. (3) Its Databricks relationship and customer case study create stronger evidence than a purely conceptual agent platform, including reported runtime and resource-consumption improvements. (4) The strongest upside would come from becoming the runtime control plane for agentic data engineering as AI multiplies the number and importance of production pipelines. The counterweights are significant: vendor-reported and anonymized traction, undisclosed revenue and retention, intense competition from Databricks and observability incumbents, operational risk from autonomous changes, and a still-unproven defense or sovereign deployment path. Priority should rise only with named references, independently verified savings, durable expansion metrics, and evidence that guardrails make autonomous operation safe in high-consequence environments.

Strategic Value to U.S.-Israel Alliance

definity's strategic value is concentrated in AI-infrastructure reliability rather than direct defense technology. (1) Trustworthy data is a prerequisite for trustworthy AI, and a runtime layer that can catch data, execution, and cost failures before downstream systems are affected addresses a foundational bottleneck. (2) The product's contextual approach may be more useful than disconnected dashboards because it links pipeline logic to infrastructure behavior and data quality. (3) Hybrid and on-premises support preserves optionality for regulated, sovereign, and critical-infrastructure operators that cannot centralize all workloads in one cloud. (4) The Databricks partnership can make definity an ecosystem component while giving it access to a large installed base. The ceiling is constrained by the absence of public-sector deployments, certifications, classified-environment evidence, or named strategic customers; for now, definity is best understood as a potentially important commercial reliability layer with resilience spillover.

Key Technologies

  • In-motion instrumentation agents embedded in Apache Spark workloads without application code changes
  • Unified runtime context joining pipeline execution, transformation behavior, data quality, lineage, configuration, and infrastructure performance
  • AI-agent-assisted anomaly detection, root-cause analysis, and guarded remediation during production execution
  • Job-level Spark and lakehouse workload profiling, resource allocation analysis, and cost optimization
  • Real-data CI validation for predicting pipeline failures, runaway compute, SLA misses, and data-integrity regressions
  • Cross-environment support for Databricks, AWS EMR, Google Cloud Dataproc, hybrid, and on-premises deployments

Use Cases & Applications

  • Preventing Spark pipeline failures before customer-facing analytics or AI feature data is delivered
  • Reducing AWS EMR or Databricks compute waste through job-level and task-level workload optimization
  • Validating data-platform code changes against real workloads before production rollout
  • Diagnosing transformation-level data-quality and performance regressions across lakehouse pipelines
  • Maintaining SLA reliability for streaming analytics, recommendations, advertising, and operational reporting
  • Supporting hybrid or on-premises data estates that cannot rely only on public-cloud monitoring
  • Improving continuity of data products used by regulated enterprises, logistics operators, and critical infrastructure
  • Prospective monitoring and controlled remediation of data pipelines supporting defense or public-sector AI systems; adjacency only

Sources and verification

This profile is based on public-source research, Claw & Talon curation, and editorial judgment. Inclusion does not imply endorsement, partnership, investment, or a recommendation to transact. Readers should still confirm current status, customers, funding, and product claims before relying on this profile. The editorial policy explains how profiles are researched, where automated drafting is used, and how corrections work; the research methodology documents how evidence is graded, what counts as an independent source, and why some profiles are excluded from search indexing.

This record lists 8 public references used for company identity, status, positioning, or material-claim review.

Public sources

The links below are visible public references used for source discipline around company identity, status, funding, customer, acquisition, public-company, or other material claims where available.

Related sector

See the Cloud & Developer Infrastructure sector page for market context, related subcategories, and other Israeli companies in this part of the database.