Tech
How Kapa uses LLMs to help companies answer users’ technical questions reliably
Generative AI and large language models (LLMs) have been all the rage in recent years, upending traditional online search via the likes of ChatGPT while improving customer support, content generation, translation and more. Now, one fledgling startup is using LLMs to build AI assistants capable specifically of answering complex questions for developers, software end-users, and employees — it’s like ChatGPT, but for technical products.
Founded in February last year, Kapa.ai is a graduate of Y Combinator’s (YC) Summer 2023 program, and it has already amassed a fairly impressive roster of customers, including ChatGPT-maker OpenAI, Docker, Reddit, Monday.com and Mapbox. Not bad for an 18-month-old business.
“Our initial concept came after several friends who ran tech companies reached out with the same problem, and after we built the first prototype of Kapa.ai to address this for them, we landed our first paid pilot within a week,” CEO and co-founder Emil Sorensen told TechCrunch. “This led to organic growth through word-of-mouth — our customers became our biggest advocates.”
To build on that early traction, Kapa.ai has now raised $3.2 million in a seed round of funding led by Initialized Capital.
Getting technical
In the broadest terms, companies feed their technical documentation into Kapa.ai, which then serves up an interface using which developers and end-users can ask questions. Docker, for example, recently launched a new documentation assistant called Docker Docs AI, which provides instant responses to Docker-related questions from within its documentation pages — this is built using Kapa.ai.

But Kapa.ai can be used for myriad use-cases such as customer support, community engagement, and as a workplace assistant to help employees query their company’s knowledge base.
Under the hood, Kapa.ai is based on several LLMs from different providers and leans on a machine learning framework called Retrieval Augmented Generation (RAG), which enhances the performance of LLMs by enabling them to easily draw from relevant external data sources to provide richer responses.
“We’re model-agnostic — we work with multiple providers, including using our own models, in order to use the best-performing stack and retrieval techniques for each specific use case,” Sorensen said.
It’s worth noting that there are a number of similar tools out there already, including venture-backed startups such as Sana and Kore.ai, which are substantively about bringing conversational AI to enterprise knowledge bases. Kapa.ai, for its part, fits into that bucket, but the company says its main differentiator is that it largely focuses on external users rather than employees — and that has had a big influence on its design.
“When deploying an AI assistant externally to end-users, the level of scrutiny jumps ten-fold,” Sorensen said. “Accuracy is the only thing that matters, because companies are worried about AI misleading customers, and everyone has tried having ChatGPT or Claude hallucinate. A few bad answers and a company will immediately lose trust in your system. So that’s what we care about.”
Accuracy
This focus on providing accurate responses about technical documentation, with minimal hallucinations, highlights how Kapa.ai is a different kind of LLM animal — it is built for a much narrower use-case.
“Optimizing a system for accuracy naturally comes with trade-offs, as it means we have to design the system to be less creative than what other LLM systems can afford to be,” Sorensen said. “This is to guarantee the answers are only generated from the universe of content they provide.”
Then there is the thorny issue of data privacy — one of the major deterrents for enterprises that may want to adopt generative AI but are wary about exposing sensitive data to third-party systems. As such, Kapa.ai includes PII (personally identifiable information) data-detection and masking, which goes some way toward ensuring private information is neither stored nor shared.
This includes real-time PII scanning: When a message is received by Kapa.ai, it’s scanned for PII data, and if any personal data is detected, then the message is rejected and not stored. Users can also configure Kapa.ai so that any PII data detected in a document will be anonymized.
Businesses can, of course, assemble something akin to Kapa.ai themselves using third-party tools such as Azure’s OpenAI service or Deepset’s Haystack. But it’s a time-consuming and resource-intensive endeavor, especially when you can just tap Kapa’s website widget, deploy its bot for Slack or Zendesk, or use its API that allows companies to customize things a little with their own interfaces.
“Most of the people we work with don’t want to do all the engineering work, or don’t necessarily have the AI resources on their teams to do so,” Sorensen said. “They want an accurate and reliable AI engine that they can trust enough to expose directly to customers, and which has already been optimized for their use-case of answering technical product questions.”
In terms of pricing, Kapa.ai says it uses a SaaS subscription model, offering tiered pricing based on the complexity of the deployment and usage — though it doesn’t publish these prices.
The company has a remote team of nine spread across the globe in two main hubs in Copenhagen, where Sorensen is based, and San Francisco.
Aside from lead backer Initialized Capital, Kapa.ai’s seed round saw participation from Y Combinator and a slew of angel investors, including Docker founder Solomon Hykes, Stanford professor and AI researcher Douwe Kiela, and Replit founder Amjad Masad.
Tech
Substack’s new tool tells you who’s been writing their newsletters with AI
Substack has launched a new feature that can show you which of your favorite newsletters are being written using AI.
This week, the newsletter and writing platform announced an integration with the AI writing detection software Pangram that will allow users to scan posts, comments, and replies on Substack’s app to see an estimate of how much of the content was written by a human and how much was AI.
In the short term, the move might be bad for Substack’s business, as it could expose many of the newsletters on its platform that aren’t entirely written by people. That could potentially erode trust in the platform’s ecosystem of independent news and blogs, or even damage its reputation as a host of high-quality content.
But in the long term, AI-detection features could help keep Substack free of “AI slop” and encourage more users to trust what they’re reading was written by a person, or at least better understand when it’s not.

Substack joins several platforms that are leaning toward labeling AI content as such, especially now that AI is playing a greater role in the creation process. Photos and videos generated with AI are labeled on social media sites, while music streaming services have more recently begun labeling and, in some cases, penalizing AI-generated music.
“This is good use of AI,” Substack CEO Chris Best said.
“When I used to pitch Substack to writers, one way I would do it is … we’ll do everything for you except the hard part,” he explained in an online chat with Pangram’s founder, Max Spero. “You have to have something — an idea that’s worth reading, that’s worth caring about, that’s worth sharing. That one thing is very hard and very valuable … [S]oftware should do everything else, but I think you do want the person to do the hard part.”
The feature will be available in Substack’s app for any post, note, reply, or comment above 100 characters. Substack will also allow its writers to include an optional AI author’s note, using which creators can properly disclose their use of AI, the company told TechCrunch.
The company clarified that the tool is not meant to prohibit or penalize AI-assisted writing, but rather to encourage writers to add a “how I make this” statement, where they explain their process.
Publishers can also run Pangram on their own drafts before publication, and report and remove scans on their own work they believe are mistakes.
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.
Tech
Arcee, a US open source AI lab, says Chinese models are not inherently dangerous
As Chinese open-weight AI models grow in capability and popularity, arguments about what should be done about them have once again reached a fever pitch.
There’s talk that the Trump administration might try to ban them (though it hasn’t yet acted on the idea). Meanwhile, proprietary model makers, particularly OpenAI and Anthropic, appear increasingly concerned about them.
Open-weight models such as Moonshot AI’s Kimi K3 or Alibaba’s Qwen offer inference at a fraction of the token cost of closed source models from these large U.S. labs. The fear is that they also pose some sort of threat. Certainly they threaten the profit margins of the large proprietary AI labs.
But should enterprises running these models in their own data centers succumb to the fear that they could be a vector for Chinese hackers?
No, says Lucas Atkins, the CTO of Arcee, which is building open models to give U.S. companies a homegrown alternative to Chinese models.
If any startup would benefit from a ban on Chinese models, Arcee would. But Atkins says China’s open models are no more dangerous than any other open source software a company may use. In fact, he says, they even offer benefits even to his own company.
“A lot of people view this as similar to a Chinese software program. Like, it was coded with these x, y, z intentions” that a bad actor could simply command, he said.
“That is fundamentally not how these models are trained. There is really not any way for an Arcee, or an Alibaba, to make a model, have someone run it in their own environment and for us have any access to it whatsoever,” he explained.
While most of these models are what’s known as “open weight” and are not really fully open source software, the source code (the part that will actually run on servers), if it is downloaded from open source sites like Hugging Face, is similarly largely visible and reviewable. (What isn’t available is the methods and data used to train the models.)
Large organizations should put any model core through their security testing and inspection processes, and they will also often post-train the models for their specific uses and can examine areas like bias, toxicity, hallucinations, and sensitivity to certain topics. So they work with, optimize, and understand the models before people start sending them prompts.
Could a model that is used for coding somehow throw malicious backdoors into the code it writes? Again, while that’s theoretically possible, it would require acrobatic feats to accomplish.
“There’s no reason that a sophisticated enough actor couldn’t train a model to be a completely amazing coding model in every circumstance, but when presented with a certain type of code base … some hidden training would kick in,” Atkins, who spends his days training models, postulated. But he adds: “I don’t know how you would do this.”
Because large language models are by nature creative, the odds are slim of getting a contemporary model to spit out malware in response to a preplanned perfect storm of context and prompt. Even slimmer are the chances that any enterprise would then use that code.
Could it happen in the future? That’s anyone’s guess. But enterprises are also building their AI apps to be model-agnostic and to use multiple models. So even if Chinese models are the best for the price today, enterprises won’t be locked into using them forever.
“I think instead of the conversation being about how to ban Chinese models, it should be about how do we foster a good, open ecosystem here in the U.S.,” Atkins says.
Arcee also gains advantages from Chinese models. Because they are open, the startup “benefits from those models being good because we can learn what they did. We can build on top of them. Then they can learn what we do,” he says. “We have tremendous respect for the people building those models, the individual researchers.”
Ultimately, the way to compete with Chinese models “is to release a model that is better,” says Atkins. “We need to give them something to talk about.”
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.
Tech
Google is making it easier to switch from iPhone to Android
Google announced on Wednesday a new migration experience built directly into Android 17 that should ease the switch from iPhone to Android. The feature lets users wirelessly transfer more data types from an iPhone without needing to download a separate app, Google says.
By simplifying the onboarding process and supporting more data types, the tech giant is looking to lower the barriers to switching smartphone ecosystems as it aims to attract more iPhone users.
With this new method, users can transfer photos, videos, contacts, messages, calendars, and newly supported data types, including their Google Account, passwords, Wi-Fi credentials, and even their eSIM, when switching from an iPhone to Android.
The upgrade has already started rolling out to select Pixel devices, and it’s also available on the new Samsung Galaxy Z Flip8 and Z Fold8 series, which were unveiled today. The new migration method will also come to more Android devices soon, Google says.
