Orca reads a LinkedIn profile through a paid data feed and lets the model decide how much more to scrape
AI agent for deep LinkedIn profile analysis.
At a glance
- What is it?
- DimiMikadze/orca is a Next.js agent that pulls one LinkedIn profile through a third party RapidAPI product, extracts named insights with LangChain, and opens extra scrapes on its own. Two billable keys and no default login sit behind the whole thing.
- Who is it for?
- Orca is a working prototype for reading one person's public professional history and turning it into fields a recruiter or investor can act on, and the agent that does the reading is the interesting part. Before running it, decide who is allowed to ask questions, because the app ships without a login while holding two billable keys, and remember that the model can widen the scraping on its own.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.
Editorial analysis
One profile URL and a list of insights go in, named findings come out
Orca takes two inputs: a LinkedIn profile URL and the insights you want extracted. The baseline sweep covers the profile itself, posts, comments, reactions, and engagement on the top posts. An agent then reasons over that material and returns named fields such as pain points, current focus, values, expertise, network influence, communication style, and how interests shift over time. Results stream into the interface while the agent is still working, so a partial answer is visible before the run ends. Three jobs are named for the output: reading a prospect's real priorities before outreach, looking past a candidate's resume during recruiting, and mapping a founder's thinking when evaluating a position. The insight list is yours to set, but the wording of each finding comes out of the model rather than a fixed schema checked into the repository, so two runs over the same profile can describe the same field differently.
The agent decides when to scrape more, and the operator does not
Step three of the documented loop is the part worth reading twice. If the agent decides that a particular insight needs more data, it calls scraping tools on its own. That autonomy is the feature, and it is also where the collection boundary leaves the operator's control. You choose which fields to extract. You do not choose which endpoints get hit to fill them. No per target limit, request budget, or switch that turns the extra calls off appears in the README or the package manifest. Contrast that with the baseline pass, which is a fixed list of five data types. The follow up traffic is model initiated, so a request for network influence can turn into extra post and reaction fetches the person running it never asked for. On a metered upstream feed those fetches are billable fetches. Data volume here is a function of how confident the model feels, not of a field anyone filled in.
Two paid keys share one .env.local, and LinkedIn is not the data source
Everything passes through a third party. The required keys are a RapidAPI key for Fresh LinkedIn Profile Data plus a key for whichever LLM provider you pick, OpenAI by default:
RAPIDAPI_KEY=your_key
OPENAI_API_KEY=your_keyOrca holds no LinkedIn credentials of its own, and that is worth stating plainly. What the repository fetches is whatever that RapidAPI product returns, carrying that provider's coverage, freshness, and rate limits with it. The price of one question is two metered calls in sequence, the upstream fetch and then the model run, and both are billed to keys sitting in a single file at the project root. The stack is Next.js 16, TypeScript, and Tailwind, with LangChain driving the model side. Node.js 20 and pnpm are the stated requirements, and the repo root carries pnpm-lock.yaml and pnpm-workspace.yaml, so pnpm is the documented install path rather than npm.
Installing is four commands, and the app comes up with no login
The install is a clone and a pnpm install:
git clone https://github.com/dimimikadze/orca.git
cd orca
pnpm install
pnpm devThen the app answers on http://localhost:3000. Nothing in that sequence starts a login, and the reason is stated outright. Authentication is optional. Locking the app down takes two further values:
NEXT_PUBLIC_SUPABASE_URL=your_url
NEXT_PUBLIC_SUPABASE_ANON_KEY=your_anon_keyOnly when both are set are all pages and the API protected behind an email and password login. Leave them out and the app runs open with no auth, on the same machine that already holds the two billable keys from the previous section. These are NEXT_PUBLIC variables, so they ship inside the browser bundle by design, which is the ordinary Supabase arrangement and not a fault in itself. What it means is that the decision to expose the app is a decision about keys and quota, not only about pages.
A types script in package.json points at one hardcoded Supabase project
One line in the manifest deserves its own look. The generate types script is:
supabase gen types typescript --project-id iwzpupxauyaukuljmjjg > ./app/supabase/database.types.tsA project id is baked into a committed script, so whoever runs it regenerates types against that project rather than their own, and the shell redirection writes whatever the CLI emits into app/supabase/database.types.ts regardless. That output path sits under the same app directory the optional auth values come from, which puts the generated schema and the login in one setup where neither is required in order to start. The README does not label the id as a placeholder, and it does not label it as production either. Three more root entries go unexplained: a proxy.ts, a CLAUDE.md that the README never refers to, and a demo.gif that sits next to a video link in the header.
Anthropic is named in the tech stack, but only the OpenAI integration is declared
The tech stack section says LangChain supports OpenAI, Anthropic, and other LLM providers, and the requirements call OpenAI the default. The dependency block tells a narrower story. It declares @langchain/openai at ^1.2.7 and langchain at ^1.2.21, and no Anthropic package appears among the dependencies. Whatever route exists for the other providers, it is not wired into this manifest as declared, and the README does not say which extra package or extra environment variable a non OpenAI setup needs. Next is pinned at 16.2.6 with React at 19.2.3, TypeScript throughout, and linting through eslint backed by an eslint.config.mjs at the root. The version field reads 0.1.0 and private is set to true, and the repository publishes no GitHub releases, so there is no tagged build to compare the manifest against.
Tests run on recorded fixtures until you edit a constant by hand
The claim is that all scrapers and the analysis agent are covered by tests. Two files at the root, vitest.config.mts and vitest.setup.mts, sit behind scripts that split the suite three ways: vitest run for everything, vitest run orca-ai/__tests__/services for the service layer, and vitest run orca-ai/__tests__/agents for the agent layer. Five narrower scripts name individual paths, among them the profile scraper, the activity scraper, post comments, post reactions, the LinkedIn data formatter, and the analysis agent. Every case runs against recorded fixtures by default with no live API needed, and switching a case over to real LinkedIn data means setting USE_LIVE_DATA = true inside the test file. That is a per file edit rather than a flag or an environment variable, so a run meant to hit live data is a change to the working tree that has to be noticed and put back. No coverage numbers appear anywhere in the repo.
Editorial conclusion
Orca is a working prototype for reading one person's public professional history and turning it into fields a recruiter or investor can act on, and the agent that does the reading is the interesting part. Before running it, decide who is allowed to ask questions, because the app ships without a login while holding two billable keys, and remember that the model can widen the scraping on its own. Treat the output as a draft to check against the profile itself, and find out what your upstream data provider permits before pointing this at anyone but yourself.
Frequently asked questions
What does DimiMikadze/orca need before it can analyze a profile?
A LinkedIn profile URL, the list of insights you want extracted, a RapidAPI key for Fresh LinkedIn Profile Data, and a key for your LLM provider, with OpenAI documented as the default. Node.js 20 and pnpm are required as well.
Does Orca put a login in front of the interface?
Not by default. Authentication is optional, and only adding both NEXT_PUBLIC_SUPABASE_URL and NEXT_PUBLIC_SUPABASE_ANON_KEY protects every page and the API behind an email and password login. Without both values the app runs open with no auth.
Does Orca fetch only the data an operator asked for?
No. The baseline pass covers the profile, posts, comments, reactions, and top post engagement. Past that the agent calls scraping tools on its own when it decides a specific insight needs more data, and no bound on those extra calls is described.
Can the Orca analysis code be reused outside the web app?
The core logic sits in orca-ai/ as a standalone library, and the repo root carries a pnpm workspace file, so it is meant to plug into a Node.js project. The root manifest reads 0.1.0 with private set to true and the repository has no GitHub releases, so no published package is indicated.
Do Orca's tests contact LinkedIn when the suite runs?
No. Each test case runs against recorded fixtures, and the README says no live API is needed for that. Setting USE_LIVE_DATA = true inside the test file is what points a case at real LinkedIn data instead.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/dimimikadze-orca)