My Agentic System Outperforms Grok Bot
My agents vs Grok Bot, a better way to scrape the internet, & finding venture deals fast.

ROUND 5
Here’s edition #005 of Smoke Test.
This week was exciting because SpaceXAI, or whatever SpaceX + XAI is called nowadays, just released Grok Bot, an agent orchestration prototype built by the Cursor team.
Grok Bot is also the Cursor team’s first release post SpaceX acquisition, so needless to say, X was hyped, and I wanted to put it to the test.
Also on the docket today are a couple other tools a few friends passed our way to review. We are excited to check out, test, and review Exa, search and retrieval for agents, and Frontrun, a tool to help source early-stage venture deals.
Let’s get into it 🫡
PERSISTENT CLOUD AGENTS
Grok Bot: Grade B

USE CASES
Manage multiple workflows simultaneously across different agents
Teach agents repeatable workflows and rerun them as routines
Connect to an assortment of tools and custom plugins.
DESCRIPTION
Like I said above, Grok Bot is SpaceXAI’s crack at an agent harness… A game they’ve been losing for quite some time to teams like Open Claw, Codex, OpenCode, Perplexity, and Hermes.
Grok Bot has all of the typical features you’d expect from platforms like Codex & Claude Cowork. The main differences are, well, its powered by Grok models of course, but also it has persistent multi-agent orchestration and the agents use declarative task delegation.
What do I mean by that?:
Persistent multi-agent orchestration: agents that coordinate and hand off work between themselves, 24/7, without you in the loop
Declarative task delegation: you describe the outcome and the agent determines the steps.
These are the two pillars that my personal agentic operating system Brodus is built on and is the exact reason why I chose to build Brodus on Hermes.
So after seeing so much hype on X, I had to test and see if it could execute better than Brodus…
REVIEW
So I put Grok Bot through six real workflows to see if it could actually outperform Brodus.
The first three focused on research and discovery such as finding company and founder signals, mapping Austin AI events, and researching the agent-memory market. The results were solid but inconsistent.
Grok was fast and surfaced information I probably would have missed through normal web search. But it repeatedly failed to link and source underlying X posts I asked for. That was especially surprising because X data accessibilty is supposed to be one of Grok’s biggest advantages and practically the only reason I would ever use it… No bueno.
I then proceeded to gave Grok Bot three structured assignments pulled from real tickets within the Brodus system. This is where it got more interesting.
With only a few paragraphs of context, Grok produced a high quality product evaluation, sales offer portfolio, and client meeting pack. It caught a real diligence error, respected commercial constraints, built realistic test plans, and produced work that was close to decision-ready…. without any access to my corpus of data.
So, did it beat Brodus? Of course not.
What honestly caught me off-guard was Grok’s struggle with open-ended research and the lack of X data. Huge disappointment considering that’s what its strongest value proposition was when Grok first launched.
However, when I gave it the same kind of task contract Brodus gives its agents: a specific outcome, clear constraints, acceptance criteria, and defined human gates, the results improved drastically.
Despite not having access to my data infrastructure and tools, overall I’d say Grok did well, and I’ll give it a solid B. Did it earn the hype? Eh, I’m not sure.
But what I’ll definitely say is I still think its 100% worth it to build your own system!
> Brody
VERDICT: B
Good platform to become acquainted with agent harnesses but still worth it to build your own system.
AGENTIC SEARCH & RETRIEVAL
Exa: Grade C

USE CASES
Find recent, niche sources that general web search buries
Compare products or technical approaches across primary sources
Generate a quick cited answer before deeper verification
DESCRIPTION
A few of my friends recommended I try Exa so I thought sure, why not. Exa is a search and retrieval API built for AI agents. It combines search, page contents, grounded answers, deeper multi-step search, and an asynchronous research agent.
Official pricing is usage-based. standard search is $7 per 1,000 requests, Answer is $5 per 1,000 requests, and Contents is $1 per 1,000 pages. New accounts currently receive free credits.
REVIEW
I tested Exa’s Search, Answer, and Contents paths against research work we already perform with general web search and Hermes retrieval.
Search was the clear winner. It improved discovery and comparison depth on the AI-tool queries we ran, especially when the job was finding less obvious sources rather than summarizing the first page of results. This is exactly what I needed it for! I felt like the previous, default, agent-native search tool calling was missing a lot of depth. Exa fixes that.
I also tested Exa’s Answer endpoint by asking it to research Hark Handoff, an AI agent that operates a virtual computer to complete online tasks such as shopping. Exa returned a useful overview of the product, pricing, limitations, and ideal users with eight citations, but left out the benchmarks and competitive evidence we needed for a final evaluation.
Contents was the weakest differentiator. It successfully extracted a technical article with its code blocks and formatting intact, but our existing Hermes tool returned essentially the same content. It worked well; it just didn’t improve a job we already handle without an additional usage charge.
All-in-all, I’m not sure if the pony is worth the pennies here unless you’re specifically looking for a deeper search on your queries. These tests didn’t get me excited enough to pay for an entirely new subscription just to access the internet better with my agents.
> Brody
VERDICT: C
Use Exa when discovery depth matters but keep your existing extractor for ordinary pages.
EARLY-STAGE DEAL INTELLIGENCE
Frontrun: Grade B-

USE CASES
Find emerging AI tools through investor follow signals early
Run semantic thesis searches across the VC follow graph
Build a weekly, sector-filtered candidate pipeline for this newsletter
Run fast company, funding, founder, and activity pre-research
DESCRIPTION
Front Run turns X follow graphs into structured investor and company signals.
I’ve been following Frontrun and testing their platform for awhile now and I can say that it’s drastically improved with the addition of the API and the MCP. Its API and MCP expose new follows, convergence, trending entities, company research, classification, rules, and alerts.
Pro is $99 per month with API and MCP access plus 10,000 monthly credits.
REVIEW
I ran six tests to see if Frontrun could actually find Smoke Test candidates before they hit the same launch channels everyone else watches.
The strongest result came from its enriched-follows endpoint. Frontrun processed 3,076 recent follows from 50 tracked investor and operator accounts, then classified 339 of the accounts being followed as companies.
That is a lot of noise, but once my agents filtered it down, the signal was useful. Instead of searching for companies that had already launched, we could see which companies were quietly beginning to attract attention from people inside the market.
There’s so many other useful ways to use these signals, including deal sourcing and client prospecting, which help you Front Run (get it) other investors or service providers that are simply relying on trigger events, such as a fundraise, to source deals.
I then tested its thesis search by asking for companies building AI systems for business operations and workflow automation. It returned 15 relevant companies, including several we probably would not have found through a normal keyword search. This workflow, combined with my competitor radar workflow will help me track our competition and keep ATX Logic sharp and ahead of the curve when it comes to AI implementation services.
But then the feed got weird…
A convergence search returned 529 signals, but the classification failed and gave us investors and public figures instead of companies. A search for smaller AI companies with fewer than 10,000 followers returned only seven results, which exposed the obvious limitation: if the tracked VC network hasn’t noticed a company, Frontrun probably hasn’t either.
So, did Frontrun replace our existing sourcing process? No. But it found a useful place inside it.
We’d use Frontrun as an early-warning system for companies gaining attention before they reach the usual launch channels. Then we’d use our existing research stack to figure out whether the company is actually interesting or just having a good week on VC Twitter.
> Brody
VERDICT: B-
Helps me stay current with venture trends, keep updated on new tools, and connect with prospects early.
YOUR TURN
Where do you get an information advantage before everyone else catches up?
Hit reply and send us:
The search, agent, or signal workflow you trust
The decision it helps you make
What you want us to test, compare, or break
We’ll test the strongest ideas and report what actually holds up.
> Brody & Jake
ATX LOGIC. AUSTIN, TX
CURIOUS ABOUT MORE TOOLS?
In edition #004, we rated Monid, a tool that’s trying to become the one tool that rules them all and Neon, a solid choice for anyone looking to easily deploy a storage solution for their agents.
We also reviewed Paper, a design tool that I originally believed did not have enough “free” in their Freemium… But I messaged their team and they gave me an additional free month to continue trying it, which definitely earned them some extra brownie points. 🙂

Powered by beehiiv