Appearance
Two bugs, neither one the reason I lost
I built StacksNG for the Africa Deep Tech Challenge 2026: an offline coding assistant grounded in the docs of four Nigerian fintech APIs, Paystack, Flutterwave, Monnify, Termii. Retrieval, not memory. The pipeline pulls the real doc passage before answering and cites the source, so it can't hallucinate an endpoint in payments code.
The corpus was 780 scraped chunks. That's small enough to embed in a minute on a laptop and query offline in milliseconds, which is the whole trick of on-device RAG. The model was stock Qwen2.5-Coder-7B, quantized to Q4_K_M. That compression brings it down to roughly 4.4 GB, the difference between running on a consumer GPU and renting a cloud box. No cloud, no API keys, all through llama.cpp.
I didn't make the semifinals. No feedback came with the rejection, so I went looking.
The decision to ship the 7B was deliberate and documented. A 1.5B version of the same pipeline beats the published scoring formula by about 35 points: faster, lighter, better on the speed and memory-efficiency components. I rejected it because on one of my two registered test prompts, the 1.5B opened with:
python
import flutterwave
client = flutterwave.Client(...)That library doesn't exist. The 7B, with the same retrieved context, wrote real request calls against real documented endpoints. I picked correctness over the formula and said so in the report. That turned out to be the wrong thing to defend.
The empty accuracy array
The first useful find was in my own submission.json:
"accuracy": []
That's not a low number. No number. git log on the file shows one commit, never touched before the deadline. The profiling run either skipped the accuracy gate or lm_eval wasn't installed when it executed. Either way the failure was silent, not an error.
Sacc is 50% of the scoring formula. Half my score shipped as an empty array.
Trying to fix it, I installed lm-eval-harness, pointed it at my model through a local llama-server, and got a crash on all 200 requests: "Invalid logprobs data." Both bugs checked out against the actual source:
- adtc_profiler's accuracy.py calls lm_eval with
base_url="local". Not a URL. The GGUF backend needshttp://host:portand posts to{base_url}/v1/completions. Anyone running it literally hits the same failure. - lm-eval-harness's GGUF backend expects
echo=trueto return logprobs for the whole echoed prompt, the old OpenAI completions behavior. Current llama-server only returns logprobs for newly generated tokens. The harness hasn't caught up.
I patched both locally. Instead of trusting the broken echo behavior, I forced the model to generate the exact answer text with a grammar rule and read its confidence off that generation. Running the challenge's own default benchmark, a general-knowledge multiple-choice eval, the stock 7B scored 0.74 on arc_easy acc_norm. Competitive. Ahead of one fine-tuned semifinalist, just behind two others.
Plugged into the scoring formula, the total goes from 10.23 to 47.23. A 4.6x swing from one unmeasured number. No architecture change, no training, no cleverness.
Reading the comments on my post sharpened the whole thing for me. One reader put a rule on it that I now use when designing anything with an eval in it: a judging system should expose three states, never one score. Not measured because the evaluator failed. Measured and failed a metric. Rejected by an earlier eligibility gate. My debugging week failed because I couldn't tell which state I was in. The empty array said "not measured," but the system presented it as a score, so I went and fixed the measurement. When I did, I found out the rejection had nothing to do with the score at all.
Quick take: a judging system can have a gate that never shows up in the rubric or the tooling, and if you don't assume it exists, you'll debug the wrong layer for the whole competition.
The gate that wasn't in the rubric
The rejection email held the actual reason.
Originality Score (0-10): 3. Model Origin: Stock model, used as-is: Qwen2.5-Coder-7B-Instruct-GGUF from lmstudio-community's official quantization.
Round 1 runs an originality gate before any technical scoring happens. Template compliance: fine. The RAG architecture, the citation grounding, the 780-chunk corpus: not mentioned. The problem was never Sacc. The problem was that I never touched the model.
The evidence was sitting in my own clones. I pulled all 20 published semifinalists with full git history and read every technical report. SME-Ledger, Jamii Afya, TaxSabi, CodeFellow, ARIS, Homa: LoRA, QLoRA, distillation, or a merge. Every single one touched the base model. Mine was the only stock-model architecture in the batch, and I filed that under "interesting" instead of "alarming." Eighteen other teams fine-tuned. I read that fact and reasoned right past it.
That's the uncomfortable part. The data was in front of me before the rejection email arrived. I was looking at the parts of the system I could inspect: the formula, the profiler source, the tooling. The part that mattered was documented nowhere I had access to.
A RAG layer grounding a stock model is a defensible product decision. To an originality reviewer, it's not evidence that you built anything. Those are different bars, and clearing one says nothing about the other.
Key numbers from the postmortem
- submitted total 10.23; recomputed total 47.23, a 4.6x swing from one unmeasured number
- real accuracy 0.74 on arc_easy acc_norm, competitive with fine-tuned semifinalists
- Sacc = 50% of the scoring formula, shipped as
"accuracy": []
What this says about enterprise agents
The agent industry is replaying this exact story at production scale. Perplexity now trusts GPT-6 Astra with end-to-end systems: writing communications, changing software, monitoring production, and checking in much less frequently than it did with earlier models. Fyxer built an executive assistant that organizes inboxes and drafts emails in each user's voice using fine-tuning, memory, and real user feedback. OpenAI is shipping ChatGPT for Financial Services, pairing built-in financial data with GPT-6 Astra for research, modeling, and client-ready materials, in a vertical where one hallucinated number is a compliance incident.
Those are trust problems before they're model problems. Every one needs an answer to who verifies what actually happened, and what happens when verification fails silently. My hackathon entry died on exactly that question.
The projects that survive treat the evidence layer as the product, not a side effect. oh-my-hermes, an operating layer on top of Hermes Agent, is the clearest example I've seen recently. Every request is scored before dispatch and every signal that moved the score is named. Execution lanes return typed results with four states: process exited, schema valid, verification observed, integration ready. An exit code of zero without evidence stays "reported done" until a gate checks it. A verification receipt is reused only when the revision, command, and environment all match. Costs render as "unknown," never $0, unless the host confirmed the price.
That design language is a direct answer to my empty accuracy array.
| System | What it automates | How it earns trust | The gate it still must clear |
|---|---|---|---|
| Perplexity + Astra | Communications, code changes, production monitoring | Model owns the loop end-to-end, humans check in less | Auditing what changed in production |
| Fyxer | Inbox triage, email drafting | Fine-tuned on the user's writing, remembers corrections | Catching a drafting mistake before send |
| ChatGPT for Financial Services | Research, modeling, client-ready materials | Grounded in built-in financial data | Compliance review for client-facing output |
| oh-my-hermes | Planning, routing, coding lanes | Verification receipts tied to revision, command, environment | Matching evidence to the actual execution context |
| StacksNG | Offline Q&A over four fintech API docs | Retrieved source citations on every answer | An originality gate that ignores the application layer |
Evidence boundaries in the operating layer
The numbers behind oh-my-hermes routing are the part I'd steal first. On the same coding tasks with the same underlying model, routing by work category instead of one prompt chain produced the same answers for $0.66 instead of $4.29, in 5 minutes instead of 23. Most of that gain is mundane: short tasks land on a cheap fast model, deep reasoning goes to an expensive one, and a rejected provider falls back along an explicit chain instead of silently downgrading.
The governance details matter just as much as the savings. Model chains are plain editable JSON, seeded by setup and reordered to the providers you've actually linked. Nothing is remembered silently: a candidate memory goes on a review card, gets approved or refused with the reason written down, and ages from active to reference to archive on silence. The per-lane HUD shows model, effort, turn, tokens, and cost, updated live, and an unpriced call says "unknown," not $0.
That's the gap my hackathon entry fell into. I had the retrieval and the citations, but I never built a layer that could prove what actually happened, because I didn't know the judging process would require one. In production, the judge is a compliance reviewer, a downstream team, or the next shift of engineers. Same requirement.
Common pitfalls
Optimizing the formula you can see
The 1.5B variant wins the published formula by 35 points and hallucinates a nonexistent API library on a real prompt. The published formula also never mentioned originality, which is where the entry actually lost. Read the rubric, then map what it leaves out. If you're in a judged setting, modify the base model. A light LoRA pass on your own corpus clears an originality bar that application work can never reach.
Shipping an eval you never saw complete
"accuracy": [] is a silent non-result, committed once and never re-checked. A failed eval and a low eval are different states, and the tooling should say which one happened. Before you trust a score, verify the artifact that recorded it. Empty output file means the eval didn't run.
Assuming the eval harness agrees with your server
base_url="local" is not a URL, and the harness still expects echo=true logprobs for the entire prompt while llama-server only returns logprobs for generated tokens. If you see "Invalid logprobs data," don't start blaming the model. Check the contract between the harness and the inference server first.
Building the application layer when the gate scores the model
RAG, citations, and a curated 780-chunk corpus counted for nothing at the originality gate. If the judge's first question is "did you touch the model?" and you didn't, everything below that line is invisible. Know which artifact is being scored before you decide where the work should go.
One thing to remember
The architecture that survives is the one that makes its own evidence visible. A system judged without a decision trace can't tell you which failure killed you, and a system that reports "done" without verification is a liability either way. The agent products I trust most treat the evidence layer as the core feature: what happened, what verified it, and what never ran.
The bottom line
- If you're entering a hackathon, benchmark, or procurement where innovation is scored, assume a gate exists that isn't in the rubric and modify the base model. A light LoRA pass on your own data clears an originality bar that a RAG layer can't.
- If you're shipping an agent into production, build evidence gates, not just test suites. Exit 0 without verification should read "reported done," cost should read "unknown" not $0, and every verification receipt needs its exact execution context.
- One thing to watch: eval tooling is still lagging the servers it talks to. The logprobs and base_url mismatches are version-skew failures, and they will keep producing silent zeros until harnesses standardize on server-native contracts. Expect that within two release cycles, and treat any eval you haven't seen run as unmeasured.