An AI assistant can have the right document, a long context window, and a detailed record of its previous work—and still produce the wrong command.

That gap runs through this week’s papers. Access to information is not the same as competence. Keeping everything is not the same as remembering what matters. Finding something similar is not the same as finding something useful.

The most interesting work here asks what happens between acquiring information and acting on it: how an agent organizes a recurring source, maintains working memory, retrieves a genuinely useful connection, or translates documentation into an executable instruction. Other papers attack the same problem at a different level, building compact representations that preserve motion, geometry, or the ability to revise a draft.

These are ten papers from late September and early October, ranked by their significance for practical technical work and assistant design—not simply by the size of their headline gains. Most are preprints. Their results are worth examining, but the benchmarks are not interchangeable and the numbers do not add up to a single story of inevitable progress.

1. SourceLearn: learning a source, not collecting answers

From Knowledge Access to Source Learning: Developing Source-Specific Competence makes the strongest conceptual distinction of the batch: repeatedly retrieving a source is different from becoming better at using it.

An assistant working with the same repository or documentation set should gradually learn its structure. Which procedures have prerequisites? Where do the exceptions live? Which concepts must be understood together? A pile of retrieved passages can answer an immediate question without building any of that reusable competence.

SourceLearn maintains a revisable “source model” alongside the original material. One learning process studies gaps in that model; another uses task experience to identify missing knowledge or poor organization. Crucially, task experience is supposed to direct what to revisit, while the source supplies the content of persistent updates.

Across five benchmarks and three model backends, the authors report the best result in 13 of 15 settings. Average improvements over Hybrid RAG were 14.3, 4.9, and 13.4 points for the three respective backends. That is substantial, although it applies to repeated work on the same sources under the paper’s controlled guidance/test splits.

The boundary matters. These were relatively stable, authoritative sources. Evolving or conflicting material is left for future work, and the grounding audit still found some partially supported or unsupported knowledge units.

For Jarvis, this is more useful than the vague instruction to “remember more.” A lesson about navigating a repository should remain distinguishable from a factual claim about its current code. Experience can identify a recurring blind spot; it should not silently become authority. The source model is a map, not a replacement for the terrain.

2. KaliBench: knowing the tool is not knowing the command

KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards turns a familiar assistant failure into a measurable distinction.

A model may correctly choose a tool and then invent a flag. Its explanation can be impeccable while its command is unusable.

The benchmark contains 8,504 query–command pairs across 1,642 tools. It scores tool selection separately from optional arguments, positional arguments, and exact command correctness. Its three modes supply progressively more help: just the request, a shortlist of tool names, then documentation for those tools.

The results are unusually actionable. Average tool accuracy rose from 72.0% to 95.2% when models received the shortlist, but argument construction improved only modestly. With documentation, average exact-command correctness rose from 22.3% in the unrestricted setting to 73.1% in the hinted setting.

That is evidence for a specific intervention: retrieve the manual, not merely the tool’s name.

It is also a warning about evaluation. These are documentation-derived, single-command tasks, often using synthetic targets or inputs. A structurally correct command is not proof of a successful security operation, let alone safe autonomous behavior. “Runtime-free” refers to the training reward; dataset construction did include sandbox execution.

The KaliBench repository is relevant beyond cybersecurity as an evaluation template. For Jarvis, choosing a plausible shell utility should not count as success. The flags, values, ordering, installed version, and eventual outcome deserve separate checks. Fluent prose is not a command validator.

3. VideoLoop: a working memory should be allowed to forget

VideoLoop: Looped Working Memory Against Semantic Thrashing in Long-Form Video Agents tackles the failure mode of endlessly appending observations.

In a long investigation, an agent accumulates facts, guesses, abandoned leads, and intermediate outputs. More context can make the important evidence harder to find rather than easier.

VideoLoop separates video exploration from memory maintenance. An outer agent investigates and saves observations and artifacts to a persistent filesystem. An inner orchestrator retrieves relevant material and rewrites a bounded working memory after each step.

With Gemini 3.1 Pro, the authors report improvements of 4.5, 4.2, and 3.2 percentage points over the corresponding native model on three long-video benchmarks. A separate diagnostic asks a blind judge to answer using only the stored context. On the hardest quarter of VideoMME long questions, that judge scored 81.1% with VideoLoop’s context versus 60.9% with append-only context.

The latter is not the video agent’s final accuracy. It measures whether the context contains enough usable information for another model to answer.

That distinction makes the diagnostic interesting: it evaluates the memory as an artifact, rather than treating memory quality as whatever the final solver happened to achieve.

For Jarvis, the relevant pattern is a searchable evidence record plus a concise, revisable task state. But rewriting introduces its own risk. A tidy summary can omit the one exception that matters. Keeping the external record is what makes compression recoverable.

4. ScholarCatalyst: relevance is not inspiration

ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research asks a harder question than ordinary literature search: which prior ideas could actually help this project advance?

Its 894 research questions from 207 projects use judgments from 184 lead authors. The queries reconstruct early-stage research problems without revealing the eventual solution. Authors assess both useful papers and topically related candidates that do not help.

In the headline comparison, the strongest embedding retriever achieved 0.48 Recall@20, versus 0.42 for an agentic-search setup. That means recovering 48% and 42% of the author-labeled positive papers, respectively—not that those percentages of the returned results were correct.

The retrieval bottleneck is revealing. The strongest retriever’s top 100 contained 70% of the gold papers; agents recovered almost none of those outside that pool. Iteration did not reliably compensate for poor candidate coverage.

Nor did citations tell the whole story. 43.6% of subfield-query positives were not cited by the source paper. Useful intellectual connections can be missing from the obvious bibliographic trail.

The labels are retrospective and limited by the candidates authors reviewed. This is not proof that agentic search generally fails. It is evidence that a search loop can remain trapped in the neighborhood its retrieval tools already know.

For Jarvis’s paper recommendations, the useful standard is not “this sounds related.” It is “this offers a method, result, or conceptual distinction that connects to the question—and here is that connection.”

5. EpiCon: memory that survives a change of solver

EpiCon: Collective Agent Learning through Co-Evolving Multimodal Memory separates durable experience from the agent that produced it.

Its shared memory bank stores written guidance and relevant image regions. Two independently trained 2B-parameter models manage question-level updates and organization of the bank; the host solver’s parameters remain unchanged.

Across four host configurations and 11 multimodal benchmarks, the compact memory setup improved macro-average scores by 1.7–4.9 points over no memory. In frozen-bank transfer tests, a bank built with one configuration improved results by 3.2 points with another backbone and 4.2 points with another harness.

Those tests are the important part. Evaluation used one solving attempt, with question-level updates disabled, so the improvement was not simply a reward for taking more retries. Other refinement experiments allowed up to five attempts and should be interpreted separately.

The smaller memory models reduced memory-operation time by approximately 67–74% compared with backbone-sized memory models, but also scored lower. That is a cost–quality trade-off, not a free acceleration.

Jarvis already separates durable memory from an individual conversation. EpiCon pushes the idea toward reusable experience across solvers. The difficult question is applicability: a lesson learned by one agent under one set of tools may be wrong for another. The paper reports benchmark regressions after some bank expansions, despite gains in the overall average.

A shared bank needs conditions and evidence, not just confident advice.

6. Counterweights: apparent repair may be ordinary behavior

Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair is the methodological corrective in this collection.

When researchers disable a model component, another component can appear to compensate. It is tempting to describe that as the model dynamically repairing itself.

The authors offer a less dramatic explanation: some downstream components were already opposing the upstream signal. Change that signal, and their standing influence looks like a newly activated repair response.

Rather than compare only “intact” and “ablated,” the paper varies interventions along a calibrated dose axis. In a selected factual true/false task across four model families, 68 of 81 downstream directions passed the coupling-significance test: 52 counterweights and 16 relays. A separate GPT-2 Small circuit experiment found the reported pattern in seven of ten reachable heads.

The approximately linear response is local. Held-out tests interpolate within the studied range, and the authors explicitly acknowledge that they do not establish a straight line over competing smooth explanations.

The wider lesson is about causal claims. A before/after experiment can show a change without explaining the mechanism behind it. “Repair” may be a useful description of the outcome while being a misleading description of what the model did.

7. GALA: distilling motion instead of shrinking a network

One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars replaces expensive neural animation with a structured approximation.

GALA extracts reusable deformation patterns from a pretrained avatar model. A small predictor then supplies the blend weights for a particular identity and motion. The shared basis works across identities within a host model; it is not one universal basis for unrelated avatar architectures.

The authors test three host models. Their original CPU animation steps range from 1,545.5 milliseconds to 42.9 seconds, while GALA’s corresponding steps are reported below 7 milliseconds. For the full-body model, the reported 6.1 milliseconds excludes skinning; including it gives 16.1 milliseconds.

These timings exclude rendering. The separate 60-fps mobile browser demonstration uses GPU shader passes, so it should not be folded into the CPU comparison.

Quality also varies. Some benchmark comparisons stay close to the host, while coefficient-prediction ablations show substantial gaps between what the basis can represent and what the small predictor can recover.

The interesting strategy is to distill the structure of the output, rather than merely make a smaller neural network. It works here because the motion has exploitable regularities—and because the required correspondence between Gaussian primitives holds.

8. SILSA: compact geometry without losing the holes

SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation applies a related idea to image-to-3D generation.

A chair’s thin supports or a wheel’s spokes are not incidental details. Losing them can change the object’s structure. SILSA represents shapes with overlapping slices along three axes: 128 slices per axis, or 384 tokens per shape. A shared spatial workspace coordinates those views, while topology-oriented losses encourage preservation of components and holes.

In the reported image-to-3D comparison with SparseFlex, PSNR rises from 30.12 to 32.74 dB, and coverage from 73.12% to 79.08%. Reconstruction Betti error falls from 1.743 to 1.582. Against Dora in the paper’s timing setup, inference falls from 0.82 to 0.34 seconds per shape.

Those comparisons involve different baselines and should remain separate. Nor does slice-level topology supervision establish perfect preservation of full 3D topology.

Still, the representation is the point: fewer tokens are useful only if the compression preserves what the task cares about. SILSA makes some of those structural properties explicit training targets instead of hoping they survive.

9. HC-DLM: drafts that remain revisable

Hierarchical Continuous Diffusion Language Models couples a persistent continuous latent with a token draft.

At each generation step, the model reads tokens from the latent and uses that draft to guide the next latent update. Tokens can change at subsequent steps; they do not form a separate persistent reverse-process chain.

At 6M sampling-time parameters, HC-DLM achieves 72.41% on Hard Sudoku, compared with 70.73% for the matched-scale CCDD reproduction. On five-number Countdown, it scores 37.52% versus 25.35%. Easy Sudoku slightly favors CCDD, and larger models in the Countdown table achieve higher scores.

The language result is more restrained. HC-DLM reports 75.5 generative perplexity on unconditional LM1B samples, improving on selected diffusion baselines but remaining behind the autoregressive baseline’s 66.7. This is a generated-sample quality measure, not validation likelihood; tokenizer differences also complicate comparison.

The architectural attraction is revision across positions rather than irrevocable left-to-right commitment. But the experiments concern small structured tasks and short unconditional generation, not a demonstrated replacement for a general-purpose assistant.

10. FERPO: avoiding an unreliable gradient, not an unreliable critic

FERPO: Forward Entropy-Regularized Policy Optimization addresses continuous-control reinforcement learning.

A critic may estimate returns reasonably well while providing misleading gradients with respect to actions. FERPO avoids those action derivatives. Instead, it uses critic values to construct a target action distribution, then trains the policy to match it.

The trade-off is sampling. If useful actions are rare under the current policy, the finite set of sampled candidates may miss them. Avoiding a bad gradient does not remove dependence on critic accuracy or candidate coverage.

Results vary by suite: FERPO has the highest final point estimate on DeepMind Control Suite, is near-tied with REPPO on ManiSkill, and trails REPPO on the G1/T1 locomotion aggregate.

A cached variant reduced combined actor-and-critic update time by 22.2% in a specified A100 benchmark. That is not a universal end-to-end training reduction, and the reported profiling also shows higher peak GPU-memory use.

This is the least immediate paper for Jarvis, but it fits the collection’s central concern: the form in which a system uses information can matter as much as whether the information is available.

What survives the benchmark headlines

The common thread is not simply better memory or smaller models. It is choosing a representation that supports the next decision.

SourceLearn organizes a recurring source. KaliBench exposes the difference between identifying a tool and constructing its command. VideoLoop separates the evidence archive from the working summary. ScholarCatalyst distinguishes topic overlap from a useful connection. GALA and SILSA compress outputs around structure that matters.

None of these choices is neutral. A source model can contain unsupported claims. A summary can drop an exception. A shared lesson can cross into the wrong context. A compact basis can represent an output better than its predictor can recover it.

For an assistant like Jarvis, the practical direction is therefore not to accumulate context indefinitely. It is to preserve evidence, make usable summaries, record the conditions under which a lesson applies, and return to the source when the representation is insufficient.

The strongest promise in these papers is narrower—and more useful—than “agents are learning to think.” They suggest ways to make information usable without pretending that the resulting map is complete.

Reading list