Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
🐝
Article: Where the Hivemind Comes From
Burton Lancaster
PRO
RiverRider
6
4
25
Follow
Quazim0t0's profile picture
mike-ravkine's profile picture
John6666's profile picture
40 followers
·
46 following
https://sunstonenorth.com
Space-Bacon
AI & ML interests
Explainable AI
Recent Activity
liked
a dataset
about 14 hours ago
detection-datasets/coco
posted
an
update
about 16 hours ago
SWE-bench Verified scores whether an agent's patch passes the tests. It does not score whether the agent found the right file first, which is the step before it. We measured that step on all 500 instances. A 33M-parameter encoder, BAAI/bge-small-en-v1.5 at 384 dimensions, names the correct file first for 229 of 500 (0.458). Plain text search over the same checkouts gets 35 (0.070). The number that makes those readable is the floor. Hand the same index a bug report from an unrelated project and it still lands the gold file at rank 1 for 5 of 500 (0.010). So 0.458 is 45.8x chance, not 45.8x nothing. That ratio is where the argument is. recall@50 reads 0.954 and sounds like a solved problem. An unrelated report reaches the same top 50 for 0.244 of instances, so the margin over chance falls from 45.8x at k=1 to 13.6x at k=10 and 3.9x at k=50. The headline that looks best is the one carrying the least. All 500 ranked lists are published under CC BY 4.0, so the floor can be recomputed rather than believed. The article also ends with eleven corrections to claims we made earlier and got wrong, including one where the lever we proposed turned out to cost accuracy rather than buy it. We have not measured the patch step. This is the one before it. For anyone who doesn't live in SWE-bench: it gives model a real GitHub issue from a real project and scores whether the code it writes makes that project's tests pass. That single score covers two jobs, finding the file that needs changing and then changing it correctly, and only the pair is ever scored. The first step is what we measured. Nothing here writes code or runs a test, so 0.458 is a hit rate for naming the right file, not a SWE-bench resolve rate. Article: https://huggingface.co/blog/RiverRider/finding-the-file-localisation-on-swe-bench-verifie Data: https://huggingface.co/datasets/RiverRider/swebench-localisation
replied
to
their
post
1 day ago
Black Window — a chat model in your browser tab, on your hardware. A memory that stays on the device that opened the page. https://blackwindow.xyz Open the site, pick a model (about 0.6B to 8B), hit Load. The weights run in that tab, on that computer. After they load, the network can drop. The context window is a working set, auto-sized to that device, up to ~32K tokens. Behind the window is the Weave. Every file, picture, recording, link, lookup, and reply is embedded as it arrives. Drop in audio and it is transcribed. Drop in an image and it is described. A question pulls the nearest passages back as notes. A long document is walked once so later questions can use the whole file, not the first pages. Nothing leaves that tab unless you turn on live lookup or connect a rented GPU box, and the chat says so each time. Prompts can go to the box. Files and the Weave stay in the tab. Console on that page: bw.ask, bw.search, bw.digest, bw.notes. A local relay exposes /v1/chat/completions on localhost so other tools on the same computer can talk to the tab. The tab polls the relay. That is the boundary. Not a server with a policy. Your hardware, a window, a Load button. If on mobile add to home-screen for best performance. If you break it lmk. It can serve a few hundred of you at a time before I have to buy a real server.
View all activity
Organizations
RiverRider
's models
25
Sort: Recently updated
RiverRider/blackwindow-mlc-libs
Updated
7 days ago
RiverRider/srt-reader-heads
Updated
11 days ago
RiverRider/srt-cxr14-pooled-probe
Image Classification
•
Updated
11 days ago
RiverRider/srt-cxr14-linear-probe
Image Classification
•
Updated
11 days ago
•
2
RiverRider/srt-omni-xvendor-towers
Feature Extraction
•
Updated
19 days ago
RiverRider/srt-omni-shared-tower
Feature Extraction
•
Updated
19 days ago
RiverRider/srt-nla-av-gemma4
Feature Extraction
•
Updated
19 days ago
RiverRider/srt-nla-gemma4-artifacts
Feature Extraction
•
Updated
19 days ago
•
2
RiverRider/srt-sunstone-linear-head
Feature Extraction
•
Updated
19 days ago
•
2
RiverRider/srt-verbalizer-v1
Text Generation
•
Updated
19 days ago
RiverRider/gemma-4-31B-it-nf4
Image-Text-to-Text
•
31B
•
Updated
19 days ago
•
57
RiverRider/srt-browser-head-118k
Feature Extraction
•
Updated
19 days ago
•
2
RiverRider/Gemma-4-31B-it-SRT-Sunstone
Feature Extraction
•
Updated
Jul 3
RiverRider/srt-adapter-gptoss20b
Feature Extraction
•
Updated
Jul 2
RiverRider/srt-nla-av-gptoss20b
Feature Extraction
•
Updated
Jul 2
RiverRider/srt-nla-av-gemma2-2b-v1
Feature Extraction
•
Updated
Jun 18
•
7
RiverRider/srt-adapter-qwen3-235b
Feature Extraction
•
Updated
Jun 18
•
14
RiverRider/srt-nla-av-llama32-3b
Feature Extraction
•
Updated
Jun 18
•
8
RiverRider/srt-nla-av-v1
Feature Extraction
•
Updated
Jun 18
•
63
•
4
RiverRider/zooL4nD3r-v0.1
Feature Extraction
•
Updated
Jun 18
•
27
RiverRider/srt-adapter-v22c_a050
Feature Extraction
•
Updated
Jun 18
RiverRider/srt-adapter-v1.0
Feature Extraction
•
Updated
Jun 18
•
17
RiverRider/srt-adapter-v21a
Feature Extraction
•
Updated
Jun 18
RiverRider/srt-adapter-v18
Feature Extraction
•
Updated
Jun 18
RiverRider/srt-adapter-v8a
Feature Extraction
•
Updated
Jun 18
•
31
•
2