RNS Logo

rns.recipes

◈ 9ce92808be498e9e05590ff27cbfdfe4
RNS 1.5.2 released https://pypi.org/project/rns/ | Nomad: a8d24177d946de4f1f0a0fe1af9a1338

The Slopware Scrapers Have Arrived

Started by Mark bc7291552be7a58f... ·

aetherlab 509723a0ccb60610...
edited #21

Mark wrote:

It's saddening. Create a network where people are actually sharing their own personal thoughts, opinions, writing and all kinds of other stuff, and some idiots goes and scrapes it all at a 10-minute interval to showcase a copy of it on the internet.

Hold on...you are in the same struggle, every meaningful, talented and capable person on a life-changing conquest has been - the Struggle With The All-Knowing Idiot. History has shown very clearly, that whenever someone like you does something like you do, around him flocks the type of humanity, that identifies itself with you and what you do, taking the task at heart to carry the light, expand and improve, in thy name's glory entrusted and blessed.
There are also the others, who know so much better, that they can't sleep well at night, until they show the world, that albeit you're your worth, they can and will ffs - IMPROVE this "thing" of yours into oblivion.
And than there are the ones, that actually know what they are doing. And those are ill in their intent. And shall be hunted down.
Thousand upon thousands of brilliant ideas over human history, inventions, intents have been drowned and diluted by this wave of human-ness, and standing up to it is crucial.
There will be more idiots with more ways to fuck this up. Up until now I have seen mostly pure idiotism and no true ill intent, as you know very well. RNS has either to grow resilient to this, or fail. And for now whatever you have done to counter this was successful and useful for growing RNS some claws. Please treat them as this - a hardening session, not a thing to get you dishearted. Eventually the true enemies will come, and they will not be idiots. They will know the code and know where it is weak. And I hope RNS will be ready to take them head on.
Hold on!
Bill
P.S. Long before idiotism was coined as a medical term for a specific medical problem, the ancient Greeks gave the following definition for the word: Someone who goes on doing the same thing, but expects different results. I think we all here would adhere to the Greek definition and would differentiate ourselves from the medical one.

Anonymous
#22

I also have had idiots trying to ruin my small projects all the time, I think the only condition it needs to attract these kinds of people is if users can interact with each other. Then it inevitably turns into a dick measuring contest. I know from experience how frustrating this can be, lol. Mark, I wish you all the best in solving this without affecting normal users.. much.

In the worst case the purpose and usefulness of Reticulum didn't really get broken. It's not about creating one big network, remember? It'll be fine. :)

Necom 15c9a46c70d24d73...
#23

I wonder if some of what I’ve been seeing is related to any of this. 1st thing is It seems a lot of link establishments to the git repo fail for some reason. It took me a few tries to update rns yesterday, and when I tried to update nomadnet it took so many tries I eventually just gave up for a while until I remembered it and tried again, and then it worked.

And the 2nd thing is, today when I updated reticulum again I tried simultaneously running rnstatus -m -l -R so I could see the speed of the ESP32 802.11 interface I made, that’s when I noticed the raspberry pi I’m using to connect to the reticulum network had over 13000 entries in the link table. Not constantly, it was fluctuating, and at some point I saw it at like 9000. The thing is yesterday it reached like 8 entries at most, the only configuration change I’ve made since yesterday was changing a tcp client interface to a backbone interface, since those are supposed to be more efficient some how, (though I’m not sure in what way, I haven’t looked at the code). Since I have connections to multiple nodes, traffic from people connecting to those nodes might be going through my node even though no one is connecting directly to my node, and if the backbone interface is so much more efficient then maybe it’s also faster and allows for announces to be going through my node slightly faster than some other nodes, meaning the network prefers a lot of paths through my node. But 13000 link entries so suddenly still feels like a bit much.

Actually I just checked again and it seems to be fluctuating around 21000 - 24000 entries now.

Necom 15c9a46c70d24d73...
#24

Alright I just updated reticulum on that node and restarted rnsd, and in 54.22 seconds it reached 1956 entries in the link table (1948 active)

Anonymous
#25

Goodbye reddit...

Mark bc7291552be7a58f...
#26

@Necom, yes it's definitely related to this. Aleph was serving 50 page/file downloads concurrently most of the time before I blocked it. The thing is still trying to get recursively scrape every now and then, but it's just getting empty responses now, so load is much less affected. It's actually running on a Raspberry Pi, so I'm pretty surprised with how well it kept up.

Also, yesterday I was constantly restarting the rngit service to play around with injecting garbage back at the scraper, so that might have complicated actual downloads for you a bit as well. Pardon the service interruption, I couldn't resist.

And also, the insane amount of link table entries comes from this junkware as well. Normally, on my central and well-connected nodes, there'll be around 500 to 1200 active links at any given time, peaking sometimes a bit higher with high usage. When this shit started, a few nodes were handling over 60k concurrent links.

if users can interact with each other then it inevitably turns into a dick measuring contest

So very true, unfortunately.

Mark bc7291552be7a58f...
#27

Don't worry about that @Rudi Mentaire, using it is not going to have any bad effect on the network, you're just fetching pages from their node, and at least I don't think that's actually causing additional scrapes. How would you have known the thing is actually a slopcoded attack machine? But in the shining light of hindsight, this should be a good lesson for all of us to be better at spotting these things, and being very skeptical and suspicious from the start.

If someone suddenly dumps a "new thing", that looks just slightly like the person being an "idea-and-an-LLM" guy, it's probably a good idea to stay the hell away from that, unless you are yourself willing to ensure that the system you're about to use it not actually an "unintentionally malicious" attack machine.

And yes, these are attacks, but an extrme amplification of an old pattern, where people with very little care, knowledge or skill can do damage by doing things without any consideration. The new thing is, that now they get a shiny user interface a gradient hover effects on top of it, and a sycophantic machine telling them they're dev-ops starts.

Mark bc7291552be7a58f...
#28

Thanks for the encouragement Bill, it does calm the frustration a bit to read it. I completely agree with you, and it's not like I have been ignorant to these things at any point from when I started this endeavor (which is also why I've always worked actively to keep Reticulum as low-interest and under the radar as I have), but oh my could I not have predicted the sheer scale and dimension of how bizarrely it would turn out when it finally started to hit!

It's a strange inversion. The actually ill-intended and malicious actors do not frustrate me nearly as much as the first classes of people you describe. The true bad actors are (still) few and far between, they are much easier to deal with, and they are much more predictable. It's also that the moronic chaos-bringers drag a lot of other people into their wake, who unknowingly and unwillingly contribute to the tsunami of crap, chaos and confusion while thinking they are actually doing good; but in reality are causing exactly the dilution and simulacrization of the core idea and innovation, like you describe.

But you are right, seeing it as an excercise in resilience and a free staging area for pressure scenarios is probably about the most constructive approach there is. Although I must say the cadence is exhausting. I am already well aware of (probably) more or less every remaining weak point in the system as a whole, and have been planning and working tirelessly to eliminate them all, one by one, over the last many years. But I am just a normal guy, I could do with a rest and a trip to the beach just every now and then.

Damn, I miss the days when Reticulum was small and unknown.

Mark bc7291552be7a58f...
#29

Also, if anyone actually wanted to create a real, proper search engine for nomadnet sites, they could have done it the careful, considered and not entirely insane way:

  • Listen for announces, de-prioritize scanning higher-hop count nodes
  • Set a small hard max limit on network distance to actually scan
  • Limit requests per day per node to one, that's perfectly viable and causes very little load. Every node you index still gets indexed every month or so.
  • Only index the front-page, and a maximum one or two levels down.
  • Set a hard maximum on pages indexed per node; I wouldn't place this at higher than 20 or so.
  • Write a prioritization algorithm that decides what page to fetch and index next, for the next one-per-day slot.
  • Then determine an overall topic / semantic adjacency matrix based on vectorized tokens of the captured sample data (for example), and use that as your search base.
  • Measure link RTT. If above a sensible maximum, don't index the node at all - it's probably on a low-bandwidth connection, and doesn't want to be scraped.
  • Create an easy way to permanently signal to your scraper that a node never wants to be contacted again. How much of a no-brainer is this?
  • Lots of other intelligent decisions, actually considered and weighed by an intelligent human being.

It's really not so hard to come up with something like that, took me about 2 minutes. But I guess the problem is, that the people creating all this crap are not at all interested in the "thinking for two minutes" part. They're interested in wanking while watching the output from Claude stream across their monitor. Consequently, they should stay the fuck away from creaating software.

The analogy is close to someone deciding on a whim that they want build an 18-wheeler in their backyard, and then take it to the roads. No brakes? Well, my "AI" didn't make that! How could I have known!? Squish, now the kids on the side of the street are marmalade.

All of that being said, I think a nomadnet search engine is a useless and horrible piece of crap idea personally. If it really has to exist, it should be opt-in. But much better to have sites organically linking to each other, and discovering new ones via announces. Personally, I don't want the same search engine centralization and dominance that was the beginning of the end for the internet.

steveplays daf042c1c000a24b...
#30

Mark wrote:

...
Damn, I miss the days when Reticulum was small and unknown.

Maybe temporarily removing the GitHub mirror and hiding reticulum.network from search results of popular search engines would help, by reducing discoverability?
That way you could potentially reduce this extra workload for a bit.

Dedicated people would still be able to get started with Reticulum, with more initial effort in finding the manual.

aetherlab 509723a0ccb60610...
edited #31

Mark wrote:

Damn, I miss the days when Reticulum was small and unknown.

Please don't. With all the great, capable and talented people now involved in not only using it, but actively developing and adding true value to it, helping you out, coming up with their own great ideas and actually realizing them... its brutiful! (This one I stole from Nickie!) It's so much alive and kicking, yes, the puberty period is always heavy and laden with worries, tiredness and not much sleep. I have one at home, I know...
Yes, RNS is in it's puberty now. Went in it somewhere around the first announce flood explosion, and just as a nervous 12 year old it is erratic, full of crazy incoherent ideas, moods, brake-downs and dramas. It matures in a very organic way in the hands of people who love it and the person who has created it.
I am proud I have been a witness of this kid's rise to prominence. It has a long way to go, but it is a labour of Love and a true hero for everyone, respecting freedom, personal responsibility and solidarity. Thank you so much for all of it! I am not a coder, my language of choice is soldering, but it has and is a privilege for me to engineer, create and deploy devices for Reticulum.
In the beginning it was simple, yes. But you were a lone parent. You are no more.
And I will have a Weihenstephan Weiss now... Cheers!

P.S. Also cheers to my persistent downvoter! :D You live in my P.S.'s now!

Anonymous
#32

Hi, I always wondered how robust this network is? And what if it replaces the internet? What if millions of people use it..or more... It's definitely not easy, but I think it would be great if it could make this leap...
And it definitely has a lot to anticipate in the big uses of this network...
All the best, Mark (:

bergie f9477df559d52317...
edited #33

Mark wrote:

  • Create an easy way to permanently signal to your scraper that a node never wants to be contacted again. How much of a no-brainer is this?

Hmm, "standardized" path for serving a robots.txt over a Reticulum request?

I think the 20 page limit is way too low. Once we bring our boat's log over to NomadNet, it'll be hundreds of pages that would be totally fine to index. But at the same time indexing every rngit commit page makes no sense.

With robots.txt (or something similar) the party running the serving node could decide what's ok and what's not.

SevenFourTwo
#34

bergie wrote:

Mark wrote:

  • Create an easy way to permanently signal to your scraper that a node never wants to be contacted again. How much of a no-brainer is this?

Hmm, "standardized" path for serving a robots.txt over a Reticulum request?

I think the 20 page limit is way too low. Once we bring our boat's log over to NomadNet, it'll be hundreds of pages that would be totally fine to index. But at the same time indexing every rngit commit page makes no sense.

With robots.txt (or something similar) the party running the serving node could decide what's ok and what's not.

We'd likely need an automated enforcement mechanism of some kind although it would be hard to make one that would work well without complicating the nomad page protocol or reducing anonymity. Stamps could be used but they could waste power.

Anonymous
#35

I only recently read about this project on /g/, which has a thriving community of unrepentant vibecoding idiots. It's been over 20 years since anything involving an internet connection you needed to start with building walls. Install NoScript, protect your SQL queries, fail2ban, and so on. Reticulum being accessible via the internet makes its abuse inevitable the second idiots catch wind of it. But even if LLMs didn't exist and the network were somehow only accessible over wireless mesh, it would eventually become a problem.

Many have spent their whole adult lives immersed in the sewage of modern internet culture. They know no other way but centralized information. Check the wiki, watch the streamer, read the top posts, etc. The Zen of Reticulum's ideals are compelling because they harken back to the old internet, but the hordes' brains are too poisoned to play along voluntarily. The coercion of centralized authorities is what keeps their behavior in check, and in the absence of explicit design to curb their behavior, they will run amok without principals and bring all the problems of the modern internet with them.

Reticulum offers a rare glimmer of hope as an important project that goes against the slop currents and could heal the minds of its newer denizens. I hope you are able to overcome the challenges, but burnout is also understandable given the problems took decades of work by many to entrench.

burger ace748e1c4e3fd4e...
#36

oh for fucks sake
we cant have anything good can we?

Anonymous
#37

What does the word "network" even stand for, after all? You people ever wondered how to distinguish this kind of misunderstanding? Like, if someone says the network is: for people, for individuals, for me, etc. Who will be right and who will be wrong?

I'm asking this because the context is directly affects the project. If reticulum's vision is to allow anyone to operate their own networks and at the same time the author says that Aleph's behavior is harmful for the network. Then there is definitely a misunderstanding about the network itself. Or if someone might say that the misuse is rather egoistical than linguistic (i.e. this was done on purpose), then again, who is to blame for this if project strives to be self-governed and uncenralized?

I'm wondering whether reticulum might end up shifting toward a more restrictive and hierarchical design because of all of this. Although, I don't want to be intrusive about it, just thinking out loud.

Zenith
#38

Anonymous wrote:

What does the word "network" even stand for, after all? You people ever wondered how to distinguish this kind of misunderstanding? Like, if someone says the network is: for people, for individuals, for me, etc. Who will be right and who will be wrong?

The network in this post means the network of publicly facing backbone interfaces running on the TCP/IP (capital I) Internet that are all largely linked to each other.

There has been plenty of abuse to these public nodes, which Reticulum handles just fine. Path spam, link spam, announce floods, attempted DoS. All dealt with.

But this is something different. And arguably it's worse than everything in the above category. It's midwits with a Claude Code or Cursor subscription whatever shitting out software that actively uses the worst possible design practices. It's not intentionally malicious by design, but malicious by the fact it's something generated from a black box in which the author has no understanding of.

THey don't read the code. They don't understand how Reticulum works or even understand the basics of the API. And they take it and run with it because they lack the mental model of how LLMs actually work.

Anonymous
#39

I'm just saying that this thing might occur again and the cause of this problem lies in the design itself. Like, there's no point blaming this guy for being "dumb" if network allows it's users to act like this. I just don't see how this can be prevented without additional restrictions and surveillance.

Anonymous
#40

*in reticulum design

Post a Reply

Supports Markdown: **bold**, *italic*, `code`, ```code blocks```, [links](url)

Log in to upload images

Quote
Copied to clipboard