Hacker Newsnew | past | comments | ask | show | jobs | submit | phasefactor's commentslogin

Very cool, will there be any recordings? Assuming it is being presented.

It can be both full of fluff and have actionable take aways.

This is dozens of pages long and reading it thoroughly will net you maybe a dozen useful sentences. That ratio is the problem people are complaining about.


"Be bold" by saying nothing of any interest. I have to agree that this is basically just fluff.

Anyone put it into pangram yet?


97% human-written


I guess prolific em-dashes is not considered a sign of LLM writing anymore.


I am not an LLM and use em-dashes and they were house style in my last 20+ years of employment. The whole em dashes == AI thing needs to just stop.


Speaking about style choices, I was curious why you sometimes use the 'f' from the Latin Extended-B block, U+0192, in your comments. I've never seen that before! (I'd meant to ignore it but you seem to be soliciting typography discussions...)


Probably a fat finger of some sort or it came through a copy and paste. Nothing I do deliberately.


It is more about the document containing more than sixty of them. They are a completely legitimate piece of punctuation, but even prolific pre-LLM users did not tend to go that wild with them.


I did not notice anything about the document in that regard that was worthy of note.


It's actually not em-dashes, they are en-dashes, which is either intentionally to deceive (e.g., swapped) or used incorrectly. Both of which are bad, it's up to everyone to decide on which is worse coming from MITality.


Spaced en dashes – like this – are a correct and accepted typographical choice. Whether to use spaced en dashes or (spaced or unspaced) em dashes is a matter of taste or local convention, and varies from one publication to another. Other than a typo or two where the space was left out on one side of a dash, the dashes in this linked document are just fine. See https://en.wikipedia.org/wiki/Dash#En_dash_versus_em_dash


I was not claiming they are not correct punctuation/syntax. My claim is just no one uses them that frequently... When you see dozens of them in the first few pages of a document it raises suspicion.


This is likely an artifact of writing standards - not especially suspicious, it's like MIT Tech Review or the New Yorker using diaresis when they need to say reëducate.


Some authors use a ton of dashes, while others never use dashes.


Sad that people now consider basic knowledge of typography, punctuation, and orthography a sign of AI.


Do you think LLMs invented the em dash? Pre LLM, em dashes were a sign of pretentiousness. Which is why LLMs learned to use them, and why a document produced by a university has plenty of them.


Correct, it was switching to unlimited lending instead of one lend per physical book that got them in trouble.


Easier to send them through a duplex scanner (or put them in a flatbed one if they are fragile). Cheaper than buying the automated ones with the page turning robot arm.

I have done it at home for my books since the mid-00s.


Do you purchase every book two times, or do you have a home devoid of physical books that you enjoy? Genuinely curious


Not OP, but I do similar. I only de-spine a book if it is in poor condition (and I have a nice copy). Some of the material I scan is stapled instead of glued and I can remove the staples, scan and replace with new (not rusted!) staples to return the book to its original form.

I have books but would have 5x as many if I could not capture them digitally. (When I go to move or go through a purge, some of the books I scanned do go to a used book store.)

So I scan in part to keep my physical book-footprint smaller, but also my scans all get cleaned up and uploaded to archive.org. Mainly I scan young-adult science books from the 50's and 60's (since they were so influential and have all but disappeared except on eBay and the like).


Thank you for your service, genuinely.


Not OP, but: If it's something I want to work from, I get a printer to slice the spine off and rebind it with a ring binding. That way it lays flat on the desk.

You've trashed the book (and its lifespan) but some books are for using, not keeping.


Sometimes? But for scanning I can use a cheap worn copy as long as the pages are clean.

Most books I would rather have a digital copy. I had four or five book cases full of stuff and it makes moving difficult.

I originally started scanning 20 years ago when it was more difficult to find PDFs of basically whatever you want. Now it is only random old books I find at the used shop or online for cheap and somehow cannot get from archive.org or elsewhere.


Creating the python extractor at all was sort of bizarre. The argument given was that it is too difficult for a developer to spin up a whole dev environment just to compile a tool, but the C extractor is a single C file that only depends on the C standard library and zlib.

This feels more like your mom likes to wash dishes by hand and you have a known good dishwasher that needs to be plugged in, but instead of plugging it in you build a new one that you tested a couple times and it seems to work but who knows? At least you saved yourself all the work of plugging in that other dishwasher?

Still a cool project to publish the source code, but the python port side quest seems pointless.


> Creating the python extractor at all was sort of bizarre. The argument given was that it is too difficult for a developer to spin up a whole dev environment just to compile a tool, but the C extractor is a single C file that only depends on the C standard library and zlib.

He just wanted the extraction utility to be easily available to other people without them needing to spin up a whole development environment. That's quite community-spirited. And honestly, having tried and failed to get a number of things to successfully compile (after spending god knows how long installing different and specific Visual Studio or cmake packages) in the past, I get it.

"For this archive I wanted a Python version that could be kept with the recovered files and run without compiling the original C program. Most developers have easy access to Python, and it's quicker to run a one-off Python script than it is to spin up a dev environment, configure and compile and run a compiled native executable."

> This feels more like your mom likes to wash dishes by hand and you have a known good dishwasher that needs to be plugged in, but instead of plugging it in you build a new one that you tested a couple times and it seems to work but who knows? At least you saved yourself all the work of plugging in that other dishwasher?

If we really want to flog this analogy :) then it's more like: I rent out a holiday property. It has a dishwasher, but it takes 20 minutes of messing around for each new tenant to learn how to turn it on, and there's a non-zero failure rate. So I replace it with a different model without this difficulty.


I just wonder why he didn't boot up an Amiga emulator and copy the files out that way. Could even run DiskSalv on the disk image to look for data in the slack space. Or he could have used a tool like what comes with ADFlib (https://github.com/adflib/ADFlib) to do it. This seems like a thinly-veiled ad for AI and for this person's company, at the end.

Edit: I loaded the ADF into WinUAE and ram it through DiskSalv to check for deleted files. There were none.


> This seems like a thinly-veiled ad for AI and for this person's company, at the end.

There was a very clear (what's the opposite of thinly-veiled?) advertisment for the person's company at the end! But I didn't read is as an advert for AI - that was just a tool that was used along the way.


There are tons of tools the blog author could have used before resorting to using AI to convert an existing library to Python.

AI isn't "just a tool", the whole reason it exists is to offload thinking. I wish I'd noticed this needed "recovery" before this person did, I'd have extracted the files and been done with it. AI was not necessary for this at all

The ad was well past the lede


The training kicks in and my knee-jerk reaction to not one of the graphs starting at zero is to discount the trustworthiness of the entire article, whether that is well deserved or not...


For most of these metrics, zero is not a logically possible data point. For example, somebody with an HbA1c of 0 would be dead.


Well, that is a metric that we all achieve as we age


Registration is now open for Qiskit Global Summer School 2026: A decade on the cloud—the latest entry in an annual, global program designed to help students, researchers, and developers build practical skills in quantum computing. The two-week event will run from 13–24 July 2026.


Love it, the article referring to a statement by a LinkedIn spokesperson: "The first part of that statement is false, as you can see from the screenshot above. Given the obvious untrustworthiness of that half of the statement, we didn't bother wasting any time trying to evaluate the second part."


To be technically correct: if even a single non-premium member can, for some reason, see who viewed their profile, then the statement "only Premium members can see who has viewed their profile" is false.

So technically, you can't say that the first part of the statement is false from the screenshot.


They do say they won’t bother, but the rest of the article is actually precisely covering this second point, aka Article 15 of LK Privacy Policy


Rhetorical argument is rhetorical?


It's covering article 15 of the GDPR, not of LinkedIn's Privacy Policy.


Also, I just checked, and LinkedIn's privacy policy page doesn't contain any information about who viewed my profile in the last year. No usernames, no company names, it's just a generic privacy policy. So the data isn't there either.


But the statement wasn't false, was it? I'm not a paying member of LinkedIn, and I can see who visited my profile.


What you have is the ability to see every 5th person who entered your house.


Seriously, this is what I miss the most in legacy media. Much too often "journalists" will simply relay politicians' statements uncritically, when they're obviously fallacious or straight lies. This is very refreshing on The Register's part.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: