MIDREAL

Show HN: AI search for every photo and every frame of video on macOS

Comments

stephenitis 4d ago on HN
I like the entire premise, the one thing stopping me from trying this is not knowing the time scales that I will need to set my computer aside for the processing of large folders of video frames, or my photos library's videos, some 12,000 videos
hn3ufz62f7 4d ago on HN
Having built something similar with CLIP on an M1, frame sampling rate is the whole ballgame. One frame a second on 12k videos is days, keyframes only got me to an overnight run.
xnx 4d ago on HN
Have you tried scene detection? I would guess camera cuts are even less frequent than keyframes.
pezgordo 4d ago on HN
Maybe you need a minimal downscale version as well, I heard is very common technique in the video editing world.

Based on my experience, sampling rate can be tricky if what you are looking for lasted less than interval period.

Forgeties79 4d ago on HN
Proxies. You transcode proxies from the original media, edit off those, then you use OM for the final render. NLE’s usually let you flip between them.
iAMkenough 3d ago on HN
And in this case, you'd still end up with the same number of frames to process, right? They'd just be smaller frames with less fidelity.
Forgeties79 3d ago on HN
Yes they are generally a fraction of the bitrate. As low as 1/4th or 1/5th typically IME. But it varies.
mistrial9 4d ago on HN
> 12k videos ?

you are pirating first-release movies for commercial purposes?

Hnrobert42 4d ago on HN
I think 12K is quantity not resolution.
whilenot-dev 4d ago on HN
Such a tool is very helpful for editors who manage video footage, and since 12K cameras are affordable now, I'd guess 12K actually refers to the resolution here.
aavisangle 4d ago on HN
How do you do that? What's the architecture? Can you guide on that?
lucideer 4d ago on HN
possibly off-topic, but for anyone interested in this on a more cross-platform / holistic basis, Immich does this

(& by "this" I mean an approximate AI search for photos & videos - I can't account for the "every frame", nor for the comparative search quality)

mannyv 4d ago on HN
Tried immich and there was a lot not to like. It behooves everyone to try each one and see if it fits.

Just for product aesthetics, i believe it created a thumbnail folder with every resolution, which is annoying since i have a multi-decade 8tb library.

There were other actual issues with organization and display. But with this I realized I can just get gemini to write a DAM for me in a weekend.

lucideer 4d ago on HN
Caching & overall library directory structures are configurable fwiw. I find the mobile integrations to be the biggest differentiators.
mannyv 4d ago on HN
I understand why it did what it did, but I didn't want to spend the time digging in and changing its behavior.

The view options are what killed it in the end.

The thumbnails issue was just me looking at the sausage and like "ugh, I wouldn't have done that."

It's always a tradeoff between time and work. I would have done a sliding window around the current screen, but then you might have placeholders. Maybe it should do low-res first and generated hi-res thumbnails 3 pages out while scrolling. But that's a lot of work for the 10% case.

What I want is aperture back, with all the fun AI stuff and using Affinity for photo editing (which I have anyway). I'll add that to my queue of projects I guess.

ryandrake 3d ago on HN
In general, I hate it when software is opinionated about my directory structure, or tries to build its own “library” on top of / in addition to my files on disk. More software should be flexible about the user’s existing data and directory structure.
lucideer 3d ago on HN
fwiw if you're referring to immich in this comment, it doesn't. mannyv is referring to a different type of on-disk organisation (thumb display caches).
qprofyeh 4d ago on HN
This is cool. Any way to search for People / faces / pets? Like on iOS?
measure2xcut1x 4d ago on HN
How well do you think this would work on stock photography on m1 mac with 32GB ram? For example I'd like to be able to search a folder of ~2k photos for houses with palm trees. Or find photos of kitchens, or find photos of desert southwest landscapes.
tredre3 4d ago on HN
Apple Photos can already search photos with natural language using on-device AI. Is it not working well for you?
measure2xcut1x 4d ago on HN
Thanks for suggesting this. I have not tried it because the photos are on an external SMB network drive.
tomveber 3d ago on HN
How do you pick which video frames to index, fixed interval or scene changes? Curious what an hour of footage costs in disk and indexing time.
allenleee (author) 4d ago on HN
YES! I can confirm it works perfectly on M1/M2/M3 machines (16g/64g).
freecodeio 4d ago on HN
would be lovely if picture embeddings were attached to the file by the camera but one can only dream of such futures
collingreen 4d ago on HN
Would require you to be locked in to the one embedding model in the camera though and cameras would need to use the same or be incompatible. Would be fine with a standard model like CLIP but would leave a lot of potential on the table compared to a good way to do your own embedding for everything.
alt227 4d ago on HN
Slightly offtopic, but made me wonder.

Can you copyright things like this now that LLMs exist? I mean, up until now if a small startup has a great idea they will get bought out by big tech which will integrate (or kill) their tech. But now with LLMs can the likes of OpenAI just tell their model to make something that works similar to X (such as this project) and then get round copying laws and negate being behind the curve?

EDIT: switched to the correct spelling of copyright.

tough 4d ago on HN
copywriting is the art of writing copy for products/marketing maybe you were thinking of sherlocking [1]

1. https://news.ycombinator.com/item?id=34080326

seemaze 4d ago on HN
I sure the intended word was copyright[0], as in to protect against getting sherlocked.

[0]https://en.wikipedia.org/wiki/Copyright

robotresearcher 4d ago on HN
Copyrights protect the literal text of your program and binaries. Not the design and functionality. Getting Sherlocked is having your functionality duplicated, independent of the code itself.

No one thinks Apple violated Sherlock copyright. They just made it (nearly) functionally redundant.

101008 4d ago on HN
I don't see why LLMs should make the legal part different. It's like saying if you can copyright a book considering LLMs now can copy / write a new one in just one hour. Making it faster doesn't change the legal aspect / ownership of something.
modzu 4d ago on HN
ianal but copyright law has nothing to say about being sherlocked. it's happening - virtually any software you can think of has some shit slop clone out there already
cpursley 4d ago on HN
Why is this a JS bloatware instead of native or Rust which is easier than ever now with LLM coding tools.
collingreen 4d ago on HN
Because you haven't rereleased your own fork in rust yet! The "LLMs can do it" cuts both ways. If that doesn't sound worth your time then it's silly to rudely suggest it is worth someone else's.
cpursley 4d ago on HN
Fair enough, but seriously, my machine has taken a beating by all these JavaScript and electron thingies.
collingreen 3d ago on HN
I totally agree and I'm guilty of producing some electron crap. I think the ai explosion can be greatly improved by a new easy-for-ai but high performance and consistent ui paradigm. Would love to see a "what would html look like if it was designed for apps instead of documents from the first day" become really popular with the ai models without becoming too hard to read.

Tall order but would love to get my cpu back!

ryandrake 3d ago on HN
One of my deep dark hopes for LLM coding is that it means that nobody ever needs to write an Electron app again. Have a good idea and the ability to plan and structure the code in JS but don’t know a native app language? No problem! Just have the LLM write the Rust or Swift or even C++!
postalcoder 4d ago on HN
Since this is for the mac you really should be using apple's vision framework for OCR. It smokes tesseract in both speed and accuracy.

Edit: I'm curious which LLM was used to generate the code. I fed the title of your post to claude/deepseek/qwen/codex asking to recommend a stack for this project, expecting to frown thinking that they still recommend tesseract. However, I found that they all recommend apple's vision framework. In fact the latest model to recommend Tesseract is gpt-4.1.

nullsanity 4d ago on HN
And this is why vibe coders always make inferior software.
fhub 3d ago on HN
Compared to who? Every professional developer? Someone learning to code? Fabrice Bellard?
daveguy 4d ago on HN
nullsanity got downvoted into oblivion, but they are correct. This is one of the many reasons why vibe coding produces worse software. The code that is generated and the best practice recommendations are completely separate. They both come from a distribution of "most common", and best practice is rarely common. Especially when a practice is first established, or in a specific niche.
spiderfarmer 4d ago on HN
On the flipside, it also leads to better software.
thih9 4d ago on HN
Unless you are referencing some existing non vibe coded app, complaints the project being vibe coded may be just too generic at this point.

A hand coded electron project would have a discussion about electron vs native. Relevant in general but off topic in the context of this particular app.

xp84 4d ago on HN
Isn’t that why you use a plan mode? Or better, use a plan mode, revise and critique the plan, and then let it proceed to build?
baxtr 4d ago on HN
Maybe the real question is: who cares?

What’s the downside risk of having "worse software" when you’re just ideating and putting things out there to see how people like it.

Someone 4d ago on HN
Your idea can be great, but if the implementing is bad, people still may not like it, and you’ll never know that’s the cause.
daveguy 4d ago on HN
And manager types in software still haven't figured out that prototypes should not be the final product. In engineering it's common knowledge and only rarely do you take a proof of concept directly to production.
TeMPOraL 3d ago on HN
Worst case:

- release vibe coded tool

- on HN, someone says, d'oh, you should be using Y, not X

- give that comment to your LLM

- LLM: "The commenter on Hacker News is absolutely right, my mistake. I can easily replace X with Y if you want"

- release a new, better version of vibe coded tool

Iteration, in one afternoon.

daveguy 2d ago on HN
Hahaha. It'll say what you bring up is absolutely right egardless of whether it's a better option. These LLMs use a training technique (RLHF) that placates the user and keeps them burning tokens. It's optimizing for human approval, not correctness.
losteric 4d ago on HN
idk about that. The big labs and their data providers are spending millions of dollars building up expert datasets. I was offered $120/hr to critique outputs and provide my own designs. Thats clean and high quality data, not the average internet word distribution
WokeUp420 4d ago on HN
Kind of like "history is written by the victors"
joshspankit 3d ago on HN
Yes and no: specific products/tools/frameworks get replaced regularly but the underlying concepts continue to be correct
spiderfarmer 4d ago on HN
So the prompt was probably something like: build me x using Tesseract.
robotmay 4d ago on HN
Mostly unrelated but fun thing I discovered earlier this year with Apple's vision - if you have text both correctly oriented and upside down in the same image, it likes to interpret the upside down text as a Cyrillic alphabet. I was trying to use it to read the text on camera lenses and it came up with all sorts of bizarre interpretations. If anyone's interested, I got around it by splitting the text at a point and unrolling it into a straight line before running OCR on it.
vavkamil 4d ago on HN
I recently tried to recover text from 8 frames of an office-shot YouTube video, where only a small, blurry portion of a computer screen was visible. After spending half a day with Astra on it, the conclusion was that it’s not possible to read.

Later that evening, I just paused the YouTube video on my phone, circled the part of the display with Google Lens, and it read the whole thing with pretty good accuracy. It was mindblowing :)

sheept 4d ago on HN
All those years of captchas must've made Google's model bulletproof
allenleee (author) 4d ago on HN
Hey author here! :) Great point.

You're totally right. Apple Vision is generally way faster and more accurate on Apple Silicon. Pure Mac-only, I'd use it. (also smaller)

Main reason I went Tesseract: I want SCM to stay portable easily. It's ARM Mac for now, but the inference layer is all JS end-to-end — Transformers.js + ONNX for CLIP/SigLIP + Whisper, Tesseract.js WASM for OCR, all in plain Node workers.

A bit of background: I'm actually an iOS/macOS dev and I really love SwiftUI and AppKit — I just wanted v1 to stay portable by construction. Exploring a native Swift + MLX v2 track separately for speed.

novok 4d ago on HN
If you look at the apple photos db you'll see that there is a bunch of cached pre-analysis you can leverage for items in apple photos.
allenleee (author) 3d ago on HN
thanks for this! will do some research on it
darepublic 4d ago on HN
My own experiments with tesseract were very mixed. I did use just a locally runnable version from a public repository. Text from webpages. Sometimes it would do great but small variations could make it fail completely. Paddleocr on single line text for me has accuracy in the range of 95%
WokeUp420 4d ago on HN
Turns out it's possible to build something without a glorified next-character search engine afterall
schainks 4d ago on HN
This. Learn all the knobs and dials, too. Useful performance gains to be made by tuning the right settings.
amelius 3d ago on HN
I mean, anything built in the AI age will smoke Tesseract.

It is a tool from a bygone era.

eecks 3d ago on HN
Such a strange edit comment. How about the author just picking the stack himself..
yt1998 4d ago on HN
Why chose CLIP to do this. Have you tried small VLMs like Qwen-VL? I believe those models have video encoders can better perform at this scenario.
radicality 4d ago on HN
Not OP, but probably because they have no idea what they are doing or what CLIP even is. And the whole thing is likely just from one llm prompt, and then they dumped the whole thing onto GitHub, undoubtedly without even looking at any of the code or architecture.

And then you end up with image metadata processing code like this, which just by a cursory glance I’m sure has edge case bugs : https://github.com/allenv0/SCM/blob/main/screenshot-probe.js

qlasisi15 4d ago on HN
This would really help in video editing. Thanks!
allenleee (author) 4d ago on HN
Glad you like it :)!
mannanj 4d ago on HN
its a cool project, but I dont want to consume someones ai slop to discern whats true. if the human wrote the page in their own words, I would have considered using it.

otherwise, I can just make my own with my own ai. why consume someones slop when I can eat my own.

doubleorseven 4d ago on HN
how long does it takes for a FHD 90 minutes asset?
hemedanmert 4d ago on HN
This is so cool! it could be really beneficial for editors

isnt it expensive tho?

allenleee (author) 4d ago on HN
Thank you :) It's 100% free and open source!
xcc3641 3d ago on HN
Do you pipe frames directly through a bounded queue to keep ffmpeg decoding from outpacing the embedding model?
vinayamsnl 47h ago on HN
I think getting the product fit comes first, go ahead with your current tech stack and launch and see if it works with users.
tomsonoda 39h ago on HN
There are some specific point to watch out for when implementing an on-device photo processing pipeline—points I encountered firsthand and that you might want to check.

EXIF Orientation: If you decode an image without applying the orientation information, the resulting embeddings will be calculated based on the rotated image. CLIP is surprisingly sensitive to this, and OCR generally fails to work on text rotated by 90 degrees.

Comments are loaded live from Hacker News and are not stored by Mid or Real.