I used to think task bars were pretty neat, back in the 1990s. But once I started keeping more than a dozen windows open, it became a waste of space. And now that I typically have 200+ windows open, the concept of a taskbar is ... like trying to get a fish to ride a bicycle, or like having a rock in my shoe. It gets in my way and provides no benefit to me.
But that's okay. I don't have to use one. Different people have different needs and preferences, and their choice to use a taskbar-free workflow doesn't interfere with their ability to get serious work done.
Same... ish. 2D grid of desktops. Can add/remove them with a hotkey, but I normally have a 3x5 grid. Hyper+Arrows to navigate them. Hyper-s to stick/unstick a window, which also works for moving a window to a different desktop. Hyper-g to group/ungroup windows into a single shared frame with tabs. Hyper+Alt+Arrows to move a window, Hyper+Shift+Arrows to resize, Hyper+PgUp/PgDn to raise/lower, Hyper+Tab to change tabs in the current frame, etc. Lots of other keys, mostly Hyper+something.
Each desktop has a project or something, a related group of windows which I leave open until I'm done with it. Some desktops have remained more or less the same since the 1900s since they've been relevant and useful for that long.
Over time, my desktop hasn't gotten fancier... it has gotten simpler. Bells and whistles are inversely proportional to useful life span. Similarly, over time I've relied less and less on big, full-featured, specialized applications and databases... and more on plain text files, a text editor, short scripts, clean filesystem organization, and a whole lotta terminals.
> Don't make me explicitly keep track of things, just hold onto it all for me!
>
> There's some universe where my desktop is basically a Notion-like thing, and I can just throw things into semi-persistent workspaces (that I close out when I'm done). Desktop Environments that make this work nicely will be quite cool IMO
>
> (A simple and dumb version of this is the BeOS "group various windows together, even though they aren't from the same application" stuff. Very cool stuff IMO
What you're describing sounds like it has been possible for a long time.
I've been doing it on my Linux desktop for decades. I throw things into workspaces which just ... stay open until I'm done with it. The same windows with the same layout and the same history and state, everything right where I left it. They stick around as long as they need to, which can be weeks, months, even years. I even have some desktops which have stuck around since the 1900s, though they do gradually change and evolve over time.
... and any windows from any program can be grouped together. Grouped into a desktop, or even grouped into a single window frame with tabs.
And nearly everything is automatically logged, so I have detailed history going back well into the past.
This stuff isn't in the big mainstream UIs, but it is at least possible with a bit of work to set it all up.
- Basically never logging out. Everything just stays running.
- Lots of desktops, usually one per project or task or ... whatever grouping makes sense.
- No minimizing windows. No taskbar, tray, dock, etc. Things just stay open, and I change desktops when I want to change to a different project or task or whatever.
- Logging at multiple levels. Like, there's a script which checks the complete session state every 15 seconds, and logs it if it has changed. It's exported as a tree or outline of each desktop, each tab group, and each window... with the title and size+position. Archived daily. That way, if the computer crashes or something, I have a recent snapshot I can restore.
- Also, detailed logging of what I'm doing. Like, every 100ms or so, a daemon checks the title of the focused window. If it has changed, it logs the new title and the timestamp, and how many input events happened since the last log entry. If the time of day changes to a new minute, it logs the number of input events since the last log entry. These logs can then be used for all sorts of useful things. Like analyzing my erratic sleep schedule and predicting when I'll be awake so I can schedule appointments. Or automatic time tracking for billing clients.
- I make sure to put useful info into my window titles, in each program. This makes the logs way more useful.
- Also, whatever other logging is relevant... like zsh history, autojump for getting to oft-used directories easily, a todo list / calendar text file with one heading per day and a list of stuff relevant to that day, a new directory each day for miscellaneous files which are worth saving but don't fit elsewhere, a lot of screenshots, an entire web archiving system to save copies of pages I view, etc.
On a side note... a lot of this stuff is explicitly forbidden in Wayland. It doesn't allow these features to exist, because most of this stuff is viewed as a security risk. So if you're using Wayland, entire categories of useful features are off limits. Their reason is basically "what if a hacker collected data about you? what if a hacker wanted to remotely control your computer?" ... but they ignored the case of "what if the user wants to collect data about themself? what if the user wants to remotely control their own computer?".
I was at Canonical during the Ubuntu Phone thing, and it was wild to watch that play out. It started out with a bunch of good ideas... but somehow, every last one of those good ideas got overturned and reversed.
Like, "let's put the buttons at the bottom so the user can reach it with their thumb while holding the phone" got turned into "let's put every important button at the top left where it's literally the hardest place on the whole screen for an average user to reach during typical use".
Or, more importantly, it started out as something like... "Debian/Ubuntu has a rich repository of high-quality full-featured battle-tested popular programs curated by community experts, designed by and for users to empower themselves, so lets give users a way to run all those beloved programs on mobile devices". And then it became "let's get rid of the package repository and forbid users from running anything from it, and build a brand new system which is incompatible with everything, and instead has a commercial app store where corps can sell proprietary software and Canonical gets a big cut of each sale, and the OS has only a small set of proof-of-concept placeholder programs, and is designed to restrict the user so corporate partners can have more profit and control".
And then when it didn't work out, the company kinda doubled down on some of the worst parts of it and went all-in on Snappy.
Instead of bringing phones up to the level of a desktop, as promised... the company tried to bring desktops down to the level of a phone. And that trend has been increasingly popular in the corporate Linux world lately... "Androidifying" desktop Linux.
From an insider's perspective, the Ubuntu Touch thing looked pretty promising... and then management pushed everything sideways. The regular developers kept trying to fix it and make it good, but then upper management would step in and forbid it. The VP of "Future Engineering" in particular. Lots of stuff was forbidden because it "goes against the vision". So things didn't work, and couldn't be fixed.
The worst was one day at a big sprint... the VP of Business gave an all-hands powerpoint presentation right after the workday ended, with a catered dinner afterward. He looked and spoke like a caricature of an Italian mobster, and spent 45 minutes gleefully explaining exactly how we were going to screw the user to make more profit, by taking control away from the user and giving that control to phone corporations. Like, selling corps the ability to place ads directly onto the user's home screen. He was so excited he was practically vibrating.
Afterward, I just went straight to my hotel room and cried.
Eventually, the phone project was cancelled because it had failed pretty spectacularly, for exactly the reasons everyone said it would... and upper management laid off nearly the entire Ubuntu Engineering business unit, about a third or a quarter of the company. The people in charge had decided the problem wasn't a bad design... it was just all those uppity devs who wouldn't follow orders. They kept the managers who were responsible for making it fail, and got rid of all the devs who tried to warn management that the "vision" was broken.
At least sabdfl seemed somewhat aware of the situation... aware of how bad it was. A day or two beforehand, he started the day by giving everyone a pep talk. The talk was pretty somber, talking about how we were supposed to have a mature product to launch by now, but instead we were cancelling the launch and hoping for an optimistic scenario where some tech magazines would at least say it "showed potential". The best case scenario he could muster for the team's motivation was ... we might get a couple reviews stating that it could eventually be good someday, with enough work.
I've been using approximately the same desktop setup for ages, and yet, it often still feels like I'm living in the future because mainstream desktops still haven't really adopted some of the best features, and in some ways are moving backward. This is kind of a disorganized laundry list, but ...
For one, the ability to change how the desktop itself works while it's running, without restarting it... modifying the code on the fly. And if I manage to crash the window manager or some other desktop component, easily just restart it without having to log out or restart all the other programs.
And being able to pretty easily run and manage hundreds of windows. Log in, open one desktop for each project, and just leave it running until the project is done. Maybe reboot for upgrades once every year or two, but otherwise no disruptions.
NO DISRUPTIONS. I don't even have the underlying plumbing installed for notifications and stuff like that. They can try to interrupt me, but they'll find those API calls aren't plugged in to anything.
Freely mixing and matching tiled and floating windows, each of which can be its own tab group with multiple arbitrary windows sharing the same frame. Super handy having free-form placement and grouping of absolutely anything.
No minimized windows. No task bar. No icon tray. No desktop full of icons underneath the windows. The entire screen is one big canvas, to be used however you like. If you need more space, just hit a key to open another desktop.
Manage nearly everything with a keyboard. Mouse usable but not required.
Sloppy focus, of course. (i.e. focus follows mouse, but stays focused if mouse "falls out" of a window into empty space) Not click-to-focus. And definitely no raise-on-click or raise-on-focus.
Mouse keys built into the keyboard, for clicking, scrolling, and moving. Mouse warp too, to teleport the cursor, like having two mouse cursors and a key to swap between them.
A full set of launcher hotkeys which are easily configured and managed by making short shell scripts. Like, F1 to F12 run scripts from ~/bin named like "f1" to "f12", and "shift-f1" and "ctrl-f1" and "shift-alt-f1" and "hyper-shift-ctrl-f1", etc... if a script exists, the key is mapped. Ideally just symlink the named key to a program or script with a more descriptive name, like "shift-f3 -> open-todays-journal-in-vim.sh" or something like that. Easy to see exactly what each key does, with a "ls -l" or similar. (can send a normal f1-f12 event by holding Fn and hitting a number key or the two keys next to the numbers, but I barely ever have a reason to)
Network transparency for everything, of course. It's wild that newer desktops have chosen to remove such a useful feature.
Work at my desk with a notebook to either side, all using a single keyboard+mouse... just slide the cursor off one screen and onto the next. Pick up a notebook and move to the living room. Use its keyboard+mouse to control the notebook plus the television, in much the same way... slide the cursor off one screen and onto the next. Pack up and go to a coffee shop. Control the remote home computers from the local notebook, using their running desktop session. Or run programs natively on several remote computers, with the windows all opened together on one desktop on the notebook.
Automate inputs easily, whenever desired. Record and play macros which work in any program.
Configurations revision-controlled and synced across all my devices, privately, without any data leaking to a corporate cloud.
Organize data into a rich set of nested topic directories or project directories, instead of organizing things by file type into places like "~/Pictures", "~/Documents", etc.
But also, a directory for each day, to hold all the other misc stuff which doesn't fit elsewhere. And a key to easily access it.
Shell scripts can easily know what program I'm using in the GUI, and what directory the focused program is working in, so scripts can take context-sensitive actions.
Automatic private logging of what I do, with easy time tracking and arbitrary queries, using simple plain text log files which are small, easy to grep, and easy to build custom tooling around.
Human-readable plain text everywhere. Plain text files are versatile, powerful, portable, compact, timeless, accessible, and future-proofed.
No churn. No need to retool every few years. Long-lived tools which are mature and robust. Maybe not the flashiest or most fashionable tools, but "bells and whistles" and "useful life" are typically at odds with each other.
A rich vocabulary of small, simple tools which each do one thing well... makes for a powerful, expressive world of deep functionality thanks to the power of language. It empowers the user to do anything they can express. But a small set of huge tools, while initially a bit easier, tends to quickly become limiting and stifling.
I think a lot of the "right answers" for interface design have been figured out for a long time. That is not to say new interfaces are bad... but anything which throws out the old battle-tested solutions in favor of something radically different is going to take a long time to reach anywhere near the same level of usefulness.
It's a lot of stuff accumulated over a long period of time, sort of pieced together... but here are some of the details:
The overall setup is X11 with a Sawfish window manager. X11 provides a lot of cool features like network transparency and automation and generally being wildly flexible. Sawfish is like the Emacs of window managers, written in a flavor of Lisp, and you can literally rewrite its code while it's running. So whenever I want something to work differently... I just find that part of the code, change it, and activate it.
If I manage to crash it, that's fine... Sawfish isn't the outer process for my session. For that, I have my .Xsession just do a simple loop. It keeps running until I touch a file called ~/.xlogout, and can restart the window manager or other tools if necessary.
I pretty much never log out or reboot unless it's time to do a dist-upgrade... which means once every couple years on average. Everything just stays running.
The main limit is that Xorg is usually compiled to use an 8-bit ID for each client, so it can only have 255 clients connected at once, which usually means a limit of 255 windows open at a time. But you can push this a bit using programs which do multiple windows in a single client connection, like "urxvtcd" in rxvt-unicode can have hundreds of terminals sharing just one client slot. I've occasionally had 500+ windows open on a single computer (though this is kind of a bad habit, and I try to keep it under 200).
I configured Sawfish to do soft tiling, and also to do tabbed windowing, and made all of it work via keyboard, including things like moving and resizing windows. Sawfish is also where I implemented the "every flavor of F1-F12" hotkeys thing, and implemented mouse warp, and it already had a lot of things built in like sloppy focus and the ability to add and remove desktops with a keypress.
I added another thing to it to save the contents of every desktop, every window, size and position and title and tab grouping, etc... every 15 seconds or so if it has changed since last time. This gets archived daily. So if the whole system crashes, or if I need to actually log out, or if something bugs out, I can restore things mostly how they were.
A lot of the stuff I don't have, like a task bar, or a desktop notifications system, is simple enough to do by just not installing one or not running any of the relevant programs or services.
As for traditional-ish desktop stuff, there's not much. I set the wallpaper using a small program I made decades ago, which basically just picks a file at random from a big nested directory structure of images I've accumulated over the years. I use a 2D grid style desktop pager in the corner of the screen to see which desktop I'm on, with Hyper+Arrows keys to navigate the grid, but I barely ever look at the widget and it sticks to the bottom stacking layer so other programs are free to draw over it. Similar story for a Conky instance configured in a thin vertical strip with all sorts of realtime system stats and any other info I care to put there. It's a quick, easy way to see what's going on, but I often cover it up and use that space for other things. It's basically just system stats rendered onto the wallpaper.
I mostly use keyboards with open-source firmware like QMK, with a pretty extensive personal keymap tailored to my needs and tastes. I also added some stuff to QMK, like a better MouseKeys motion mode called "inertia" mode, to make mouseless use easier. But on notebooks and legacy keyboards, I can at least use Kanata to get most of the same features.
Macros recording/playback is built into the keyboard, or X11 in general is pretty easy to automate with small programs or shell scripts or even one-liners. Like, with xdotool.
For remote controlling other systems, x2x is great. Ancient program which "just works". Or running remote programs on the local screen, ssh + normal X11 stuff. Or for a full remote desktop, there's VNC (there's a tigervnc xorg extension, and the tigervnc viewer is good too... I patched it to make it handle larger desktops better, and the patches are upstreamed).
(continued in next comment, my post was too long to fit in one comment)
I have my dotfiles and other stuff checked into Subversion, with different programs in different modules so I can build each host's config out of a set of modules, and not have to have everything checked out on every host. There's an article about the setup on my site. Git didn't exist at the time, but you could probably do something similar with git submodules. I run a revision control server of sorts in a lxc guest on one of my computers, to make it easy to sync things. Or depending on the nature of what I'm syncing, sometimes "unison" works better, or sometimes "syncthing", or sometimes git + github, or sometimes just rsync.
I have a little python program constantly monitor the title of the focused window, and look for anything which looks like a filesystem path. This gets saved to a ramdisk, like ~/ram/last-dir and ~/ram/focused-window-title, so other programs can use this info for context-sensitive things. Hotkeys which react in different ways in different programs, or always know what directory to work in, or whatever.
I also basically keylog myself, with another little python program I wrote. It monitors the title of the focused window, logs the timestamp when each window was focused, and logs the quantity of input events... but not the actual individual content of keystrokes. So I can tell which window was used when and how active I was, but not exactly what I was typing.
To make these title-based things work better, I make sure to put useful info into my window titles. Zsh is configured to put the command line (so far) into the title, along with the current host and path, or put the finished command line into the title while a program is running. I added a browser extension to put the current URL into the window title. I configured Vim to put useful info into the title. Etc. So the logs are very detailed. This info comes in handy quite often, and is also used for other things like generating a heat map graph of my erratic sleep schedule, and predicting when I'll be awake so I can schedule appointments more easily. Can also easily tell exactly how much time I spent working on pretty much anything, for invoicing clients or working toward productivity goals or whatever.
I've put together several things which work in a daily log directory... like ~/daily/YYYY/MM/DD/ . There's a key to open a shell in today's directory, or open a minimalist filemanager there. I made a screenshot program which saves things there... like, press a key, click or drag to select what I want to save onscreen, then it shows a preview I can edit if I want, and asks for a title, with recent titles available for easy recall or searching or tab completion so I can save a series of images without having to retype the title... and then it saves to a timestamped file in today's dir. Shift+Screenshot to do the whole screen, Ctrl+Screenshot to turn off all controls and just save automatically using the most recent title, Meta+Screenshot if I want to add metadata (like open "hacker-news.2026-09-24_11:59:23.png.md" in Vim to save extra notes). Or any combination of modifiers to combine effects.
And, of course, a nice terminal and shell. I've found rxvt-unicode is lightweight and can run hundreds of terminals with near zero resource use. And zsh is very powerful and customizable. But YMMV, use whatever you like. Build up your own vocabulary of tools over time, and the shell becomes incredibly powerful.
I've had the same $HOME, basically, since I first installed Linux in the 1900s. It occasionally gets moved from one hard drive to another, and now is scattered across an entire network of many hosts, with pretty much everything replicated and synced and backed up... but it's still the same data, collected and curated and organized and optimized over decades. When one's computing environment is stable and built for user empowerment instead of corporate growth... one can really settle in and get everything "just right".
I've also had mostly the same browser session for like 15 years. Even when switching between browsers. The session is much like my $HOME filesystem, just an organized hierarchy of arbitrary stuff, whatever I found useful or notable or worth keeping. At first, I did this using Tabs Outliner in Chrome, but it has degraded over the years and has a lot of the best features paywalled, and it doesn't work on Firefox, and it didn't adapt well to Manifest V3, and it's buggy and unreliable and awkward now... so I made my own replacement. I suck at naming things though, so I just called it "TK's Tree Style Tab Outliner". Still working on the feature to auto-sync between browsers, but at least it can export and import manually and do automatic backups, to make sure data won't be lost.
... which reminds me. It's also very helpful to run Chrome (or Firefox, or whatever) in a memory jail cgroup so it won't grow until it squeezes every other program into swap space. A cgroup like this also makes it easy to pause and unpause the entire browser session so it won't do anything while you're not actively using it. Slow down the memory leaking, reduce CPU use to zero to make the rest of the system run faster and cooler and make the battery last longer, and completely stop data leaking out to corporations while the browser is idle. Yet another thing I made a little script for.
And a ton of other things which didn't even come to mind, because I'm so accustomed to it that I don't even think about it any more.
Most of this is niche and weird and probably has very few users, and took a lot of time to discover and integrate... but a lot of the stuff people want in a desktop is already possible and has been for a long time, if you're willing to look around and try things until you find what works for you. Just try to use systems which are designed to be powerful and flexible, and you'll never hit the ceiling. It may be harder initially, compared to something mainstream, but you'll be able to go a lot farther with it.
Thanks so much for this detailed write-up! It's always enlightening to learn about different setups, especially when they have been painstakingly crafted over multiple decades.
I know someone with aphantasia. It's absolutely wild to me that he can't see things in his head. Totally alien from my perspective, since I'm at the opposite end of the spectrum. I see what's in my head so vividly that I often won't see things in real life which are right in front of me.
But what's even more wild to me is the people leaving comments refusing to believe that different people are actually different. Refusing to believe aphantasia or hyperphantasia exists. Refusing to believe that other people are telling the truth when they describe their experiences. Insisting that everyone is the same, actually, and anyone who thinks differences are real must be confused or misunderstanding something.
I guess these are both types of cognitive disability... but I know what aphantasia is called. Not sure what the other thing is called, but I'd imagine it's something which sounds pretty unflattering.
im like you where i become so caught up in the image in my head that i lose track of my surroundings
when engaging in small talk i make eye contact and have been told im very easy to talk to
when discussing something im interested in, i stare into the distance and start waving my hands around gesturing at the diagrams im describing. its to the degree where my managers prefer i not meet with one specific customer who commented on this
at my first job my coworkers pointed out that i stare directly up to the ceiling for a period of time, then will type furiously, and repeat. id never known that was different before. assumed everyone constructed their ideas mentally before getting it down on paper
im reminded of my high school AP physics teacher who would close her eyes while lecturing because i assume she was picturing things. the funniest students knew how to get some physical comedy into those windows of obliviousness
when it comes to aphantasia it reminds me of Stan in South Park: "I get it now! I don't get it!" the best i can do be nice and accommodating and acknowledge i simply cant relate
I believe some people have aphantasia and hyperphantasia, but a lot of people who claim they have either are somewhere on the bell curve between them.
I also suspect that some of the people who think they have aphantasia could learn. I personally am not good at visualising and don't naturally do it, but these days I can imagine a not very good apple if I try hard, and I trained this by building on my existing ability to imagine the tactile feel of holding an apple and on spatial awareness.
In fact, the term typical mind fallacy was born from aphantasia [1]:
> Galton gave people some very detailed surveys, and found that some people did have mental imagery and others didn't. [...] Though another psychologist, Dr. Berman, did, at that time, come up with a name for the masking behavior, called the Typical Mind Fallacy: the human tendency to believe that one's own mental structure can be generalized to apply to everyone else's.
I'm one of the naysayers denying aphantasia in the comments.
Why would we default to believing people, though? I get your point. Believe people when they tell you their experience. Sure, that seems like a reasonable thing to do. But in this case, when there is zero evidence besides a self-reported experience, why do we default to believing that there are suddenly two types of people, instead of assuming that semantics has created this illusion?
You're defaulting to being nice, which is fine, but it doesn't make you any more right than me and my skepticism. We both have no idea what the truth is.
That's a bold claim. It doesn't tell me there is zero evidence... hardly anything has zero evidence. It just tells me you either reject the evidence or haven't looked. It's not hard to check. For the lazy, Wikipedia provides a nice summary and 82 references to get more details. But that's just a starting point.
> there are suddenly two types of people
This comment is pretty revealing. For almost any trait, there are way more than two types of people. Usually there's a whole spectrum with lots of diversity. It's a beautiful thing. What would be very unusual is if there was only one type of person, or only two types.
Even distinct "types" is a bit of a misleading concept. Instead of distinct types or buckets, most things have more of a cluster or nebula structure. Each individual member of a cluster is different, but has overall similarities, on average, to other members. There will usually be a fairly dense core, and then it gets more sparse with distance. And there will often be several clusters near each other, overlapping, so there is no clear boundary between one cluster and another. That is how descriptive categories work... nebulous clusters in thingspace, where that space is a many-dimensional abstract space with each trait having its own axis, and many degrees of each trait.
For convenience and efficiency, we use labels to refer to these clusters, but when deeper understanding is needed, or more nuanced communication, the labels stop being a useful simplification. It's like trying to dig out a splinter with a shovel; shovels are very useful, but they're way too coarse and blunt for something as precise as a splinter. A deep enough understanding requires us to more explicitly recognize the cluster structure and the multidimensional spectrum of diversity.
Don't mistake the map (the labels, or the "types" of people) for the territory (the rich variety of life).
> why do we default to believing that there are suddenly two types of people
Because we know that there are already more than two types of people, even when we ignore aphantasia? So there is nothing new here.
Also if you read the FA, it explains that there are objective differences:
> People have taken various physiological and behavioral tests in a research lab. Researchers have observed that people who report vivid imagery respond differently to these more objective measures than aphantasics. Although this doesn't let us know definitively that they vividly see an image, it all points to the fact that something is truly different between people who claim to see vividly and those who don't.
Here’s some evidence for you: when I close my eye, I see black. That’s it. No amount of effort conjures an image. That’s not what most people I talk to about it report, so I am going to assume there’s a functional difference between me and people with functioning visualization.
I believe you. Why wouldn't someone believe you? There's not much reason to claim "I cannot do X" maliciously. I'm hopeless at catching a ball, can't play tennis to save my life, for instance... believe me?
I just want to throw my two cents in: I think I have a pretty good level of visualization ability, but it's not actually projected into my field of vision. It's like having a second monitor plugged in somewhere... _else_. It doesn't overlap with the 3D space of our world. I don't SEE it, but I do "see" it, not as concepts but as something akin to generated memories.
It does often help to look away from visually distracting things so that the mind is more free to use its capacity to generate things.
I have the same. I'm looking at my laptop's screen as I type this, and if I close my eyes while typing I cannot even visualise said screen. Now, I can blind-type, recognise faces, navigate aso without any problem. I can even play chess quite well (fide ~2050, 2200 blitz on chess.com), but I cannot visualise a chess board, nor can I visualise my moves or the moves of my opponent.
However, I don't "see" (pun intended) it as a problem. On the contrary, I'm not wasting cycles on irrelevant details while I'm thinking.
If it makes you feel any better, I'm on the hyperphantasia end of the spectrum and it is honestly actively problematic sometimes. For me, it can be very difficult to work with things that I can't translate to visuals, mentally; and daydreaming, and memories can be so vivid they are distracting or disturbing. It's very much an asset in certain circumstances (like debugging code or fixing an engine), but there's a reason most people don't have that fidelity.
>Here’s some evidence for you: when I close my eye, I see black.
This, I can imagine things, but it's like the idea of something, not a picture of it. Blows my mind when people tell me their imagination is like watching a movie.
Same! My wife is at the opposite end and can visualize these wild scenes and describe them in detail and I’m like “I can tell you roughly where the parts of a cat go”.
That's so funny, my wife and I have the opposite dynamic. She rolls her eyes at me when I insist "no the couch would definitely fit through that doorway if we rotate it like that, I don't need to measure", and after she gets a tape measure it's about an inch of room to spare.
One, a study with more than 32 people. Two, replications with equally strong observations.
I've read a lot of research on fMRIs and mental disorders. Something like this pops up every other day. Not a single one has really study the test of time ion terms of diagnostic utility.
I meant definitely proven with in a realistic sense. Smoke causes lung cancer, albeit only something like 20% of smokers get lung cancer. It's been proven, but not in the mathematical sense of the word.
It was funny to watch this discussion turn from skepticism at the top to then a bunch of proponents moving in and crowding out the rational voices. It feels a bit religious to me, at least unscientific, there are obviously no testable hypotheses about subjective descriptions, and so people seem to dig in on their sides more than if there was a real basis for comparison. I default to not believing in it until there is actual strong evidence, same as ghosts and the Loch Ness monster and whatnot.
I'm not sure how the examples given in this thread, such as brain scans, are "unscientific" while simply saying "nah, I believe everybody else is exactly like me" is "rational voice"?
> This brings so many advantages, such as true tear-free graphics and lower input latency.
Curiously, people always seem to list the same advantages, and it's a very short list. Fewer (but still non-zero) torn frames are the top of the list pretty much every time, since that was the very first thing its creator listed in his original goals... but that's a pretty small benefit in exchange for breaking entire categories of functionality. Like, it looks smoother when I scroll, but features I rely on heavily every day are forbidden.
As for input latency... that doesn't seem like it was ever a problem. Using X11, I'm able to get 500 to 1000 inputs per second even on a potato PC. That's faster than the frame rate of pretty much any screen, and fast enough even for audio / midi purposes. Reducing input latency from ~1.5ms to ~1.0ms doesn't really matter when the timeslice scheduling has ~6ms of jitter on an average system, a common screen only draws a frame every ~16ms, and many input devices have 50+ ms of their own additional latency.
> if you want to take a screenshot in a generic way, you have to go through the XDG portal API
This is a symptom of what's wrong with Wayland. Wayland doesn't do screenshots. Like many things users need, they decided it was someone else's problem, and threw it over the fence for the fragmented ecosystem of downstream projects to solve. So each downstream project came up with their own workarounds for essential features not existing. The solutions had to be built entirely outside of Wayland, and even after years of development, the solutions are still incomplete, unreliable, complex, and full of caveats. It required (and still requires) everyone except Wayland's core devs to write a lot more code for less functionality than they used to get with a couple of simple API calls. They had to architect entire complex infrastructure layers in order to work around a missing feature in the core protocol, since the core devs stubbornly refused to allow it.
Similar situation for input automation and remote control. It's a common thing people need. I use it every day and can't use the notebook at my desk without it... but the Wayland folks refused to solve it, so it had to be done outside of Wayland. For example, one workaround is to give the user direct access to the kernel so they can create fake input devices at a kernel level, and generate the inputs they need, which Wayland then sees as a local physical keyboard or mouse. So... problem solved, from Wayland's point of view. The user gets what they need, sort of, and it's implemented outside of Wayland, so Wayland doesn't have any security issues. But... and this is a big "but"... the solution involves giving users device-level kernel access. Which seems significantly worse than the issue it was originally trying to solve.
> I understand where you're coming from, but there is no turning back at this point.
A position of "sure it has major problems, but it's too late now" is not a position of progress. Much like the situation with pulseaudio being deployed everywhere then replaced with pipewire, it's almost never too late to fix bad software architecture. As you said, this whole ecosystem is still very much a moving target.
The ideal solution would be the creation of a new system which supports the features, protocols, and APIs of the older system(s), in a way which "just works". But that requires a very different mindset from the people behind it. Instead of "not my problem, someone else can deal with it" like the Wayland policy, a proper solution needs people to adopt a "the buck stops here" approach, and take responsibility for making the entire system work. Things like accessibility, network transparency, automation, and legacy support... need to be built in from the ground up, not rejected or treated as an afterthought for someone else to handle.
> Curiously, people always seem to list the same advantages, and it's a very short list.
The things I listed are just the tip of the iceberg. Having studied earlier versions of X like X10, X11, and now Wayland, I could go on and on. What I'm trying to emphasize here is not to look at Wayland through an Xorg lens. I know it's easy to look at Wayland and immediately think it's a rigid protocol that breaks everything, but the reality is that it's a lot more flexible than people realize. Once you understand it, I think the closest thing that comes to mind is "There is no spoon." And just like there is no spoon, there are also no windows, only surfaces, but again, that's only the tip of the iceberg.
If you want an example of what I'm talking about, take a look at projects like p9wl. [0] This is using wlroots to create a proxy between Linux and Plan 9 in order to display windows remotely. Such flexibility is only possible when the protocols are small and composable, which Wayland is.
> As for input latency... that doesn't seem like it was ever a problem.
I never said it's impossible to get acceptable input latency on Xorg. I'm just saying that when your input and graphics go directly through your compositor, the experience is on a completely different level. This is even more true with features like direct scanout, where the compositor steps aside to send graphics directly to the GPU and input directly to the game as it goes fullscreen. It simply makes using your computer so much more pleasant. These input latency issues were exactly what Kristian Høgsberg tried to fix in Xorg, and addressing them was one of the primary motivations behind creating Wayland. [1]
> This is a symptom of what's wrong with Wayland. Wayland doesn't do screenshots.
You're thinking through your Xorg lens again. That's understandable, since you're probably used to it. However, to really understand Wayland, you have to look at it through a different lens. Wayland isn't defective, and it's not trying to be Xorg 2. It's a protocol. It doesn't do your laundry, nor does it take screenshots. It's just there to provide the mechanism for sending buffers to your compositor.
You can think of Wayland as thin building blocks to build compositors, just as you can build window managers on X, except this time, the core protocol does one thing and the libraries already exist. Wayland exists because all the infrastructure is already there in the kernel and in libraries.
The screenshot feature is the job of the compositor. Most compositors already have this feature built in. As I mentioned before, XDG portals exist, and interoperability will only keep getting better.
> Similar situation for input automation and remote control.
I agree input automation and remote desktop are still problematic. GNOME and KDE have their own remote desktop solutions, as do most of the others, but I haven't had a need to use them myself. All in all, Wayland really isn't trying to prevent you from doing things, it's just a protocol.
I'm pretty sure all this will be solved eventually, and I understand the frustration of those who need these features. However, I also understand Wayland and compositor developers taking a careful approach here, and it shows.
> A position of "sure it has major problems, but it's too late now" is not a position of progress.
That's not what I intended to say. The intended message was: the core Wayland protocol is set in stone and won't change, but we can build these things that aren't working and make them work. That is already happening every day. Progress has been steady, it's a lot better than it used to be, and it will only keep getting better. Going from Wayland into a whole different windowing system, I don't see that happening, at least not in the next 60+ years, but I definitely see the whole Wayland ecosystem improving.
So I keep hearing. But hardly anyone ever talks about what's underneath the tip... and when they do, it turns out to be a big mess under there, not a mountain of advantages.
> just like there is no spoon, there are also no windows, only surfaces
Yes, I understand. That's why programs can't position their own windows, or get or set the mouse position. It doesn't even assume the coordinates are 2-dimensional. It's all very futuristic and makes a lot of interesting things possible, like using Minecraft as a compositor, or a Facebook style Metaverse compositor. But you're not hearing what I'm saying:
Wayland is solving the wrong problems. What it can do is pretty neat, but what it does is not what people need. It's good at party tricks, but bad at solving the real needs of daily life. It largely solves problems that nobody had, and to do this, it breaks major features people need. Much like how Facebook's Metaverse failed, nobody really wants to do their spreadsheets in Minecraft.
> It doesn't do your laundry
Yeah, and that's part of the problem. To use your analogy, it's like removing someone's washer and dryer, and replacing them with a box of tools which don't do laundry. It may be a really fancy foundation for building efficient laundry appliances, but that's not what people need. People may have used it to build an entire ecosystem of laundry appliances, but they couldn't agree on basic things so the ecosystem became fragmented, such that P-brand soap doesn't work in Q-brand washers, and clothes washed in Q-brand washers can't be dried in R-brand dryers, and if you use a P-brand washer you need to take fifteen extra steps to prepare it for a drying cycle, and T-brand dryers can only be used for pants, and ... etc etc. It's a big complex mess which, after 18 years of development, still doesn't let me get my socks clean. And most of it doesn't work at all for people who use a wheelchair.
> when your input and graphics go directly through your compositor ... input directly to the game
When they cut out the middleware, the middleware stops working. And when they make it impossible to add middleware, it stops being possible to do all sorts of useful things. Sure, it may be a millisecond faster, but in exchange, all my keys get mapped wrong, or I can't use my notebook without physically touching it (and causing repetitive strain injury), or my autoclicker stops working, or my accessibility tools become impossible to fix, or my automatic time tracker is treated as a security violation, or the solution which works for me suddenly won't work for my friend who uses a different compositor, or ... etc etc etc.
It architecturally eliminates entire categories of useful things... in order to make the simplest case slightly faster. This provides a more pleasant experience for average able-bodied normal people with no uncommon needs or preferences, while making things worse for everyone else. In particular, the complete lack of accessibility in Wayland is very able-ist and makes it unusuable for anyone with disabilities. That sort of thing needs to be built deep into the core, but it was instead rejected outright and left as an afterthought exercise for Someone Else to solve.
> [this] feature is the job of the compositor. Most compositors already have this feature ... GNOME and KDE have their own solutions, as do most of the others
This really gets at the nature of Wayland's biggest problem. A ton of important things are rejected and declared to be Someone Else's Problem. The compositors have attempted to deal with the aftermath of this mistake, but they all do their own thing and refuse to agree on a lot, so instead of one robust solution, we end up with an entire fragmented ecosystem of partial solutions which aren't compatible with each other.
It's even more of a nightmare for application developers. Instead of needing to support Windows, MacOS, and X11, suddenly they need to support Windows, MacOS, X11, GNOME, KDE, Sway, Weston, Hyprland, Enlightenment, Niri, etc. Every time two compositors disagree on something, it means every application must add support for both of the ways of doing it, like how they are now required to optionally draw their own title bar and window frame, depending on which compositor the user has... and good luck getting two programs from two different developers to draw their frames in the same visual style.
> Wayland really isn't trying to prevent you from doing things
I agree with your pain points. But the reality is that GNOME is already Wayland-only, KDE will be dropping the X11 session in 6.8, and GTK has deprecated X11 support, so going back isn't realistic.
> so the ecosystem became fragmented
The current fragmentation is frustrating, but Linux and Unix have gone through this exact phase before. Unix and especially Linux have never been static, they are more like living organisms. Back in the late '80s and early '90s, PC Unix suffered from the same growing pains. Multiple commercial and open-source X servers popped up for different hardware, until things eventually settled around X386 and XFree86. Even X developers in 1987 were saying "old window systems don't die" [0], sounds familiar?
We'll likely see compositors come and go, along with new protocols to address today's pain points. But the core protocol is here to stay, and wider adoption will naturally push the ecosystem toward shared solutions, just like we're seeing with River and wlroots. The fragmentation will fade over time.
In the meantime, anyone who still relies on X11 can simply stay on it until Wayland covers their use cases. Standalone X11 window managers aren't going away anytime soon.
Rather than dwelling on the friction, it's probably best to focus on what you can control: test software, file good bug reports, and help resolve the remaining edge cases.
I sympathize with the headache, but many Wayland developers came straight from X.org and know what they're doing. It's an uncomfortable adjustment period, but we'll get through it.
Metux (Enrico Weigelt) was banned from the Xorg project because his patches kept breaking things, often in ways which demonstrated he didn't even do the bare minimum before pushing commits. Like, adding code which didn't even compile. It's fascinating to look through the huge pile of commits they had to revert after banning him, to see how bad a lot of it is. I picked a few at random, and ... wow.
So he started his own fork.
The ban didn't seem to have any relation to politics or personal behavior. However, he was widely known for being ... how to put it in a way which is acceptable here ... uh, difficult to deal with. Like, after Linus Torvalds made a vow to be nicer, he made an exception just one time... for metux. He was the only person obnoxious enough to get Linus to break his vow. Which, if I recall correctly, was how he ended up focusing on the Xorg project. After getting booted from Linux, he picked a different project.
That's who is in charge of XLibre.
Things didn't get political until he announced the fork. Because, although politics had nothing to do with him getting banned (either time), he framed it as if he was being targeted for political reasons, and used inflammatory political language in the project's documentation. This gained a lot of attention and caused a lot of controversy, and also had the effect of ensuring the contributors all had mostly the same political views.
Anyway, the number of commits doesn't tell much of the story. The content of the commits (especially the reverted ones at Xorg) are far more enlightening.
Exactly. It's hard to be enthusiastic about something which breaks a lot of stuff I rely on daily, when the benefit is that it fixes problems I've never had.
My theory about the motivation behind the extra security is that it's largely driven by corporations wanting to make desktop Linux less free and less open, and normalize proprietary software instead of open-source. Because profit. Proprietary software is inherently not trustworthy, so the execution environment needs extra security and restrictions, and must be generally less powerful to reduce the damage it can do. Essentially, proprietary software needs the same precautions as malware. So the corps needed to "androidify" desktop Linux. Hence the change from curated distro package repositories to corporate app stores, and the reason why there's so much money pushing to replace X11 with Wayland.
In an open-source ecosystem, users and developers are one and the same, or at least on the same "side", cooperating with each other to make tools which work as well as possible for everyone. Each big program tends to be a collaborative effort where a lot of people contribute to make things better for everyone. Things mostly "just work" and people can typically trust their computers not to do anything weird or hostile.
Very different than a proprietary commercial ecosystem, where users and developers have more of an adversarial relationship. Each program tends to be created in a closed silo by one person or a relatively small team, and is designed primarily to extract money from users, with all other concerns being secondary. It is very common for profit-driven developers to engage in deceptive practices, or do things the user doesn't want, like showing advertisements, collecting and selling data, sabotaging products from competitors, using the device as a node in a botnet or secret compute farm, forcing unwanted updates, microtransactions or subscription fees, etc. So nothing can be trusted, and the entire system needs extensive protections against every type of misbehavior imaginable... even if that means reducing the power and features available to the user.
I've really enjoyed the past few decades of living entirely in the open-source world, where those problems pretty much just don't exist. But with corps pushing the androidification of desktop Linux, I fear those days may be coming to an end.
Wayland security is a side effect of how the core protocol works. In X, you have one giant buffer, while in Wayland, each application renders to its own buffer and pushes that buffer to the compositor when ready.
X has had compositing for ages as well. That didn't require preventing applications from moving the cursor or simulating keyboard input and even screen grabs still work fine.
But that's okay. I don't have to use one. Different people have different needs and preferences, and their choice to use a taskbar-free workflow doesn't interfere with their ability to get serious work done.
reply