C++ is getting std::simd in the latest version and I am all aboard writing the vectorization with the least amount of intrinsic builtins I am able to. Even if not optimal, it's far better than the scalar ops.
This is why I love Go. Nobody was asking for this, but they took the time to do it right and continue to Push go as a memory safe, high-level systems language.
There is no memory safety without freedom from data races. One is a prerequisite of the other. This is why languages like C# throw exceptions on unsynchronized concurrent access to some container types, and treat all property accesses as atomic.
I suppose that if you create a map in one thread, and then access it from another thread, then that might cause segmentation faults? Because map is a type that is implemented in C.
(The gist: memory safety is a term of art coined by security practitioners. Go, Python, Rust, Java, others: memory safe. C/C++: memory unsafe. Periodically, people in different slices of industry or academia come up with new definitions of memory safety that declare Rust or Go or other languages to be memory unsafe, but that is not by the broadly accepted definition across industry.)
You can write unsafe code in Go (import unsafe), but then, you can do the same in Rust. Unsafe code is not the default, and in day to day Go i rarely see the use of the unsafe package.
What he probably means is data-races in go can result in memory/type unsafe accesses -- I suspect, likely due to slice types -- not sure if that is true/false.
Sure, but a data race is, IMHO not the same as memory safety. A data race, can be 100% memory safe, but just cause a logic bug in some program. I often see people mixing memory safety with racing. Go has bounds checks so you end up with a panic either way. Not UB.
As an (outside go) example, Ocaml (5) promises strong memory safety, but not to be data race free. A data race is not something we can prevent, because its usually not bound by code, but by time and the race-source rarely in source-code.
This means we have data races in http, database inserts etc. The source is usually not a concurrent task in source code-land.
This is not generally considered part of the "memory safety" contract. You can not lift a nil pointer exception into a replacement for Go's "unsafe" library.
When we finally rid ourselves of C and C++ is so larded over with extensions and additions and features that we can finally plausibly say the C subset is just not in use anymore, we can perhaps consider as a community expanding what "memory safe" means, but in the meantime it has some very important meanings and we should not try to augment the term. Memory safety doesn't mean anything like "forcing exhaustiveness into sum type deconstructions" or "never has a race condition" (though it does mean said race condition shouldn't be something that allows you to escape out of an array or forcibly change the type on something in a way the language doesn't normally permit) or any of several other things that may be very nice to have indeed, but are not part of the definition of "memory safe".
Memory safe is a very old concept, and almost everything is memory safe now. But not quite, and as such the term still has use. And also zig for some reason gave it up so it won't be disappearing as soon as I'd like.'
This feature opens many doors for optimizing low-level performance in Go projects, that are already running multicore. IIRC there aren’t a lot of languages with built-in std lib support for SIMD and variants. Love the way Go is trying new stuff lately.
Besides the usual C and C++, we have Java, .NET, D, Zig, Julia, Swift, Rust.
So yeah, also appreciate having Go in the group instead of manually having to write Assembly.
However not many languages adopt ways to manually write SIMD, because most of us have no idea how to write good SIMD code in first place, I surely don't.
Even with languages that adopt ways to manually write SIMD, it’s mostly left to library maintainers rather than application developers.
I work for a C++ timeseries database startup that leverages SIMD about as much as we possibly can, and except for some extremely rare places we just use libraries.
Yeah, that is what I have heard from some NVidia folks as well, like Bryce Adelstein, use the libraries as much as possible, and leave the kernels for experts.
However even then, it depends on how the libraries API surface looks like.
But it’s not necessary at all, the whole point is that these utility libraries bring you more elegant code that work on all platforms without having to pollute your codebase with SIMD intrinsics.
Unless this was tongue in cheek, because this is in fact a problem with AI that it degrades your codebase in these types of ways.
> The interface conversion and type switch look like they should be inefficient, but the compiler-side implementation of simd specializes code and optimizes away the type switch.
I don’t understand this - how is it able to if the same go binary might run on unknown types? I’m assuming what it means is that the switch is implemented efficiently due to CPU branch prediction? I know fearless SIMD is doing cool stuff with static dispatch so that the feature set is checked just once at program start - is that what it means it’s doing under the hood? Very unclear.
It creates multiple versions of functions referencing SIMD and lifts the dispatch switching cost to their callers.
> The AST rewrite creates multiple specialized copies of functions, variables, and types that mention simd types, where simd types are replaced with references to size-specialized types in simd/internal/bridge. Each of these bridge types is defined as an archsimd type, but with a restricted set of methods. The specialized functions, variables, and types acquire a suffix of the form @simdNNN, where NNN is either a vector length (128, 256, or 512) or 0, indicating emulation. Functions that mention simd internally, but not in their signature, are converted to wrappers that switch on the SIMD level detected at program start, and call the appropriate specialized version of that function. Specialized functions call other specialized functions directly without dispatch overhead (and perhaps with inlining). This rewrite strategy was chosen as a compromise between code duplication and SIMD performance; the overhead is hoisted as high as necessary to avoid dispatch within SIMD computations, but not higher. If SIMD dispatch appears “too low” in a computation, a gratuitous mention of a simd type will move it upwards, as in this example:
The problem with Go isn't performance but with the C/C++ interop overhead, even with the "30% less overhead" from a few updates ago which isnt true for 99% of cases, it isnt enough
Use Assembly instead of CGO, isn't that scary, back in the 8 bit days we were coding Assembly aged 10, on our Spectrum, C64, Atari, Apple, Acorn, MSX,....
its not limited but it has overhead because of the memory model of go doesnt match the C one so there has to be some sort of rerodering being done, that's what i understood atleast, and theres also the go concurrency
It's hard to make predictions with an open source project, but my personal guess is some flavor of it will land (including it is already demonstrating good results without an enormous level of code complexity in the compiler and without overly slowing down compile speeds), but I guess we'll see.
It's being driven by an external contributor who has landed some good changes in the past to the Go compiler. (I think the autovectorization work might be part of their PhD or other academic research, but not sure.)
As a first step, it might be possible to write a linter rule that rewrites suitable numeric loops to SIMD. There are already rules to rewrite several loop types, so that should be doable.
Already using this for foreground estimation of cutouts in my project, around 30% speedup over non-SIMD, but the algorithm is probably not very optimised yet.
Portable SIMD is ~11% slower than non-portable SIMD in this case, but both are ~5x faster than non-SIMD.
So whats your point here? Haskell?
See for example comments from tptacek like:
https://news.ycombinator.com/item?id=43335748
https://news.ycombinator.com/item?id=46028232
https://news.ycombinator.com/item?id=44672371
(The gist: memory safety is a term of art coined by security practitioners. Go, Python, Rust, Java, others: memory safe. C/C++: memory unsafe. Periodically, people in different slices of industry or academia come up with new definitions of memory safety that declare Rust or Go or other languages to be memory unsafe, but that is not by the broadly accepted definition across industry.)
At some point in the future, a fully backwards compatible Rust compiler will report an error when you try to compile cve-rs.
As an (outside go) example, Ocaml (5) promises strong memory safety, but not to be data race free. A data race is not something we can prevent, because its usually not bound by code, but by time and the race-source rarely in source-code.
This means we have data races in http, database inserts etc. The source is usually not a concurrent task in source code-land.
That throws a NullReferenceException
When we finally rid ourselves of C and C++ is so larded over with extensions and additions and features that we can finally plausibly say the C subset is just not in use anymore, we can perhaps consider as a community expanding what "memory safe" means, but in the meantime it has some very important meanings and we should not try to augment the term. Memory safety doesn't mean anything like "forcing exhaustiveness into sum type deconstructions" or "never has a race condition" (though it does mean said race condition shouldn't be something that allows you to escape out of an array or forcibly change the type on something in a way the language doesn't normally permit) or any of several other things that may be very nice to have indeed, but are not part of the definition of "memory safe".
Memory safe is a very old concept, and almost everything is memory safe now. But not quite, and as such the term still has use. And also zig for some reason gave it up so it won't be disappearing as soon as I'd like.'
So yeah, also appreciate having Go in the group instead of manually having to write Assembly.
However not many languages adopt ways to manually write SIMD, because most of us have no idea how to write good SIMD code in first place, I surely don't.
I work for a C++ timeseries database startup that leverages SIMD about as much as we possibly can, and except for some extremely rare places we just use libraries.
However even then, it depends on how the libraries API surface looks like.
Unless this was tongue in cheek, because this is in fact a problem with AI that it degrades your codebase in these types of ways.
"CGO 2022 Keynote: Compiler 2.0"
https://www.youtube.com/watch?v=w_sX9aZoZxg
I'm grateful that Go a non-proprietary language offers these features.
Personally I will implement it in https://github.com/viggy28/streambed
This naturally is an easier way.
I don’t understand this - how is it able to if the same go binary might run on unknown types? I’m assuming what it means is that the switch is implemented efficiently due to CPU branch prediction? I know fearless SIMD is doing cool stuff with static dispatch so that the feature set is checked just once at program start - is that what it means it’s doing under the hood? Very unclear.
> The AST rewrite creates multiple specialized copies of functions, variables, and types that mention simd types, where simd types are replaced with references to size-specialized types in simd/internal/bridge. Each of these bridge types is defined as an archsimd type, but with a restricted set of methods. The specialized functions, variables, and types acquire a suffix of the form @simdNNN, where NNN is either a vector length (128, 256, or 512) or 0, indicating emulation. Functions that mention simd internally, but not in their signature, are converted to wrappers that switch on the SIMD level detected at program start, and call the appropriate specialized version of that function. Specialized functions call other specialized functions directly without dispatch overhead (and perhaps with inlining). This rewrite strategy was chosen as a compromise between code duplication and SIMD performance; the overhead is hoisted as high as necessary to avoid dispatch within SIMD computations, but not higher. If SIMD dispatch appears “too low” in a computation, a gratuitous mention of a simd type will move it upwards, as in this example:
The one negative I'd say is that often autovectorisation is 'good enough' and this doesn't really tackle that gap.
There's a CL stack here:
https://go.dev/cl/791740
It's hard to make predictions with an open source project, but my personal guess is some flavor of it will land (including it is already demonstrating good results without an enormous level of code complexity in the compiler and without overly slowing down compile speeds), but I guess we'll see.
It's being driven by an external contributor who has landed some good changes in the past to the Go compiler. (I think the autovectorization work might be part of their PhD or other academic research, but not sure.)
While reaching out to CGO is the easier way, it doesn't mean it is the only tool available in Go.
You'd think these people would know the meaning of API, no?