| Age | Commit message (Collapse) | Author |
|
|
|
|
|
|
|
|
|
+ Found it in newlib. This allows the compiler to more effectively
optimize away unnecessary register operations. Applied to both kernel
and init.c for a slight speed increase.
I really should move init.c to some other repository though, the
binary files are starting to get annoying
|
|
+ Tests will have to come later
|
|
|
|
|
|
+ User can now specify debug, release and assert handling
|
|
|
|
+ Try to write arguments to memory as fast as possible to free up some
registers that the compiler can then play with.
Qemu runs ~1 700 000 rpc's per second, visionfive2 @1GHz (I think) ~800 000.
|
|
|
|
|
|
+ Make 8250 serial driver more generic
+ Still TODO: write a tutorial on how to boot on the visionfive2
|
|
|
|
+ Not bootable quite yet. Among other things, I couldn't
get the current starfive u-boot fork to boot, so
try adding support for booting with precompiled u-boot
via the `go` command. Initial testing with qemu shows that
this should be possible, and if it works, might be useful
in the (far) future with other slightly janky SBCs.
Also, NS16550 is 8250-based, and I'm really only using the base
8250, so rename and add visionfive 2 uart to list of compatibles.
Visionfive 2 is still completely untested.
|
|
+ CSR_STVEC requires four byte alignment
|
|
|
|
+ RELEASE=1 doesn't work for some reason, though
|
|
+ Both kind of go hand in hand, made sense to do both at the same time.
Some parts feel slightly hacky, the loader works by placing everything
on the stack and avoiding global values. Still, seems to work?
|
|
|
|
|
|
|
|
|
|
|
|
+ Now 10M alloc/frees succeed, which totals more memory than the virtual
machine has, so no too obvious leaks are occuring. For future
debugging speed, changed 10M to 1M.
+ Also quick fix to rpcs, stacks are now assigned. Not entirely sure why
they worked before this, but good that I found it.
|
|
|
|
|
|
|
|
|
|
+ The secret is using gravestones. I'll have to write up full
documentation for the feature but essentially riscv lets us encode
whatever we want into page table entries, as long as they're not
active. We use this to encode highest user address that is not a zero,
in that all entries in the top level are either active or gravestones.
When an active entry is removed, it is either a gravestone (if there
are other active entries above it) or it starts a cascade of removing
entries that have been previously removed
Slight runtime overhead to page mapping, pretty major advantage in rpc
calls. Feature will need to be tested more thorougly, and the init
program is sort of a best scenario with just one top level userspace
page table entry active at a time, leading to incredibly fast context
switches.
Current implementation limits a process' max virtual memory to 248 GiB
(in Sv39), but I don't think the missing 8 GiB is that big of a deal.
|
|
+ Eliminated need for save/load_regs, now the register save area is
behind a pointer which _slightly_ increases overall syscall delay, but
speeds up ipc requests by a fair bit.
|
|
+ It really only works with Sv39, I suppose similar structures should be
added for Sv4} etc. if I ever get around to it.
|
|
+ Both easier to look at and now the userspace doesn't play as big of a
role in the 'benchmarking' with optimizations turned on.
|
|
|
|
|
|
|
|
|
|
+ Not strictly done yet, but we can swap between two processes :D
|
|
+ Not actually working (at least fully), need to continue debugging when
I have time.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
+ Slightly unsure if it should be a syscall, or if assign would be
enough. Also, which thread should run in a fork/spawn? For now, old
thread, though this might increase latency. Or add swap flag to these
calls?
|
|
|
|
|
|
+ In total, a syscall is built up of 6 values, with the first being the
syscall number. Symmetrically, the first value is now a status and the
following five values return "values". This allows us to cram in more
info into the ipc_* functions. The performance difference is
absolutely minimal, at least from my testing in qemu.
|