Systemd

Author	SHA1	Message	Date
Lennart Poettering	bbb4e7f39f	core: hide /run/credentials whenever namespacing is requested Ideally we would like to hide all other service's credentials for all services. That would imply for us to enable mount namespacing for all services, which is something we cannot do, both due to compatibility with the status quo ante, and because a number of services legitimately should be able to install mounts in the host hierarchy. Hence we do the second best thing, we hide the credentials automatically for all services that opt into mount namespacing otherwise. This is quite different from other mount sandboxing options: usually you have to explicitly opt into each. However, given that the credentials logic is a brand new concept we invented right here and now, and particularly security sensitive it's OK to reverse this, and by default hide credentials whenever we can (i.e. whenever mount namespacing is otherwise opt-ed in to). Long story short: if you want to hide other service's credentials, the most basic options is to just turn on PrivateMounts= and there you go, they should all be gone.	2020-08-25 19:45:38 +02:00
Lennart Poettering	bb0c0d6f29	core: add credentials logic Fixes: #15778 #16060	2020-08-25 19:45:35 +02:00
Lennart Poettering	45374f6503	Merge pull request #15662 from Werkov/fix-cgroup-disable Fix unsetting cgroup restrictions	2020-08-25 17:36:07 +02:00
Lennart Poettering	f053c9477b	core: drop redundant comment Since `625a164069` we don't need to update analyze-condition.c separately anymore, hence drop the comment suggesting otherwise.	2020-08-25 07:47:50 +02:00
Lennart Poettering	4e39995371	core: introduce ProtectProc= and ProcSubset= to expose hidepid= and subset= procfs mount options Kernel 5.8 gained a hidepid= implementation that is truly per procfs, which allows us to mount a distinct once into every unit, with individual hidepid= settings. Let's expose this via two new settings: ProtectProc= (wrapping hidpid=) and ProcSubset= (wrapping subset=). Replaces: #11670	2020-08-24 20:11:02 +02:00
Lennart Poettering	df6b900a1b	namespace: assert() first, use second	2020-08-24 20:10:58 +02:00
Lennart Poettering	52b3d6523f	namespace: move protect_{home\|system} into NamespaceInfo it's not entirely clear what shall be passed via parameter and what via struct, but these two definitely fit well with the other protect_xyz fields, hence let's move them over. We probably should move a lot more more fields into the structure actuall (most? all even?).	2020-08-24 20:10:30 +02:00
Lennart Poettering	9aab8d7a98	Merge pull request #16804 from keszybz/conditionals-and-spelling-fixes Conditionals and spelling fixes	2020-08-21 13:36:30 +02:00
Zbigniew Jędrzejewski-Szmek	3fb01017ee	Merge pull request #16686 from bluca/mount_images_opts core: add mount options support for MountImages	2020-08-21 10:11:08 +02:00
Zbigniew Jędrzejewski-Szmek	990307c3da	Merge pull request #16803 from poettering/analyze-condition-rework support missing conditions/asserts everywhere	2020-08-20 18:18:13 +02:00
Zbigniew Jędrzejewski-Szmek	2aed63f427	tree-wide: fix spelling of "fallback" Similarly to "setup" vs. "set up", "fallback" is a noun, and "fall back" is the verb. (This is pretty clear when we construct a sentence in the present continous: "we are falling back" not "we are fallbacking").	2020-08-20 17:45:32 +02:00
Lennart Poettering	5b14956385	Merge pull request #16543 from poettering/nspawn-run-host nspawn: /run/host/ tweaks	2020-08-20 16:20:05 +02:00
Luca Boccassi	427353f668	core: add mount options support for MountImages Follow the same model established for RootImage and RootImageOptions, and allow to either append a single list of options or tuples of partition_number:options.	2020-08-20 14:45:40 +01:00
Luca Boccassi	9ece644435	core: change RootImageOptions to use names instead of partition numbers Follow the designations from the Discoverable Partitions Specification	2020-08-20 13:58:02 +01:00
Luca Boccassi	bc8d56d305	core: use strv_split_colon_pairs when parsing RootImageOptions	2020-08-20 13:24:32 +01:00
Luca Boccassi	c20acbb2bd	core: cleanup unused variables Leftovers from previous implementation of MountImages feature, unused now	2020-08-20 13:24:32 +01:00
Lennart Poettering	476cfe626d	core: remove support for ConditionNull= The concept is flawed, and mostly useless. Let's finally remove it. It has been deprecated since `90a2ec10f2` (6 years ago) and we started to warn since `55dadc5c57` (1.5 years ago). Let's get rid of it altogether.	2020-08-20 14:01:25 +02:00
Lennart Poettering	4f55a5b0bf	core: add missing conditions/asserts to unit file parsing	2020-08-20 13:56:14 +02:00
Zbigniew Jędrzejewski-Szmek	ec673ad4ab	Merge pull request #16559 from benzea/benzea/memory-recursiveprot mount-setup: Enable memory_recursiveprot for cgroup2	2020-08-20 13:05:07 +02:00
Lennart Poettering	3242980582	core: create per-user inaccessible node from the service manager Previously, we'd create them from user-runtime-dir@.service. That has one benefit: since this service runs privileged, we can create the full set of device nodes. It has one major drawback though: it security-wise problematic to create files/directories in directories as privileged user in directories owned by unprivileged users, since they can use symlinks to redirect what we want to do. As a general rule we hence avoid this logic: only unpriv code should populate unpriv directories. Hence, let's move this code to an appropriate place in the service manager. This means we lose the inaccessible block device node, but since there's already a fallback in place, this shouldn't be too bad.	2020-08-20 10:18:02 +02:00
Lennart Poettering	9fac502920	nspawn,pid1: pass "inaccessible" nodes from cntr mgr to pid1 payload via /run/host Let's make /run/host the sole place we pass stuff from host to container in and place the "inaccessible" nodes in /run/host too. In contrast to the previous two commits this is a minor compat break, but not a relevant one I think. Previously the container manager would place these nodes in /run/systemd/inaccessible/ and that's where PID 1 in the container would try to add them too when missing. Container manager and PID 1 in the container would thus manage the same dir together. With this change the container manager now passes an immutable directory to the container and leaves /run/systemd entirely untouched, and managed exclusively by PID 1 inside the container, which is nice to have clear separation on who manages what. In order to make sure systemd then usses the /run/host/inaccesible/ nodes this commit changes PID 1 to look for that dir and if it exists will symlink it to /run/systemd/inaccessible. Now, this will work fine if new nspawn and new pid 1 in the container work together. as then the symlink is created and the difference between the two dirs won't matter. For the case where an old nspawn invokes a new PID 1: in this case things work as they always worked: the dir is managed together. For the case where different container manager invokes a new PID 1: in this case the nodes aren't typically passed in, and PID 1 in the container will try to create them and will likely fail partially (though gracefully) when trying to create char/block device nodes. THis is fine though as there are fallbacks in place for that case. For the case where a new nspawn invokes an old PID1: this is were the (minor) incompatibily happens: in this case new nspawn will place the nodes in the /run/host/inaccessible/ subdir, but the PID 1 in the container won't look for them there. Since the nodes are also not pre-created in /run/systed/inaccessible/ PID 1 will try to create them there as if a different container manager sets them up. This is of course not sexy, but is not a total loss, since as mentioned fallbacks are in place anyway. Hence I think it's OK to accept this minor incompatibility.	2020-08-20 10:17:52 +02:00
Zbigniew Jędrzejewski-Szmek	2eecdd1d69	Merge pull request #16790 from poettering/core-if-block-merge core: merge a few if blocks	2020-08-20 10:15:01 +02:00
Lennart Poettering	1f894e682c	machine-id-setup: don't use KVM or container manager supplied uuid if in chroot env Fixes: #16758	2020-08-19 18:23:11 +02:00
Lennart Poettering	4428c49db9	mount-setup: drop pointless zero initialization	2020-08-19 18:11:00 +02:00
Lennart Poettering	3196e42393	core: merge a few if blocks arg_system == true and getpid() == 1 hold under the very same condition this early in the main() function (this only changes later when we start parsing command lines, where arg_system = true is set if users invoke us in test mode even when getpid() != 1. Hence, let's simplify things, and merge a couple of if branches and not pretend they were orthogonal.	2020-08-19 18:06:12 +02:00
Michal Koutný	d9ef594454	cgroup: Cleanup function usage Some masks shouldn't be needed externally, so keep their functions in the module (others would fit there too but they're used in tests) to think twice if something would depend on them. Drop unused function cg_attach_many_everywhere. Use cgroup_realized instead of cgroup_path when we actually ask for realized. This should not cause any functional changes.	2020-08-19 11:41:53 +02:00
Michal Koutný	12b975e065	cgroup: Reduce unit_get_ancestor_disable_mask use The usage in unit_get_own_mask is redundant, we only need apply disable_mask at the end befor application, i.e. calculating enable or target mask. (IOW, we allow all configurations, but disabling affects effective controls.) Modify tests accordingly and add testing of enable mask. This is intended as cleanup, with no effect but changing unit_dump output.	2020-08-19 11:41:53 +02:00
Michal Koutný	4c591f3996	cgroup: Introduce family queueing instead of siblings The unit_add_siblings_to_cgroup_realize_queue does more than mere siblings queueing, hence define a family of a unit as (immediate) children of the unit and immediate children of all ancestors. Working with this abstraction simplifies the queuing calls and it shouldn't change the functionality.	2020-08-19 11:41:53 +02:00
Michal Koutný	f23ba94db3	cgroup: Implicit unit_invalidate_cgroup_members_masks Merge members mask invalidation into unit_add_siblings_to_cgroup_realize_queue, this way unit_realize_cgroup needn't be called with members mask invalidation. We have to retain the members mask invalidation in unit_load -- although active units would have cgroups (re)realized (unit_load queues for realization), the realization would happen with potentially stale mask.	2020-08-19 11:41:53 +02:00
Michal Koutný	fb46fca7e0	cgroup: Eager realization in unit_free unit_free(u) realizes direct parent and invalidates members mask of all ancestors. This isn't sufficient in v1 controller hierarchies since siblings of the freed unit may have existed only because of the removed unit. We cannot be lazy about the siblings because if parent(u) is also removed, it'd migrate and rmdir cgroups for siblings(u). However, realized masks of siblings(u) won't reflect this change. This was a non-issue earlier, because we weren't removing cgroup directories properly (effectively matching the stale realized mask), removal failed because of tasks left by missing migration (see previous commit). Therefore, ensure realization of all units necessary to clean up after the free'd unit. Fixes: #14149	2020-08-19 11:41:53 +02:00
Michal Koutný	7b63961415	cgroup: Swap cgroup v1 deletion and migration When we are about to derealize a controller on v1 cgroup, we first attempt to delete the controller cgroup and migrate afterwards. This doesn't work in practice because populated cgroup cannot be deleted. Furthermore, we leave out slices from migration completely, so (un)setting a control value on them won't realize their controller cgroup. Rework actual realization, unit_create_cgroup() becomes unit_update_cgroup() and make sure that controller hierarchies are reduced when given controller cgroup ceased to be needed. Note that with this we introduce slight deviation between v1 and v2 code -- when a descendant unit turns off a delegated controller, we attempt to disable it in ancestor slices. On v2 this may fail (kernel enforced, because of child cgroups using the controller), on v1 we'll migrate whole subtree and trim the subhierachy. (Previously, we wouldn't take away delegated controller, however, derealization was broken anyway.) Fixes: #14149	2020-08-19 11:41:53 +02:00
Benjamin Berg	56f47800d8	mount-setup: Enable memory_recursiveprot for cgroup2 When available, enable memory_recursiveprot. Realistically it always makes sense to delegate MemoryLow= and MemoryMin= to all children of a slice/unit. The kernel option is not enabled by default as it might cause regressions in some setups. However, it is the better default in general, and it results in a more flexible and obvious behaviour. The alternative to using this option would be for user's to also set DefaultMemoryLow= on slices when assigning MemoryLow=. However, this makes the effect of MemoryLow= on some children less obvious, as it could result in a lower protection rather than increasing it. From the kernel documentation: memory_recursiveprot Recursively apply memory.min and memory.low protection to entire subtrees, without requiring explicit downward propagation into leaf cgroups. This allows protecting entire subtrees from one another, while retaining free competition within those subtrees. This should have been the default behavior but is a mount-option to avoid regressing setups relying on the original semantics (e.g. specifying bogusly high 'bypass' protection values at higher tree levels). This was added in kernel commit 8a931f801340c (mm: memcontrol: recursive memory.low protection), which became available in 5.7 and was subsequently fixed in kernel 5.7.7 (mm: memcontrol: handle div0 crash race condition in memory.low).	2020-08-19 11:17:01 +02:00
Alyssa Ross	556a7bbed6	load-fragment: fix grammar in error messages	2020-08-18 20:56:59 +00:00
Lennart Poettering	3f181262f4	namespace: fix minor memory leak	2020-08-14 15:33:04 +02:00
Lennart Poettering	0a388dfcc5	core,home,machined: generate description fields for all groups we synthesize	2020-08-07 08:39:52 +02:00
Luca Boccassi	b3d133148e	core: new feature MountImages Follows the same pattern and features as RootImage, but allows an arbitrary mount point under / to be specified by the user, and multiple values - like BindPaths. Original implementation by @topimiettinen at: https://github.com/systemd/systemd/pull/14451 Reworked to use dissect's logic instead of bare libmount() calls and other review comments. Thanks Topi for the initial work to come up with and implement this useful feature.	2020-08-05 21:34:55 +01:00
Axel Rasmussen	a119185c02	selinux: improve comment about getcon_raw semantics This code was changed in this pull request: https://github.com/systemd/systemd/pull/16571 After some discussion and more investigation, we better understand what's going on. So, update the comment, so things are more clear to future readers.	2020-08-05 20:20:45 +02:00
Zbigniew Jędrzejewski-Szmek	d06bd2e785	Merge pull request #16596 from poettering/event-time-rel Conflict in src/libsystemd-network/test-ndisc-rs.c fixed manually.	2020-08-04 16:07:03 +02:00
Zbigniew Jędrzejewski-Szmek	94efaa3181	core: reset bus error before reuse From a report in https://bugzilla.redhat.com/show_bug.cgi?id=1861463: usb-gadget.target: Failed to load configuration: No such file or directory usb-gadget.target: Failed to load configuration: No such file or directory usb-gadget.target: Trying to enqueue job usb-gadget.target/start/fail usb-gadget.target: Failed to load configuration: No such file or directory Assertion '!bus_error_is_dirty(e)' failed at src/libsystemd/sd-bus/bus-error.c:239, function bus_error_setfv(). Ignoring. sys-devices-platform-soc-2100000.bus-2184000.usb-ci_hdrc.0-udc-ci_hdrc.0.device: Failed to enqueue SYSTEMD_WANTS= job, ignoring: Unit usb-gadget.target not found. I think this is the place where the reuse occurs: we call bus_unit_validate_load_state(unit, e) twice in a row.	2020-08-03 17:54:32 +02:00
Zbigniew Jędrzejewski-Szmek	7e62257219	Merge pull request #16308 from bluca/root_image_options service: add new RootImageOptions feature	2020-08-03 10:04:36 +02:00
Zbigniew Jędrzejewski-Szmek	b67ec8e5b2	pid1: stop limiting size of /dev/shm The explicit limit is dropped, which means that we return to the kernel default of 50% of RAM. See `362a55fc14` for a discussion why that is not as much as it seems. It turns out various applications need more space in /dev/shm and we would break them by imposing a low limit. While at it, rename the define and use a single macro for various tmpfs mounts. We don't really care what the purpose of the given tmpfs is, so it seems reasonable to use a single macro. This effectively reverts part of `7d85383edb`. Fixes #16617.	2020-07-30 18:48:35 +02:00
Luca Boccassi	18d7370587	service: add new RootImageOptions feature Allows to specify mount options for RootImage. In case of multi-partition images, the partition number can be prefixed followed by colon. Eg: RootImageOptions=1:ro,dev 2:nosuid nodev In absence of a partition number, 0 is assumed.	2020-07-29 17:17:32 +01:00
Michal Koutný	30ad3ca086	cgroup: Add root slice to cgroup realization queue When we're disabling controller on a direct child of root cgroup, we forgot to add root slice into cgroup realization queue, which prevented proper disabling of the controller (on unified hierarchy). The mechanism relying on "bounce from bottom and propagate up" in unit_create_cgroup doesn't work on unified hierarchy (leaves needn't be enabled). Drop it as we rely on the ancestors to be queued -- that's now intentional but was artifact of combining the two patches: `cb5e3bc37d` ("cgroup: Don't explicitly check for member in UNIT_BEFORE") v240~78 `65f6b6bdcb` ("core: fix re-realization of cgroup siblings") v245-rc1~153^2 Fixes: #14917	2020-07-28 15:49:24 +02:00
Michal Koutný	a479c21ed2	cgroup: Make realize_queue behave FIFO The current implementation is LIFO, which is a) confusing b) prevents some ordered operations on the cgroup tree (e.g. removing children before parents). Fix it quickly. Current list implementation turns this from O(1) to O(n) operation. Rework the lists later.	2020-07-28 15:49:24 +02:00
Lennart Poettering	39cf0351c5	tree-wide: make use of new relative time events in sd-event.h	2020-07-28 11:24:55 +02:00
Axel Rasmussen	199a892218	selinux: handle getcon_raw producing a NULL pointer, despite returning 0 Previously, we assumed that success meant we definitely got a valid pointer. There is at least one edge case where this is not true (i.e., we can get both a 0 return value, and also a NULL pointer): `4246bb550d/libselinux/src/procattr.c (L175)` When this case occurrs, if we don't check the pointer we SIGSEGV in early initialization.	2020-07-24 13:34:27 +09:00
Lennart Poettering	8047ac8fdc	core: clean more env vars from env block pid1 receives We generally clean all env vars we use ourselves to communicate with out childrens. We forgot some more recent additions however. Let's correct that.	2020-07-23 18:30:15 +02:00
Lennart Poettering	00b868e857	Merge pull request #16542 from keszybz/make-targets-fail-again Make targets fail again	2020-07-23 08:37:47 +02:00
Lennart Poettering	c3f8a065e9	execute: take ownership of more fields in ExecParameters Let's simplify things a bit, and take ownership of more fields in ExecParameters, so that they are automatically freed when the structure is released.	2020-07-23 08:37:21 +02:00
Zbigniew Jędrzejewski-Szmek	94d1ddbd7c	pid1: target units can fail through dependencies Fixes #16401. `c80a9a33d0` introduced the .can_fail field, but didn't set it on .targets. Targets can fail through dependencies. This leaves .slice and .device units as the types that cannot fail. $ systemctl cat bad.service bad.target bad-fallback.service [Service] Type=oneshot ExecStart=false [Unit] OnFailure=bad-fallback.service [Service] Type=oneshot ExecStart=echo Fixing everythign! $ sudo systemctl start bad.target systemd[1]: Starting bad.service... systemd[1]: bad.service: Main process exited, code=exited, status=1/FAILURE systemd[1]: bad.service: Failed with result 'exit-code'. systemd[1]: Failed to start bad.service. systemd[1]: Dependency failed for bad.target. systemd[1]: bad.target: Job bad.target/start failed with result 'dependency'. systemd[1]: bad.target: Triggering OnFailure= dependencies. systemd[1]: Starting bad-fallback.service... echo[46901]: Fixing everythign! systemd[1]: bad-fallback.service: Succeeded. systemd[1]: Finished bad-fallback.service.	2020-07-22 17:58:12 +02:00

1 2 3 4 5 ...

5606 commits