From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from atuin.qyliss.net (localhost [IPv6:::1]) by atuin.qyliss.net (Postfix) with ESMTP id D9EACA13B; Sun, 30 Aug 2026 02:54:24 +0000 (UTC) Received: by atuin.qyliss.net (Postfix, from userid 993) id 5973CA0CB; Sun, 30 Aug 2026 02:54:22 +0000 (UTC) X-Spam-Checker-Version: SpamAssassin 4.0.1 (2024-03-26) on atuin.qyliss.net X-Spam-Level: X-Spam-Status: No, score=-0.1 required=3.0 tests=DKIM_SIGNED,DKIM_VALID, DKIM_VALID_AU,DMARC_PASS,FREEMAIL_FROM,RCVD_IN_DNSWL_NONE, SPF_HELO_NONE autolearn=unavailable autolearn_force=no version=4.0.1 Received: from mail-yx1-xb131.google.com (mail-yx1-xb131.google.com [IPv6:2607:f8b0:4864:20::b131]) by atuin.qyliss.net (Postfix) with ESMTPS id 5059AA0C9 for ; Sun, 30 Aug 2026 02:54:20 +0000 (UTC) Received: by mail-yx1-xb131.google.com with SMTP id 956f58d0204a3-66c70f69d3fso2647239d50.1 for ; Sat, 29 Aug 2026 19:54:20 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1788058457; x=1788663257; darn=spectrum-os.org; h=content-type:in-reply-to:autocrypt:from:content-language:references :cc:to:subject:user-agent:mime-version:date:message-id:from:to:cc :subject:date:message-id:reply-to:content-type; bh=m8BzYuz67Bk3Z7rmWn/4rc5r5iWEW+WfPYt+Kn5Yvrg=; b=nxCrr2hc3KNd/58BGONLL7o89hhqJ8vE2c/94F7G0gGZKB2wMBAfWJTaZ5UXmOMtPE D9gd6kwWqZIPO/7v6R8MwOfrvEvhe5T1hwK7rLEy/tUIG4bUoYStQAE6Ee311BLvmKGZ R8Yi3n9Xv8DUb8fwZg6B477VoPBfDnJiMoX6lu/xEoKx+xydqdr8jLX2JnYxSwTFKrdV uPcsP38pzLqKpnmCU2tTU0yGBmmDzf+WcaoGK4Tr/rqx5aSoTtN6IVLZy1m5nkty2LGx ONXaHnJ+VErv639glmGsvwrqjEr0ZGKxpEdx53X6A0rdis5GPGZiVgfBOY2pV8JuvaaP XN4g== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788058457; x=1788663257; h=content-type:in-reply-to:autocrypt:from:content-language:references :cc:to:subject:user-agent:mime-version:date:message-id:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=m8BzYuz67Bk3Z7rmWn/4rc5r5iWEW+WfPYt+Kn5Yvrg=; b=Ebi8oXK0AcfKMnZXBTNHirMLXHb41XpjawXS7a+kcZ94ah9pMYOFkCd2iC1HHbPOWs thfIe/jIdhN+muNYWCBadeVDwJauecFDyzPYT28Ded9nzUJYeNyFJQ7tyheSEulTgtS2 t3hPtjm/qrErTbBksV8QQYhdIACjcQikKTUkuaIX6knI6BSmCRPncOWFfN7R1+tMVjvt nk5O2+tPIwYLDMAn3YnhfZtmSpv8EV/ywgO5BuxSiZLRWm6vN2klZbbzsIrCmSU08TPR bff31b7i33+9k/auyK9KpOdXEAfSQF97wWVAS1FEvS3xj3Jchp2ePQPb3QkIo4UYKmGb kryQ== X-Gm-Message-State: AFuF++m/S1pt7qudJKGjlXLUjlxnO+b47hzgw5mfRaYaRPnRWjbw53DA 8uaxgjYoKfbteXct5Y2eaMOYZ7u3BP+H1wdQZxsircGPn7XxPmaQn51B9ZqnhA== X-Gm-Gg: AYBFou0QRy75sOlz3ZVCSUSIBaMl4cjMFv+jKncCPmiSBbPhkNV7gAuwQ4QDDJ8C65j 51AnIv59Ck8u/z5WtuCn6NT7/Q5bD9HyY/mLsTDwJHV6e8cc0fiyCQ4fNhsRjguYKArANRgEY4L DV8r6FmzPVya/GRD/rtP6qR4/cFvWBaZrbuZewFr17OQFrAr0vGsuURPJ2tE3DJAT2q0z9LBgq6 E6L1zL1W/51jURRyUfeZK7wfl7oy27gfr0XQ8XykiVTK6m91Lp28LP+wQgESz0C260Zvh/9yUcT RR+5RagqbJtW8ZFqBrVty5SHIRPEZmYT6oZlFswHMSsNXrQ+TIrKHlOtvB6bKHX3FJuSlZHiieq iI+caKPtiFPgzQ70ybfWGa0m6pZdETQ/ZaBrj4MULMjcmdPvZ5R1aFcBd386nZ/1h9GesGinxsH aPX5PJLx7bV6ep9XjLLfh12Na0FLEWdXOILk8QC8WzaPCXV1kNUrAJc0T6uwr5nrg= X-Received: by 2002:a05:690e:d49:b0:667:c7ba:99ee with SMTP id 956f58d0204a3-66e4c681b94mr4893106d50.2.1788058456855; Sat, 29 Aug 2026 19:54:16 -0700 (PDT) Received: from [10.138.10.6] ([185.98.168.14]) by smtp.gmail.com with ESMTPSA id 956f58d0204a3-66e4ed39d4csm3498426d50.21.2026.08.29.19.54.14 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Sat, 29 Aug 2026 19:54:15 -0700 (PDT) Message-ID: Date: Sat, 29 Aug 2026 22:54:10 -0400 MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH v7 02/19] tools: Add control group manager To: Alyssa Ross References: <20260821-cgroups-v7-0-7f1870dedefc@gmail.com> <20260821-cgroups-v7-2-7f1870dedefc@gmail.com> <87y0dt7z7e.fsf@alyssa.is> Content-Language: en-US From: Demi Marie Obenour Autocrypt: addr=demiobenour@gmail.com; keydata= xsFNBFp+A0oBEADffj6anl9/BHhUSxGTICeVl2tob7hPDdhHNgPR4C8xlYt5q49yB+l2nipd aq+4Gk6FZfqC825TKl7eRpUjMriwle4r3R0ydSIGcy4M6eb0IcxmuPYfbWpr/si88QKgyGSV Z7GeNW1UnzTdhYHuFlk8dBSmB1fzhEYEk0RcJqg4AKoq6/3/UorR+FaSuVwT7rqzGrTlscnT DlPWgRzrQ3jssesI7sZLm82E3pJSgaUoCdCOlL7MMPCJwI8JpPlBedRpe9tfVyfu3euTPLPx wcV3L/cfWPGSL4PofBtB8NUU6QwYiQ9Hzx4xOyn67zW73/G0Q2vPPRst8LBDqlxLjbtx/WLR 6h3nBc3eyuZ+q62HS1pJ5EvUT1vjyJ1ySrqtUXWQ4XlZyoEFUfpJxJoN0A9HCxmHGVckzTRl 5FMWo8TCniHynNXsBtDQbabt7aNEOaAJdE7to0AH3T/Bvwzcp0ZJtBk0EM6YeMLtotUut7h2 Bkg1b//r6bTBswMBXVJ5H44Qf0+eKeUg7whSC9qpYOzzrm7+0r9F5u3qF8ZTx55TJc2g656C 9a1P1MYVysLvkLvS4H+crmxA/i08Tc1h+x9RRvqba4lSzZ6/Tmt60DPM5Sc4R0nSm9BBff0N m0bSNRS8InXdO1Aq3362QKX2NOwcL5YaStwODNyZUqF7izjK4QARAQABzTxEZW1pIE1hcmll IE9iZW5vdXIgKGxvdmVyIG9mIGNvZGluZykgPGRlbWlvYmVub3VyQGdtYWlsLmNvbT7CwXgE EwECACIFAlp+A0oCGwMGCwkIBwMCBhUIAgkKCwQWAgMBAh4BAheAAAoJELKItV//nCLBhr8Q AK/xrb4wyi71xII2hkFBpT59ObLN+32FQT7R3lbZRjVFjc6yMUjOb1H/hJVxx+yo5gsSj5LS 9AwggioUSrcUKldfA/PKKai2mzTlUDxTcF3vKx6iMXKA6AqwAw4B57ZEJoMM6egm57TV19kz PMc879NV2nc6+elaKl+/kbVeD3qvBuEwsTe2Do3HAAdrfUG/j9erwIk6gha/Hp9yZlCnPTX+ VK+xifQqt8RtMqS5R/S8z0msJMI/ajNU03kFjOpqrYziv6OZLJ5cuKb3bZU5aoaRQRDzkFIR 6aqtFLTohTo20QywXwRa39uFaOT/0YMpNyel0kdOszFOykTEGI2u+kja35g9TkH90kkBTG+a EWttIht0Hy6YFmwjcAxisSakBuHnHuMSOiyRQLu43ej2+mDWgItLZ48Mu0C3IG1seeQDjEYP tqvyZ6bGkf2Vj+L6wLoLLIhRZxQOedqArIk/Sb2SzQYuxN44IDRt+3ZcDqsPppoKcxSyd1Ny 2tpvjYJXlfKmOYLhTWs8nwlAlSHX/c/jz/ywwf7eSvGknToo1Y0VpRtoxMaKW1nvH0OeCSVJ itfRP7YbiRVc2aNqWPCSgtqHAuVraBRbAFLKh9d2rKFB3BmynTUpc1BQLJP8+D5oNyb8Ts4x Xd3iV/uD8JLGJfYZIR7oGWFLP4uZ3tkneDfYzsFNBFp+A0oBEAC9ynZI9LU+uJkMeEJeJyQ/ 8VFkCJQPQZEsIGzOTlPnwvVna0AS86n2Z+rK7R/usYs5iJCZ55/JISWd8xD57ue0eB47bcJv VqGlObI2DEG8TwaW0O0duRhDgzMEL4t1KdRAepIESBEA/iPpI4gfUbVEIEQuqdqQyO4GAe+M kD0Hy5JH/0qgFmbaSegNTdQg5iqYjRZ3ttiswalql1/iSyv1WYeC1OAs+2BLOAT2NEggSiVO txEfgewsQtCWi8H1SoirakIfo45Hz0tk/Ad9ZWh2PvOGt97Ka85o4TLJxgJJqGEnqcFUZnJJ riwoaRIS8N2C8/nEM53jb1sH0gYddMU3QxY7dYNLIUrRKQeNkF30dK7V6JRH7pleRlf+wQcN fRAIUrNlatj9TxwivQrKnC9aIFFHEy/0mAgtrQShcMRmMgVlRoOA5B8RTulRLCmkafvwuhs6 dCxN0GNAORIVVFxjx9Vn7OqYPgwiofZ6SbEl0hgPyWBQvE85klFLZLoj7p+joDY1XNQztmfA rnJ9x+YV4igjWImINAZSlmEcYtd+xy3Li/8oeYDAqrsnrOjb+WvGhCykJk4urBog2LNtcyCj kTs7F+WeXGUo0NDhbd3Z6AyFfqeF7uJ3D5hlpX2nI9no/ugPrrTVoVZAgrrnNz0iZG2DVx46 x913pVKHl5mlYQARAQABwsFfBBgBAgAJBQJafgNKAhsMAAoJELKItV//nCLBwNIP/AiIHE8b oIqReFQyaMzxq6lE4YZCZNj65B/nkDOvodSiwfwjjVVE2V3iEzxMHbgyTCGA67+Bo/d5aQGj gn0TPtsGzelyQHipaUzEyrsceUGWYoKXYyVWKEfyh0cDfnd9diAm3VeNqchtcMpoehETH8fr RHnJdBcjf112PzQSdKC6kqU0Q196c4Vp5HDOQfNiDnTf7gZSj0BraHOByy9LEDCLhQiCmr+2 E0rW4tBtDAn2HkT9uf32ZGqJCn1O+2uVfFhGu6vPE5qkqrbSE8TG+03H8ecU2q50zgHWPdHM OBvy3EhzfAh2VmOSTcRK+tSUe/u3wdLRDPwv/DTzGI36Kgky9MsDC5gpIwNbOJP2G/q1wT1o Gkw4IXfWv2ufWiXqJ+k7HEi2N1sree7Dy9KBCqb+ca1vFhYPDJfhP75I/VnzHVssZ/rYZ9+5 1yDoUABoNdJNSGUYl+Yh9Pw9pE3Kt4EFzUlFZWbE4xKL/NPno+z4J9aWemLLszcYz/u3XnbO vUSQHSrmfOzX3cV4yfmjM5lewgSstoxGyTx2M8enslgdXhPthZlDnTnOT+C+OTsh8+m5tos8 HQjaPM01MKBiAqdPgksm1wu2DrrwUi6ChRVTUBcj6+/9IJ81H2P2gJk3Ls3AVIxIffLoY34E +MYSfkEjBz0E8CLOcAw7JIwAaeBT In-Reply-To: <87y0dt7z7e.fsf@alyssa.is> Content-Type: multipart/signed; micalg=pgp-sha256; protocol="application/pgp-signature"; boundary="------------2WLnCRpR0JNoIko5fBqrA06A" Message-ID-Hash: 7GRIMYBOHZYVUKH3B3A6QQMOPGG7TQQI X-Message-ID-Hash: 7GRIMYBOHZYVUKH3B3A6QQMOPGG7TQQI X-MailFrom: demiobenour@gmail.com X-Mailman-Rule-Hits: member-moderation X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; loop; banned-address; header-match-devel.spectrum-os.org-0; header-match-devel.spectrum-os.org-1; header-match-devel.spectrum-os.org-2; header-match-devel.spectrum-os.org-3; header-match-devel.spectrum-os.org-4; emergency CC: Spectrum OS Development X-Mailman-Version: 3.3.10 Precedence: list List-Id: Patches and low-level development discussion Archived-At: List-Archive: List-Help: List-Owner: List-Post: List-Subscribe: List-Unsubscribe: This is an OpenPGP/MIME signed message (RFC 4880 and 3156) --------------2WLnCRpR0JNoIko5fBqrA06A Content-Type: multipart/mixed; boundary="------------V404WttBfSpXI9pQnNBJbV8X"; protected-headers="v1"; hp="clear" Message-ID: Date: Sat, 29 Aug 2026 22:54:10 -0400 MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH v7 02/19] tools: Add control group manager To: Alyssa Ross Cc: Spectrum OS Development References: <20260821-cgroups-v7-0-7f1870dedefc@gmail.com> <20260821-cgroups-v7-2-7f1870dedefc@gmail.com> <87y0dt7z7e.fsf@alyssa.is> Content-Language: en-US From: Demi Marie Obenour Autocrypt: addr=demiobenour@gmail.com; keydata= xsFNBFp+A0oBEADffj6anl9/BHhUSxGTICeVl2tob7hPDdhHNgPR4C8xlYt5q49yB+l2nipd aq+4Gk6FZfqC825TKl7eRpUjMriwle4r3R0ydSIGcy4M6eb0IcxmuPYfbWpr/si88QKgyGSV Z7GeNW1UnzTdhYHuFlk8dBSmB1fzhEYEk0RcJqg4AKoq6/3/UorR+FaSuVwT7rqzGrTlscnT DlPWgRzrQ3jssesI7sZLm82E3pJSgaUoCdCOlL7MMPCJwI8JpPlBedRpe9tfVyfu3euTPLPx wcV3L/cfWPGSL4PofBtB8NUU6QwYiQ9Hzx4xOyn67zW73/G0Q2vPPRst8LBDqlxLjbtx/WLR 6h3nBc3eyuZ+q62HS1pJ5EvUT1vjyJ1ySrqtUXWQ4XlZyoEFUfpJxJoN0A9HCxmHGVckzTRl 5FMWo8TCniHynNXsBtDQbabt7aNEOaAJdE7to0AH3T/Bvwzcp0ZJtBk0EM6YeMLtotUut7h2 Bkg1b//r6bTBswMBXVJ5H44Qf0+eKeUg7whSC9qpYOzzrm7+0r9F5u3qF8ZTx55TJc2g656C 9a1P1MYVysLvkLvS4H+crmxA/i08Tc1h+x9RRvqba4lSzZ6/Tmt60DPM5Sc4R0nSm9BBff0N m0bSNRS8InXdO1Aq3362QKX2NOwcL5YaStwODNyZUqF7izjK4QARAQABzTxEZW1pIE1hcmll IE9iZW5vdXIgKGxvdmVyIG9mIGNvZGluZykgPGRlbWlvYmVub3VyQGdtYWlsLmNvbT7CwXgE EwECACIFAlp+A0oCGwMGCwkIBwMCBhUIAgkKCwQWAgMBAh4BAheAAAoJELKItV//nCLBhr8Q AK/xrb4wyi71xII2hkFBpT59ObLN+32FQT7R3lbZRjVFjc6yMUjOb1H/hJVxx+yo5gsSj5LS 9AwggioUSrcUKldfA/PKKai2mzTlUDxTcF3vKx6iMXKA6AqwAw4B57ZEJoMM6egm57TV19kz PMc879NV2nc6+elaKl+/kbVeD3qvBuEwsTe2Do3HAAdrfUG/j9erwIk6gha/Hp9yZlCnPTX+ VK+xifQqt8RtMqS5R/S8z0msJMI/ajNU03kFjOpqrYziv6OZLJ5cuKb3bZU5aoaRQRDzkFIR 6aqtFLTohTo20QywXwRa39uFaOT/0YMpNyel0kdOszFOykTEGI2u+kja35g9TkH90kkBTG+a EWttIht0Hy6YFmwjcAxisSakBuHnHuMSOiyRQLu43ej2+mDWgItLZ48Mu0C3IG1seeQDjEYP tqvyZ6bGkf2Vj+L6wLoLLIhRZxQOedqArIk/Sb2SzQYuxN44IDRt+3ZcDqsPppoKcxSyd1Ny 2tpvjYJXlfKmOYLhTWs8nwlAlSHX/c/jz/ywwf7eSvGknToo1Y0VpRtoxMaKW1nvH0OeCSVJ itfRP7YbiRVc2aNqWPCSgtqHAuVraBRbAFLKh9d2rKFB3BmynTUpc1BQLJP8+D5oNyb8Ts4x Xd3iV/uD8JLGJfYZIR7oGWFLP4uZ3tkneDfYzsFNBFp+A0oBEAC9ynZI9LU+uJkMeEJeJyQ/ 8VFkCJQPQZEsIGzOTlPnwvVna0AS86n2Z+rK7R/usYs5iJCZ55/JISWd8xD57ue0eB47bcJv VqGlObI2DEG8TwaW0O0duRhDgzMEL4t1KdRAepIESBEA/iPpI4gfUbVEIEQuqdqQyO4GAe+M kD0Hy5JH/0qgFmbaSegNTdQg5iqYjRZ3ttiswalql1/iSyv1WYeC1OAs+2BLOAT2NEggSiVO txEfgewsQtCWi8H1SoirakIfo45Hz0tk/Ad9ZWh2PvOGt97Ka85o4TLJxgJJqGEnqcFUZnJJ riwoaRIS8N2C8/nEM53jb1sH0gYddMU3QxY7dYNLIUrRKQeNkF30dK7V6JRH7pleRlf+wQcN fRAIUrNlatj9TxwivQrKnC9aIFFHEy/0mAgtrQShcMRmMgVlRoOA5B8RTulRLCmkafvwuhs6 dCxN0GNAORIVVFxjx9Vn7OqYPgwiofZ6SbEl0hgPyWBQvE85klFLZLoj7p+joDY1XNQztmfA rnJ9x+YV4igjWImINAZSlmEcYtd+xy3Li/8oeYDAqrsnrOjb+WvGhCykJk4urBog2LNtcyCj kTs7F+WeXGUo0NDhbd3Z6AyFfqeF7uJ3D5hlpX2nI9no/ugPrrTVoVZAgrrnNz0iZG2DVx46 x913pVKHl5mlYQARAQABwsFfBBgBAgAJBQJafgNKAhsMAAoJELKItV//nCLBwNIP/AiIHE8b oIqReFQyaMzxq6lE4YZCZNj65B/nkDOvodSiwfwjjVVE2V3iEzxMHbgyTCGA67+Bo/d5aQGj gn0TPtsGzelyQHipaUzEyrsceUGWYoKXYyVWKEfyh0cDfnd9diAm3VeNqchtcMpoehETH8fr RHnJdBcjf112PzQSdKC6kqU0Q196c4Vp5HDOQfNiDnTf7gZSj0BraHOByy9LEDCLhQiCmr+2 E0rW4tBtDAn2HkT9uf32ZGqJCn1O+2uVfFhGu6vPE5qkqrbSE8TG+03H8ecU2q50zgHWPdHM OBvy3EhzfAh2VmOSTcRK+tSUe/u3wdLRDPwv/DTzGI36Kgky9MsDC5gpIwNbOJP2G/q1wT1o Gkw4IXfWv2ufWiXqJ+k7HEi2N1sree7Dy9KBCqb+ca1vFhYPDJfhP75I/VnzHVssZ/rYZ9+5 1yDoUABoNdJNSGUYl+Yh9Pw9pE3Kt4EFzUlFZWbE4xKL/NPno+z4J9aWemLLszcYz/u3XnbO vUSQHSrmfOzX3cV4yfmjM5lewgSstoxGyTx2M8enslgdXhPthZlDnTnOT+C+OTsh8+m5tos8 HQjaPM01MKBiAqdPgksm1wu2DrrwUi6ChRVTUBcj6+/9IJ81H2P2gJk3Ls3AVIxIffLoY34E +MYSfkEjBz0E8CLOcAw7JIwAaeBT In-Reply-To: <87y0dt7z7e.fsf@alyssa.is> --------------V404WttBfSpXI9pQnNBJbV8X Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: quoted-printable On 8/26/26 10:07, Alyssa Ross wrote: > Demi Marie Obenour writes: >=20 >> This program has two modes: >> >> 1. Create a control group if it doesn't exist, optionally wait for oth= er >> programs in it to exit, and exec another program in it. >> >> 2. Purge a control group: kill all programs in it, then delete it. >> >> Locking is used to ensure that concurrent invocations are safe. >> >> Signed-off-by: Demi Marie Obenour >=20 > The structure of the program is looking really good now. All remaining= > comments are minor, except for making sure we're doing the right thing > with enabling controllers. Pay attention to naming =E2=80=94 good name= s are > really important for making it clear to readers what a program does. That it is! >> diff --git a/tools/cgroup-setup/src/cgroup.rs b/tools/cgroup-setup/src= /cgroup.rs >> new file mode 100644 >> index 0000000000000000000000000000000000000000..463d3fc87ee34ccfbf399a= 337e2f49d9031728ee >> --- /dev/null >> +++ b/tools/cgroup-setup/src/cgroup.rs >> @@ -0,0 +1,289 @@ >> +// SPDX-FileCopyrightText: 2026 Demi Marie Obenour >> +// SPDX-License-Identifier: EUPL-1.2+ >> + >> +use std::ffi::OsStr; >> +use std::fs::File; >> +use std::io::{Read as _, Seek as _, Write as _}; >> +use std::os::unix::prelude::*; >> + >> +use std::path::{Component, Path, PathBuf}; >> + >> +use rustix::fs::{AtFlags, CWD, Dir, FlockOperation}; >> +use rustix::path; >> +use rustix::{ >> + fs::{Mode, OFlags, ResolveFlags}, >> + io::Errno, >> +}; >> + >> +pub enum OpenFlags { >> + Read, >> + Write, >> + Directory, >> +} >> + >> +#[derive(Debug)] >> +pub(crate) struct Cgroup { >> + fd: Vec, >> +} >> + >> +impl AsFd for Cgroup { >> + fn as_fd(&self) -> BorrowedFd<'_> { >> + self.fd.last().unwrap().as_fd() >> + } >> +} >> + >> +fn assert_single_component(component: &Path) { >> + match component.as_os_str().as_bytes() { >> + b"" | b"." | b".." =3D> panic!("bad component"), >> + c if c.contains(&b'\0') =3D> panic!("NUL in component"), >> + c if c.contains(&b'/') =3D> panic!("/ in component"), >> + _ =3D> {} >> + } >> +} >> + >> +// Wrapper around openat2() with better defaults. >> +pub fn openat2_simple( >> + fd: impl AsFd, >> + path: impl path::Arg, >> + flags: OpenFlags, >> +) -> Result { >> + rustix::fs::openat2( >> + fd.as_fd(), >> + path, >> + OFlags::CLOEXEC >> + | match flags { >> + OpenFlags::Read =3D> OFlags::RDONLY | OFlags::NOCTTY,= >> + OpenFlags::Write =3D> OFlags::WRONLY | OFlags::NOCTTY= , >> + OpenFlags::Directory =3D> OFlags::RDONLY | OFlags::DI= RECTORY, >> + }, >> + Mode::empty(), >> + ResolveFlags::NO_SYMLINKS | ResolveFlags::NO_MAGICLINKS | Res= olveFlags::NO_XDEV, >> + ) >> +} >> + >> +pub const DEFAULT_LEAF: &str =3D "$inner.service"; >> + >> +pub fn check_path(path: &Path) -> Result<(), String> { >> + let bytes =3D path.as_os_str().as_bytes(); >> + // Path::components() skips ., so use string manipulation instead= =2E >> + for component in bytes[path.is_absolute() as usize..].split(|&b| = b =3D=3D b'/') { >> + if matches!(component, b"" | b"." | b"..") { >> + return Err(format!("cgroup path {path:?} isn't canonical"= )); >> + } >> + } >> + Ok(()) >> +} >> + >> +// Remove all subdirectories of the given directory recursively, but = not the >> +// directory itself. The directory file descriptor is closed. >=20 > [nit] That's clear from the type signature, so probably doesn't need to= > be explicitly documented. Will delete. >> +// >> +// This isn't the most efficient possible algorithm, but simplicity i= s more >> +// important than performance in this case. Also, it keeps open more= file >> +// descriptors than strictly necessary, but Spectrum runs with a very= high limit >> +// for the number of open file descriptors, and it uses shallow contr= ol group >> +// hierarchies. >> +// >> +// This uses a recursive algorithm, but so does std::fs::remove_dir_a= ll(). Trying >> +// to be more robust than the standard library is not worthwhile. In= particular, >> +// the standard library function must be safe on systems where untrus= ted users (or >> +// even network endpoints!) can create deeply nested directory trees,= whereas in >> +// Spectrum cgroups are only writeable by root. >> +fn remove_recursively(mut dirfd: Dir, remaining_depth: usize) -> Resu= lt<(), Errno> { >> + if remaining_depth < 1 { >> + panic!("control groups too deeply nested"); >> + } >> + while let Some(element) =3D dirfd.next() { >=20 > "element" is a bit of an odd name for a directory entry, no? I will rename this to "entry". >> + let parent_fd =3D dirfd.fd().unwrap(); >=20 > Could be lifted out of the loop, right? That doesn't compile: dirfd.fd() takes an immutable borrow, while dirfd.next() takes a mutable one. >> + let element =3D element.expect("Iterating through a cgroup di= rectory failed?"); >> + let path =3D element.file_name(); >=20 > "name" would probably be clearer than "path", since we know it's a > single component (and it's consistent with the file_name method). Will fix. >> + if element.file_type() !=3D rustix::fs::FileType::Directory |= | path =3D=3D c"." || path =3D=3D c".." { >> + continue; >> + } >> + let fd =3D openat2_simple(parent_fd, path, OpenFlags::Directo= ry)?; >> + remove_recursively(Dir::new(fd).unwrap(), remaining_depth - 1= )?; >> + rustix::fs::unlinkat(parent_fd, path, AtFlags::REMOVEDIR)?; >> + } >> + Ok(()) >> +} >> + >> +fn assert_simple_path(current_cgroup: &Path) { >> + let current_cgroup =3D current_cgroup.as_os_str().as_bytes(); >> + if !matches!(current_cgroup, b"" | b".") { >> + for component in current_cgroup.split(|&b| b =3D=3D b'/') { >> + assert_single_component(Path::new(OsStr::from_bytes(compo= nent))); >> + } >> + } >> +} >=20 > Very non-obvious what this does =E2=80=94 of course you'll get a lot of= "single > component"s if you split a path on /. What you actually want to do is > just check for no null bytes or .. components, right? Given > assert_single_component is doing a more specific check than just single= > components, it should be renamed accordingly. (Although I'd struggle t= o > think of a name, because I still find what it's checking, and where we > check it, to be a bit arbitrary, especially when it's a path that's com= e > from the kernel=E2=80=A6) It's checking that a path component is =E2=80=9Csimple=E2=80=9D, meaning = no NUL or / and not =E2=80=9C.=E2=80=9D or =E2=80=9C..=E2=80=9D. Those are the = criteria for =E2=80=9Cthe path is valid and its lookup does not cross any directories=E2=80=9D. > With check_path as well we have a confusing collection of subtly > different, inconsistently named path checking functions. These should > be unified if possible, named systematically if not, and in either case= > it should be clear from the name what the function is for. I will inline check_path into its single caller, and delete the assert_* functions. >> +// Convert the cgroup path to one relative to /sys/fs/cgroup. >> +// >> +// If the path starts with /, the leading / is removed and the result= is returned >> +// without further processing. Otherwise, the current cgroup is read= from >> +// /proc/thread-self/cgroup. If its last component is $inner.service= , that is >> +// removed. Finally, the current cgroup is prepended to the provided= cgroup path, >> +// with a single / as separator. The result of this operation is ret= urned. >> +fn prepend_current_cgroup_if_needed(path: &Path) -> Result { >> + if let Ok(suffix) =3D path.strip_prefix("/") { >> + return Ok(suffix.to_owned()); >> + } >> + assert_simple_path(path); >> + // /proc/thread-self is the same as /proc/self, except for the cu= rrent thread >> + // instead of the initial thread. In this case, the two are iden= tical, but >> + // using /proc/thread-self is better practice as it is correct in= more cases. >> + // Reading /proc/thread-self/cgroup should never fail unless the = system is >> + // seriously broken. >> + let current_cgroup =3D >> + std::fs::read("/proc/thread-self/cgroup").expect("cannot read= /proc/thread-self/cgroup"); >> + // Using this on a system without cgroups v2 mounted is user erro= r >> + // and not supported. >> + let current_cgroup =3D current_cgroup >> + .strip_prefix(b"0::/") >> + .and_then(|e| e.strip_suffix(b"\n")) >> + .ok_or_else(|| { >> + "/proc/thread-self/cgroup doesn't start with 0::/ or does= n't end with a newline.\n\ >> + Either cgroups aren't in use at all, or you are using cgr= oups v1." >> + .to_owned() >> + })?; >> + let mut current_cgroup =3D PathBuf::from(OsStr::from_bytes(curren= t_cgroup)); >> + assert_simple_path(¤t_cgroup); >> + // Strip the implied $inner.service suffix. >> + // This is used to satisfy the "no internal processes" rule. >> + if current_cgroup.ends_with(Path::new(DEFAULT_LEAF)) { >> + assert!(current_cgroup.pop()); >> + } >> + // "." refers to the current cgroup. >> + if path !=3D Path::new(".") { >> + current_cgroup.push(path); >> + } >> + Ok(current_cgroup) >> +} >> + >> +pub(crate) fn write_value(fd: &dyn AsFd, name: &Path, value: &[u8]) -= > Result<(), String> { >> + let fd =3D openat2_simple(fd, name, OpenFlags::Write) >> + .map_err(|e| format!("Cannot open {name:?}: {e}"))?; >> + File::from(fd).write_all(value).map_err(|e| { >> + format!( >> + "Cannot write {:?} to {name:?}: {e}", >> + OsStr::from_bytes(value) >> + ) >> + }) >> +} >> + >> +impl Cgroup { >> + pub fn new(path: &Path) -> Result { >> + let cgroup_root =3D rustix::fs::openat2( >> + CWD, >> + Path::new("/sys/fs/cgroup"), >> + OFlags::CLOEXEC | OFlags::DIRECTORY | OFlags::RDONLY, >> + Mode::empty(), >> + ResolveFlags::NO_SYMLINKS | ResolveFlags::NO_MAGICLINKS, >> + ) >> + .map_err(|e| format!("Cannot open /sys/fs/cgroup: {e}"))?; >> + // It's simpler to always have the root cgroup at the bottom = of the stack, >> + // even though no lock needs to be taken on it. Otherwise, o= ne would need >> + // to special-case the cgroup root. One could remove the fir= st element if >> + // there is more than one element in the vector, but that's n= ot worth it. >> + // cgroup-setup doesn't operate in an environment where FDs a= re a limited >> + // resource. >> + let mut cgroup =3D Self { >> + fd: vec![cgroup_root], >> + }; >> + >> + let path =3D prepend_current_cgroup_if_needed(path)?; >> + for component in path.components() { >> + let Component::Normal(component) =3D component else { >> + unreachable!() >> + }; >> + let sub_fd =3D openat2_simple(&cgroup, component, OpenFla= gs::Directory) >> + .map_err(|e| format!("Cannot open sub-cgroup {compone= nt:?}: {e}"))?; >> + // Take a shared lock on the cgroup. >> + rustix::fs::flock(&sub_fd, FlockOperation::LockShared) >> + .map_err(|e| format!("Cannot lock sub-cgroup {compone= nt:?}: {e}"))?; >> + cgroup.fd.push(sub_fd); >> + } >> + Ok(cgroup) >> + } >> + >> + pub fn wait_for_empty(fd: &dyn AsFd) -> std::io::Result<()> { >> + let wait_file =3D openat2_simple(fd, c"cgroup.events", OpenFl= ags::Read)?; >> + let mut wait_fd =3D File::from(wait_file); >> + let mut v =3D vec![]; >> + loop { >> + v.clear(); >> + wait_fd >> + .seek(std::io::SeekFrom::Start(0)) >> + .expect("Seek on control group file should succeed");= >> + wait_fd >> + .read_to_end(&mut v) >> + .expect("reading from control group should work"); >> + // Check that the cgroup isn't already empty. If it was,= >> + // the kernel would not send an event and poll() would wa= it >> + // forever. >> + if v.split(|&c| c =3D=3D b'\n').any(|line| line =3D=3D b"= populated 0") { >> + break; >> + } >> + let mut fds =3D libc::pollfd { >> + fd: wait_fd.as_raw_fd(), >> + events: libc::POLLPRI | libc::POLLERR, >> + revents: 0, >> + }; >> + // SAFETY: FFI call, valid arguments, fds contains 1 elem= ent >> + if unsafe { libc::poll(&raw mut fds, 1, -1) } !=3D 1 { >> + panic!("poll failed"); >> + } >> + } >> + Ok(()) >> + } >> + >> + pub fn purge_child(&mut self, path: &Path) -> Result<(), String> = { >> + assert_single_component(path); >> + // See if we can just delete the child directly. >> + match rustix::fs::unlinkat(&self, path, AtFlags::REMOVEDIR) {= >> + // If the cgroup was successfully deleted, or if it >> + // has already been deleted, we are done. >> + Ok(()) | Err(Errno::NOENT) =3D> return Ok(()), >> + // If this cgroup is in use, keep going. >> + Err(Errno::BUSY) =3D> {} >> + Err(e) =3D> return Err(format!("Cannot purge {path:?}: {e= }")), >> + } >> + >> + let sub_fd =3D match openat2_simple(&self, path, OpenFlags::D= irectory) { >> + Ok(sub_fd) =3D> sub_fd, >> + Err(Errno::NOENT) =3D> return Ok(()), >> + Err(e) =3D> { >> + return Err(format!("Cannot open sub-cgroup {path:?}: = {e}",)); >> + } >> + }; >> + >> + // Take an exclusive lock on the cgroup that is about to be r= emoved. This >> + // avoids concurrent executions of this program operating on = deleted >> + // sub-cgroups. Dir::new() doesn't expose a reference to its= internal FD >> + // so it must be delayed until later. >=20 > Yes it does? It's Dir::fd. You used it elsewhere already. It's fine > to delay Dir::new but this comment is not correct. Will delete. I think I missed this because my IDE didn't include in its completions. >> + rustix::fs::flock(&sub_fd, FlockOperation::LockExclusive) >> + .map_err(|e| format!("Cannot lock sub-cgroup: {e}"))?; >> + >> + // Kill all processes in the child cgroup. >> + write_value(&sub_fd, Path::new("cgroup.kill"), b"1")?; >> + >> + // Wait for the child cgroup to become empty. >> + Self::wait_for_empty(&sub_fd) >> + .map_err(|e| format!("Cannot wait for cgroup to become em= pty: {e}")) >> + .inspect_err(|_| { >> + self.fd.pop().unwrap(); >=20 > Why do we need to do this? What's the matching push? Why should > failing to wait for a child cgroup to be empty mean we unlock its > parent? We definitely do not need to do it. It's stale code from when this function did a lot of pushes and pops. >> + })?; >> + >> + // Remove the child cgroup and its contents recursively. >> + remove_recursively(Dir::new(sub_fd).unwrap(), 1000) >> + .map_err(|e| format!("Cannot remove: {e}"))?; >> + >> + // Delete the cgroup. If it's been re-created in the meantim= e and is >> + // currently in use, this is not an error. Another process d= eleting the >> + // cgroup is also not an error. Both of these can happen bec= ause of the >> + // time period between remove_child_directories() closing the= file >> + // descriptor (releasing its lock) and the above call to floc= k(). >> + match rustix::fs::unlinkat(&self, path, AtFlags::REMOVEDIR) {= >> + Ok(()) | Err(Errno::BUSY) | Err(Errno::NOENT) =3D> Ok(())= , >> + Err(e) =3D> Err(format!("Cannot delete: {e}")), >> + } >> + } >> +} >> diff --git a/tools/cgroup-setup/src/main.rs b/tools/cgroup-setup/src/m= ain.rs >> new file mode 100644 >> index 0000000000000000000000000000000000000000..c757a4ad37c812ef5ce524= 9dc1bc3104ec246eef >> --- /dev/null >> +++ b/tools/cgroup-setup/src/main.rs >> @@ -0,0 +1,168 @@ >> +// SPDX-FileCopyrightText: 2026 Demi Marie Obenour >> +// SPDX-License-Identifier: EUPL-1.2+ >> + >> +mod cgroup; >> + >> +use cgroup::{Cgroup, OpenFlags, openat2_simple, write_value}; >> +use rustix::{ >> + fs::{FlockOperation, Mode, XattrFlags}, >> + io::Errno, >> +}; >> +use std::{ >> + env::ArgsOs, >> + fs::File, >> + io::Read as _, >> + os::unix::prelude::*, >> + path::{Path, PathBuf}, >> +}; >> + >> +fn enable_subtree_control(fd: &dyn AsFd) -> Result<(), String> { >> + rustix::fs::fsetxattr(fd, c"user.delegate", b"1", XattrFlags::emp= ty()) >> + .map_err(|e| format!("Cannot enable cgroup delegation: {e}"))= ?; >> + let mut buf =3D Vec::new(); >> + File::from( >> + openat2_simple(fd, c"cgroup.controllers", OpenFlags::Read) >> + .map_err(|e| format!("Cannot open cgroup.controllers: {e}= "))?, >> + ) >> + .read_to_end(&mut buf) >> + .map_err(|e| format!("Cannot read cgroup.controllers: {e}"))?; >> + let mut subtree =3D vec![]; >> + if buf.is_empty() { >> + return Ok(()); >> + } >> + for controller in buf.split(|&b| b =3D=3D b' ') { >> + if !subtree.is_empty() { >> + subtree.push(b' '); >> + } >> + subtree.push(b'+'); >> + subtree.extend_from_slice(controller); >> + } >> + if !subtree.is_empty() { >> + write_value(&fd, Path::new("cgroup.subtree_control"), &subtre= e)?; >> + } >> + Ok(()) >> +} >=20 > My memory of our conversation on a call last week is that we found it > undesirable to enable every controller, since that causes behaviour > surprising action-at-a-distance behaviour changes. Rather specific > requested controllers should be enabled when necessary, right? Yup! I'll move this to the next patch series that enables limits. >> + >> +fn spawn_in_cgroup( >> + mut args: std::iter::Peekable, >> + cgroup: Option<&dyn AsFd>, >> +) -> Result<(), String> { >=20 > We're not spawning anything if all we're doing is an exec. It should b= e > called exec_in_cgroup. Will fix. >> + let Some(program_name) =3D args.next() else { >> + return Ok(()); >> + }; >> + if let Some(cgroup) =3D cgroup { >> + let pid =3D std::process::id().to_string(); >> + write_value( >> + cgroup, >> + Path::new("$inner.service/cgroup.procs"), >> + pid.as_bytes(), >> + ) >> + .map_err(|e| format!("Cannot move process to child cgroup: {e= }"))?; >> + } >> + let e =3D std::process::Command::new(&program_name).args(args).ex= ec(); >> + Err(format!("Cannot spawn child {program_name:?}: {e}")) >=20 > Cannot *exec*. Will fix. >> +} >> + >> +// Check that the path is canonical, >> +// then split it into basename and filename. >> +fn split_path(path: &Path) -> Result<(&Path, &Path), String> { >> + cgroup::check_path(path)?; >> + Ok((path.parent().unwrap(), Path::new(path.file_name().unwrap()))= ) >> +} >> + >> +fn cgroup_setup(args: ArgsOs) -> Result<(), String> { >> + let mut wait =3D true; >> + let mut args =3D args.peekable(); >> + while let Some(arg) =3D args.peek() { >> + if !arg.as_bytes().starts_with(b"-") { >> + break; >> + } >> + let arg =3D args.next().unwrap(); >=20 > I'd find let _ =3D args.next() slightly clearer, because then it's clea= r > we're not interested in the value, since we already have it. That fails to compile (args mutably borrowed more than once). >> + let Some(option) =3D arg.as_bytes().strip_prefix(b"--") else = { >> + return Err("takes no short options".to_owned()); >> + }; >> + match option { >> + b"" =3D> break, >> + b"no-wait" =3D> wait =3D false, >> + _ =3D> return Err(format!("unknown long option {arg:?}"))= , >> + } >> + } >> + let Some(cgroup_path) =3D args.next().map(PathBuf::from) else { >> + return Err("have no positional arguments, expected at least 1= ".to_owned()); >> + }; >> + >> + let (parent_cgroup_path, child_cgroup_path) =3D split_path(&cgrou= p_path)?; >> + let cgroup =3D Cgroup::new(parent_cgroup_path)?; >> + match rustix::fs::mkdirat(&cgroup, child_cgroup_path, Mode::from_= raw_mode(0o755)) { >> + Ok(()) | Err(Errno::EXIST) =3D> {} >> + Err(e) =3D> { >> + return Err(format!( >> + "Cannot create child cgroup {child_cgroup_path:?}: {e= }" >> + )); >> + } >> + } >> + >> + let child =3D openat2_simple(&cgroup, child_cgroup_path, OpenFlag= s::Directory) >> + .map_err(|e| format!("Cannot open child cgroup: {e}"))?; >> + >> + // While waiting, hold an exclusive lock on the child. >> + // This avoids two processes both waiting for the same cgroup to = become >> + // empty, then spawning processes in the same cgroup. >> + rustix::fs::flock(&child, FlockOperation::LockExclusive) >> + .map_err(|e| format!("Cannot take an exclusive lock on child = cgroup: {e}"))?; >> + if wait { >> + Cgroup::wait_for_empty(&child) >> + .map_err(|e| format!("Cannot wait for {parent_cgroup_path= :?} to be empty: {e}"))?; >> + } >> + >> + // Spectrum's programs (such as this one) expect cgroup.subtree_c= ontrol >> + // to be set by the program that created the cgroup. systemd-awa= re >> + // programs, like systemd-udevd, expect user.delegate=3D1 to be s= et. >> + enable_subtree_control(&child)?; >> + >> + // If the child process will need to manage cgroups itself, it wi= ll need >> + // to set up a sub-cgroup due to the "no internal processes" rule= =2E It's >> + // simplest to just do it automatically. If the cgroup already e= xists, >> + // that isn't an error. >> + match rustix::fs::mkdirat(&child, cgroup::DEFAULT_LEAF, Mode::fro= m_raw_mode(0o755)) { >> + Ok(()) | Err(Errno::EXIST) =3D> {} >> + Err(e) =3D> return Err(format!("Cannot create $inner.service = cgroup: {e}")), >> + } >> + >> + spawn_in_cgroup(args, Some(&child)) >> +} >> + >> +fn cgroup_purge(mut args: ArgsOs) -> Result<(), String> { >> + if args.len() !=3D 1 { >> + return Err("usage: cgroup-purge CGROUP_TO_PURGE".to_owned());= >> + } >> + let arg =3D args.next().unwrap(); >> + let (parent, child) =3D split_path(Path::new(&arg))?; >> + Cgroup::new(parent)?.purge_child(child) >> +} >> + >> +fn run(prog_name: &Path, args: ArgsOs) -> Result<(), String> { >> + match prog_name.file_name().map(|f| f.as_bytes()) { >> + Some(b"cgroup-setup") =3D> cgroup_setup(args), >> + Some(b"cgroup-purge") =3D> cgroup_purge(args), >> + _ =3D> Err(format!( >> + "must be invoked as \"cgroup-setup\" or \ >> + \"cgroup-purge\", got {prog_name:?}", >> + )), >> + } >> +} >> + >> +fn main() { >> + let mut args =3D std::env::args_os(); >> + let Some(prog_name) =3D args.next() else { >> + eprintln!("No command line arguments (argv[0] is NULL)"); >> + std::process::exit(1); >> + }; >> + match run(Path::new(&prog_name), args) { >> + Ok(()) =3D> {} >> + Err(e) =3D> { >> + eprintln!("{prog_name:?}: {}", e); >> + std::process::exit(1); >> + } >> + } >> +} --=20 Sincerely, Demi Marie Obenour (she/her/hers) --------------V404WttBfSpXI9pQnNBJbV8X-- --------------2WLnCRpR0JNoIko5fBqrA06A Content-Type: application/pgp-signature; name="OpenPGP_signature.asc" Content-Description: OpenPGP digital signature Content-Disposition: attachment; filename="OpenPGP_signature.asc" -----BEGIN PGP SIGNATURE----- iQIzBAEBCgAdFiEEopQtqVJW1aeuo9/sszaHOrMp8lMFAmqTm1MACgkQszaHOrMp 8lN/RBAAgoZ3uHBY+Jyav5D4Yf4BuovtP2+x+KZaGdo+2kDlotiwrpFMzTd94g91 L3w2qhJjvz8HVD436y7w2pnTgW41sp43XJCQgnIwLhdIm1x1G6sAt8dKtwCUjJ55 GyRnbQ6mC/N/eCS77dV4ai6t4815KywfMP3mpkYg9jFHFyZQcRcJxKaGBKa0LeSO C+hapNnrKWWUFEM8bKWK7F+7oAJJcGgLMoWqn1GBAgCrjxz1POq4K6/pySmbhIv9 Ccg5poQC1DI+eh69cWYUZjlOgm5J/u60dMs7xnu7Kag2Wf2Q0VUoJSrWVG/3IY2x hb00MmOlBc1PdE679EevllAsWNRwbdnOlJ7Uv0QS43XqTpI4p/pMfq51FhDZWzkE a8a40LkabXJSdYVBT+vVVitLapIdpsniEFwTSEKGVKXycKUuGxZH8dLtxFwouI02 KCidMX1HjtBEfChUFZCMEdGxNAEJkzcjkdRQxGLXhfABScLu+wV57jj3ykhSmkHi mvVjA5LYL2OXGnJMOXYVB4fKIysW4344b9zxoDLiH+X+fnCKJbGSQ47JcL1A9Z6G 7zq/VIQzkkNYA+CNmCWoDpvfB430XdH+nR5G5/2GM3lzZXyz6gn7oRhmjv6rv1zf PcG4AoLHIKjzFGsM8RH3SolYOupsbpllq0MacS5jKVCAEa2HodU= =B3N8 -----END PGP SIGNATURE----- --------------2WLnCRpR0JNoIko5fBqrA06A--