2026-07-01 12:42:38 hello, what can we do to unlock build-edge-riscv64 ? 2026-07-01 12:47:14 Find stable, fast and available rv64 hw 2026-07-01 12:47:24 oh, and affordable 2026-07-01 12:49:21 The builders are locking up / freezing 2026-07-01 13:04:03 i have ordered k3 machine(s). not sure if/when they arrive 2026-07-01 13:04:21 I guess you want to own a server, not renting one at providers like Scaleway, right? 2026-07-01 13:04:47 how many rv64 sbcs did you order so far? 2026-07-01 13:15:50 raspbeguy: we have 2 servers on scaleway, but they are used for CI 2026-07-01 13:16:01 ok 2026-07-01 13:21:05 2x K3. one from sipeed.com and one milk-v jupiter 2 2026-07-01 13:23:11 i also have a banapi bpi-f3 machine on my desk here, as CI, and a hifive pioneer p550 which is the current build-edge-riscv64 and build-3-20-riscv64. I think CI is disabled on the p550 for now 2026-07-01 13:25:20 i also have a stafive sifive 2 as a router, and an orangepi rv2 that is not set up yet 2026-07-01 13:25:50 oh also a starpro64, which also is waiting for a kernel 2026-07-02 05:01:30 great 2026-07-02 20:52:27 smth is borked with gitlab ci aarch64: https://gitlab.alpinelinux.org/dne/aports/-/jobs/2419937 2026-07-02 21:20:52 dne: I added a new runner to help with the load, but it needs to be tuned a abit 2026-07-02 21:44:29 ikke: Is there anything more I should do with https://gitlab.alpinelinux.org/alpine/infra/mirrors/-/merge_requests/13? 2026-07-02 21:45:39 No, looks good 2026-07-02 21:45:50 I've added the mirror to our monitoring to see how stable it was 2026-07-02 21:47:19 dne: I've paused the runner for now 2026-07-02 21:47:48 Thanks. Appreciated. 2026-07-03 05:06:16 fyi, I've enabled the runner again, it should work better now, but please let me know if it still has issues 2026-07-03 07:01:58 ikke: I see, seems better now 2026-07-06 08:37:08 Hey! Quick question: Is it possible that the riscv64 builder skipped the latest build of 'baikal' from 2026-06-23? 2026-07-06 08:38:27 I can see that e.g. 'rust' has been built for 'risc' yesterday, 2026-07-06 (triggered for all architectures originally on 2026-07-03), so rust was triggered after 'baikal' has been 2026-07-06 08:38:54 Are perhaps 'community' and 'main' aports prefered over 'testing' aports or has there been a skip? 2026-07-06 08:39:08 see https://pkgs.alpinelinux.org/packages?name=rust&branch=edge&repo=&arch=&origin=&flagged=&maintainer= 2026-07-06 08:39:22 sorry, that's https://pkgs.alpinelinux.org/packages?name=baikal&branch=edge&repo=&arch=&origin=&flagged=&maintainer= 2026-07-06 08:40:43 Yeah the builder build main, push the repo, community, push the repo, and testing and push the repo. currently its on its way building community (see https://build.alpinelinux.org/), so it didnt got to baikal yet, but rust is already pushed 2026-07-06 08:41:58 I see :) Then just patience is the way. Thanks for confirming! 2026-07-06 09:37:02 we have a problem with the riscv64 builder. new builds come in faster than the builder manages to build 2026-07-06 09:37:29 maybe I should move it back to the pioneer machine, now that 3.24 is done 2026-07-06 10:02:40 I think that would help 2026-07-06 10:26:33 BTW, thought I'd mention, I sometimes get in an infinite loop with go-away 2026-07-06 10:30:07 f_: hmm, on gitlab? 2026-07-06 10:31:04 yeah 2026-07-06 10:31:27 Not sure it's something you can do much about, this isn't exclusive to alpine gitlab 2026-07-06 10:31:45 (e.g. I sometimes get that behavior on fdo..) 2026-07-07 07:19:58 Hmm 2026-07-07 07:20:28 Server is slighly overloaded 2026-07-07 07:21:29 Yeah very much noticeable 😅 2026-07-07 07:48:31 Ok, should be better now 2026-07-07 07:48:51 lots of requests from python-requests/2.34.2 from amazon 2026-07-07 08:04:45 Smells like LLM training 2026-07-07 14:38:37 ikke: just block that 2026-07-07 14:38:46 many websites block it for these reasons 2026-07-07 14:39:02 Legitimate software will set their UA correctly anyway 2026-07-07 14:39:18 f_: I have 2026-07-07 14:40:02 good 2026-07-07 14:40:15 blocked up to 80 r/s at some point 2026-07-07 14:40:20 woah 2026-07-08 09:22:35 More crawlers... 2026-07-08 09:22:42 Need to send abuse mails to AWS 2026-07-08 09:28:31 maybe send an invoice instead ;) 2026-07-08 09:28:40 If only 2026-07-08 10:47:35 It's so anoying, the load is quite low and easy to handle under normal circumstances. We are forced to make everything a lot more complicated just to handle this abuse 2026-07-08 10:48:31 Server load is below 5 now. While under load, it was 150+ 2026-07-08 15:14:21 and here we go again 2026-07-08 15:17:50 I could use some help with improving analysis 2026-07-08 15:17:52 I have an idea 2026-07-08 15:18:23 I have most components already, but need them glued together 2026-07-08 15:22:56 ikke: For what it's worth, we've been suffering from the same thing on the Arch Linux side for some time 😕 Pretty annoying indeed. 2026-07-08 15:23:43 Antiz1: Yeah, can imagine 2026-07-08 15:23:56 Some traffic is easy to identify, sometimes it's harder 2026-07-08 15:23:58 ikke: For what it's worth, crawlers are presenting themselves with a Windows NT or Macintosh User agent 🤷 2026-07-08 15:24:02 On our side that is. 2026-07-08 15:24:06 https://gitlab.archlinux.org/archlinux/infrastructure/-/merge_requests/1217/diffs 2026-07-08 15:24:39 The above seems to work for reducing the load on our side. 2026-07-08 15:24:43 Yeah, I've seen that coming around 2026-07-08 15:24:54 Ok, will apply a similar filter on our go-away setup, thanks! 2026-07-08 15:25:08 We had to do the same for the Arch Wiki as well. 2026-07-08 15:25:16 Yeah, similar here 2026-07-08 15:25:22 ikke: Alright, hopefully it'll help 🤞 2026-07-08 15:25:27 You're welcome! :) 2026-07-08 15:26:10 That looks like a smart way to bump the filter 2026-07-08 15:26:48 ikke: I used https://gitlab.com/anarcat/asncounter/-/blob/main/asncounter.1.md on openmw.org to ban problematic ASN 2026-07-08 15:27:37 I know that Ubuntu is using the https://dustri.org/b/trivial-anti-crawler-with-caddy.html trick for some of its infra 2026-07-08 15:27:42 ^ https://github.com/canonical/ubuntu-autopkgtest-operators/pull/127/changes 2026-07-08 15:27:45 jvoisin: interesting, I have something similar to that, didn't know about asncounter 2026-07-08 15:27:57 at least, what I have was just meant to augment asn data 2026-07-08 15:29:16 we don't have it yet in alpine 2026-07-08 16:04:58 oof, 27 armv7 jobs pending 2026-07-09 04:50:56 wow, if the numbers are correct, it had a small peak of 400+ r/s yesterday 2026-07-09 04:51:01 or tonight 2026-07-09 04:51:17 how are you measuring the numbers? 2026-07-09 04:51:59 I'm reading out go-away metrics 2026-07-09 04:52:19 nginx itself received a peak of 50/s 2026-07-09 04:54:33 https://paste.pictures/fGtMkREsMG.png 2026-07-09 04:55:19 that's something :) 2026-07-09 04:56:44 lotheac: btw, found this yesterday: https://devopsbeast.com/blog/oomkilled-wrong-pod-qos 2026-07-09 04:56:52 And saw that the runner pods have no requests or limits 2026-07-09 04:57:04 So they are the first ones to get evicted 2026-07-09 04:57:22 re: improving analysis that you mentioned... could be a good idea to have horizontally scalable reverse proxies that log all requests (including select headers) and sprinkle some visualization on top. my stack is usually traefik -> victorialogs, grafana 2026-07-09 04:57:25 ikke: heh 2026-07-09 04:57:59 i'm a bit surprised that they do not have requests/limits, i thought we did set those 2026-07-09 04:58:13 I thought so too, but not on the runners themselves 2026-07-09 04:58:21 the job pods have 2026-07-09 04:58:41 Which explains why they are killed so frequently 2026-07-09 04:58:51 right 2026-07-09 04:59:05 i guess i was confused about which gitlab options control which pods 2026-07-09 04:59:46 https://gitlab.alpinelinux.org/alpine/infra/k8s/ci-cplane-1/-/blob/master/kustomize/apps/gitlab-runner/base/small-x86/config.toml like... runners.cpu_request does actually not add cpu reqs to the runner pod? 2026-07-09 05:00:38 No, the config.toml is consumed by the runner 2026-07-09 05:00:50 so it can only apply to jobs it spawns 2026-07-09 05:01:18 The operator controls the runners 2026-07-09 05:02:30 It would have to be set in runner.yaml I believe 2026-07-09 05:02:42 ah, i see what you're saying 2026-07-09 05:03:06 okay, yes, that makes sense 2026-07-09 05:04:47 generally speaking, the idea is that if you have memory requests and limits set on every pod, then it should not be possible for the oom killer to kill random processes (because none of the nodes should have been able to be overprovisioned wrt memory) 2026-07-09 05:05:01 We deployed an aarch64 runner, but because the runner gets killed, it looses track of the build pods, which then linger around, preventing new pods to get scheduled 2026-07-09 05:05:33 Yeah, but the nature of build jobs is that you have to over-provision 2026-07-09 05:05:49 not if you have autoscaling 2026-07-09 05:06:21 but if you do overprovision, then it might be a good idea to dedicate the build jobs to specific nodes 2026-07-09 05:06:31 as in, prevent other pods from scheduling there 2026-07-09 05:06:47 To be honest, with the limits / requests in kubernetes we have more issues then without it on docker 2026-07-09 05:07:37 yeah, you do kinda have to apply the same logic to everything in order for it to work 2026-07-09 05:08:01 On docker, we just let everything do it's thing and it somehow works out 2026-07-09 05:08:27 Only very occasianally with extreme large builds, running 6 times in parallel, it overwhelems one host 2026-07-09 05:08:43 (oh, and that is not even CI) 2026-07-09 05:09:01 if you add a taint to the nodes (say dedicated=build), and a label of the same name, and then add a toleration plus nodeselector correspondingly to the build jobs, then they will only run on those specifid nodes (and nothing else will) 2026-07-09 05:09:29 All our nodes in the CI cluster except for the control plane are build hosts 2026-07-09 05:10:25 well, certainly the infrastructure for creating the build pods are different though :) 2026-07-09 05:10:43 you could also have them be scheduled on the cplane nodes i suppose 2026-07-09 05:11:25 for x86_64 and x86 things seem to be stable right now 2026-07-09 05:11:48 But due to the lack of requests/limits for the runner pods, they easily get OOM killed if something heavy is running 2026-07-09 05:12:10 if the runner pod is only responsible for creating the job pods and interacting with them, that in itself does not need to run on the target architecture, right? 2026-07-09 05:12:24 although i didn't check how exactly gitlab does this... 2026-07-09 05:12:29 correct 2026-07-09 05:13:01 well, then it might make sense to run them on cplane nodes (which probably have some spare memory/cpu anyway) 2026-07-09 05:13:07 not a lot 2026-07-09 05:13:18 I picked the smallest instances available 2026-07-09 05:13:50 i suppose adding request/limit to the runner pods should be the simplest solution. do we have metrics on their usage? 2026-07-09 05:15:36 they look pretty light 2026-07-09 05:15:51 https://tpaste.us/moeg 2026-07-09 05:16:04 Yeah, I don't expect them to use much 2026-07-09 05:16:38 I was trying to find if gitlab recommended anything 2026-07-09 05:17:14 heh Error from server (Forbidden): nodes is forbidden: User "lotheac" cannot list resource "nodes" in API group "" at the cluster scope 2026-07-09 05:18:13 https://tpaste.us/R9oo 2026-07-09 05:18:37 yeah that's not too much spare memory 2026-07-09 05:20:11 On the docker nodes, we have no issues with the runners running on them 2026-07-09 05:21:05 at one of my clients -- though github, not glab -- the idea is that there's very little spare capacity for build CI in the cluster. but, the job pods have tactical nodeselectors and taints on them which allow them only to be scheduled on a specific type of VM. there are normally 0 in the cluster, but if there's a pending build job, cluster-autoscaler will talk to the cloud provider API to create a new node and cloud-init it into the cluster 2026-07-09 05:21:26 and then cluster-autoscaler will delete autoscaled nodes after 10 mins of not having any job run on them 2026-07-09 05:21:42 which is cost-efficient, but a bit complex 2026-07-09 05:22:58 We have a fixed set of nodes that are always-on 2026-07-09 05:23:01 i understand your frustration. kube has many footguns 2026-07-09 05:24:51 how do the docker runners handle scheduling many concurrent jobs? 2026-07-09 05:25:19 We limit the runners to max 2 concurrent builds 2026-07-09 05:25:24 generally with kube one should set mem limit == mem request 2026-07-09 05:26:12 with the 48Gi limit but 24Gi request, the scheduler will happily run as many jobs on the same node at the same time as fit in there, calculated with 24G per job 2026-07-09 05:26:38 yes, that's the idea 2026-07-09 05:26:57 build jobs are not fixed size memory workloads 2026-07-09 05:27:44 if you did the same with docker, i don't see how you would avoid oomkills either 2026-07-09 05:28:23 Sometimes build jobs would get oomkilled 2026-07-09 05:28:35 but not frequently 2026-07-09 05:28:44 and what would prevent dockerd from getting killed? 2026-07-09 05:29:02 Nothing, but the OOM killer would always target gcc or something like that 2026-07-09 05:30:01 But because kubernetes activelly adjusts the oom score based on limits/requests, now the runner becomes #1 target 2026-07-09 05:30:42 right 2026-07-09 05:31:09 adding a priorityclass to the runner would probably help too 2026-07-09 05:32:18 Only against eviction apparently 2026-07-09 05:33:18 yeah. the other part of the story is adding requests/limits 2026-07-09 05:34:44 and also adding monitoring for pods that do get oomkilled... 2026-07-09 05:38:09 One reason we need to set requests so low is to make sure the cluster can actually accept jobs that the runner picks 2026-07-09 05:38:23 if the runner says it can accept a job, but the cluster says there is no space, then the job fails 2026-07-09 05:39:00 right. and it wouldn't know before it tried to schedule the pod 2026-07-09 05:39:06 correct 2026-07-09 05:39:52 So requests has been set to the minimum value to allow 2 jobs, but not 3 jobs 2026-07-09 05:40:17 (per node targetted by a specific runner) 2026-07-09 05:40:49 this might be helpful if you are looking for other pods in the cluster that should probably get limits/requests added: kubectl get pod -A -o custom-columns=NAME:.metadata.name,QOS:.status.qosClass,NODE:.spec.nodeName 2026-07-09 05:41:41 quite some BestEffort pods 2026-07-09 05:41:52 Most all kube-system 2026-07-09 05:42:12 fun :) 2026-07-09 05:43:24 but those most likely have priorityClassName set though 2026-07-09 05:43:29 add PRIO:.spec.priorityClassName 2026-07-09 05:44:07 yes 2026-07-09 05:44:27 system-node|cluster-critical 2026-07-09 05:44:55 yeah, that's probably how *they* avoided being killed 2026-07-09 05:45:05 The problem with the runner getting killed is that it looses track of the pods 2026-07-09 05:45:18 if it would gracefully recover, it would not even be an issue in the first place 2026-07-09 05:45:24 that sounds like an implementation issue 2026-07-09 05:45:31 yes, there is a gitlab issue about it 2026-07-09 05:47:13 if the runner pods supported multiple replicas that would help too 2026-07-09 05:47:21 but i suspect that they do not 2026-07-09 05:49:30 each would act as a separate runner accepting jobs 2026-07-09 05:49:41 right, so no 2026-07-09 05:50:35 Easiest fix is setting reqs+limits 2026-07-09 05:50:53 i suppose we could try to hack together some cleanup thing that runs in an initcontainer for the runner pod or something that looks for orphaned jobs and deletes them :p 2026-07-09 05:51:15 yeah, that helps for getting targeted by OOM killer, but does not help for any other situation in which the runner exits or gets killed 2026-07-09 05:51:19 crashes or exceeding its own mem limit 2026-07-09 05:51:45 https://gitlab.com/gitlab-org/gitlab-runner/-/work_items/27333 2026-07-09 05:52:15 old issue :X 2026-07-09 05:52:24 i'll take a proper look later, gotta run to a meeting now 2026-07-09 05:53:19 Have to go as well, talk to you later 2026-07-09 05:53:23 later! 2026-07-09 13:25:37 The aarch64 builder seems to need a bump 2026-07-09 14:16:54 lotheac: I've now set requests+limits on both the container as the init container, it's now has a QOS for guaranteed 2026-07-09 14:17:03 Sertonix[m]: ftr, I bumped it 2026-07-10 01:03:59 ikke: nice :) 2026-07-10 05:32:02 Wow, received spikes of over 100r/s last night 2026-07-10 05:38:25 Even with the modifications you put in the go-away setup? That's got to be frustrating 2026-07-10 05:38:40 Most have been blocked, no impact 2026-07-10 05:38:49 Oh, nice! 2026-07-10 05:42:24 I'm working on some python scripts to help me try to track mirrors that are not responding, hugely out of date, and eventually to verify that certain files are valid, like any APKINDEX's they have. 2026-07-10 05:43:20 Ok, we do track which mirrors our out of date in our monitoring 2026-07-10 05:44:01 When I have time, I'll give you access to that 2026-07-10 05:44:34 Yeah, I'm just looking to put the results in to a db so I can look it heuristics 2026-07-10 05:45:07 No hurry. You are always so busy 2026-07-10 05:45:28 yes, but I keep being busy if I don't delegate things :-) 2026-07-10 05:46:07 Well, if you think of something else for me, don't hesitate to ask 2026-07-10 05:47:34 Anyway, even if I don't really need to do the scripts, I am improving my python foo 2026-07-10 05:48:22 Class inheritance is much easier for me to understand in python that it ever has been for c++ 2026-07-10 05:48:26 Heh 2026-07-10 05:48:50 Having something to practice with really helps 2026-07-10 05:49:12 That is very true 2026-07-10 05:49:54 I will eventually try to tackle go at some point, but for now python is nice because it is so quick to get moving in 2026-07-10 10:08:53 Is build-edge-ppc64le stuck? 2026-07-10 17:51:26 Sertonix[m]: it was