infrastructure
LOGS
<@nirik:matrix.scrye.com>
15:00:47
!startmeeting Infrastructure (2026-08-27)
<@meetbot:fedora.im>
15:00:50
Meeting started at 2026-08-27 15:00:47 UTC
<@meetbot:fedora.im>
15:00:50
The Meeting name is 'Infrastructure (2026-08-27)'
<@nirik:matrix.scrye.com>
15:00:58
!meetingname infrastructure
<@nirik:matrix.scrye.com>
15:00:58
!chair @nirik:matrix.scrye.com @zlopez:fedora.im @jnsamyak:matrix.org @james:fedora.im @gwmngilfen:fedora.im @patrikp:matrix.org
<@nirik:matrix.scrye.com>
15:00:58
!info Agenda is at: https://board.net/p/fedora-infra
<@nirik:matrix.scrye.com>
15:00:58
!info About our team: https://docs.fedoraproject.org/en-US/cle/
<@nirik:matrix.scrye.com>
15:00:58
!info Fedora Infra documentation: https://docs.fedoraproject.org/en-US/infra
<@nirik:matrix.scrye.com>
15:00:58
!topic Hola y bienvenido
<@meetbot:fedora.im>
15:01:01
The Meeting Name is now infrastructure
<@zlopez:fedora.im>
15:01:35
!hi
<@zodbot:fedora.im>
15:01:37
Michal Konečný: Michal Konecny (zlopez)
<@nirik:matrix.scrye.com>
15:02:05
morning everyone.
<@sheikhlimon:matrix.org>
15:02:22
hello
<@nirik:matrix.scrye.com>
15:02:54
will wait a few for more folks to wander in
<@gwmngilfen:fedora.im>
15:03:43
!hi
<@zodbot:fedora.im>
15:03:49
Gwmngilfen: Greg Sutcliffe (gwmngilfen) - he / him / his
<@gwmngilfen:fedora.im>
15:03:58
oh hey, zodbot likes me this week
<@nirik:matrix.scrye.com>
15:04:24
lucky
<@nirik:matrix.scrye.com>
15:04:30
ok, I guess lets go ahead...
<@nirik:matrix.scrye.com>
15:04:34
!topic New folks introductions
<@nirik:matrix.scrye.com>
15:04:34
!info This is a place where people who are interested in Fedora Infrastructure can introduce themselves.
<@nirik:matrix.scrye.com>
15:04:34
!info Join Guide: https://docs.fedoraproject.org/en-US/infra/join_guide/
<@nirik:matrix.scrye.com>
15:04:38
any new folks today?
<@patrikp:matrix.org>
15:05:14
Hi. 👋
<@nirik:matrix.scrye.com>
15:06:33
ok, seems not.
<@nirik:matrix.scrye.com>
15:06:35
!topic Next chair
<@nirik:matrix.scrye.com>
15:06:35
!info Magic eight ball says:
<@nirik:matrix.scrye.com>
15:06:35
!info chair 2026-09-03 - gwmngilfen
<@nirik:matrix.scrye.com>
15:06:35
!info chair 2026-09-10 - ?
<@nirik:matrix.scrye.com>
15:06:50
anyone want the 10th? or shall we leave it to next time?
<@zlopez:fedora.im>
15:07:54
I can take it
<@nirik:matrix.scrye.com>
15:08:42
excellent!
<@nirik:matrix.scrye.com>
15:08:55
!topic Announcements
<@nirik:matrix.scrye.com>
15:08:55
!info CLE Infra&Releng NA-hours team has a Monday through Thursday 30 minute meeting going through tickets at 1900 UTC in https://matrix.to/#/#meeting-3:fedoraproject.org
<@nirik:matrix.scrye.com>
15:09:15
!info tomorrow (2026-08-28) is a Red Hat recharge day. Many Red Hat folks will be away
<@nirik:matrix.scrye.com>
15:09:37
!info we are in Fedora 45 Beta infrastructure freeze
<@nirik:matrix.scrye.com>
15:09:49
Any other announcements?
<@gwmngilfen:fedora.im>
15:10:32
nothing here
<@nirik:matrix.scrye.com>
15:10:58
ok, on to our fav section...
<@nirik:matrix.scrye.com>
15:11:01
!topic Monitoring discussion [nirik / gwmngilfen]
<@nirik:matrix.scrye.com>
15:11:01
!info https://zabbix.fedoraproject.org (top 100 triggers: https://zabbix.fedoraproject.org/zabbix.php?action=toptriggers.list)
<@nirik:matrix.scrye.com>
15:11:01
!info Go over existing items and fix them.
<@gwmngilfen:fedora.im>
15:11:35
so we've had a few fun ones today for things that look related to the rhel upgrades
<@gwmngilfen:fedora.im>
15:11:50
i think i sorted most of them, there's a bunch of PRs open we can get to at some point
<@gwmngilfen:fedora.im>
15:12:04
i do not know what is going on with db-datanommer
<@gwmngilfen:fedora.im>
15:12:26
it appears to be reporting that zabbix can't connect to it, so I'll try to troubleshoot that later
<@nirik:matrix.scrye.com>
15:12:51
so... it's postgresql had been oom killed yesterday. It was restarted, but then it died again...
<@nirik:matrix.scrye.com>
15:12:58
I gave the host more memory for now.
<@gwmngilfen:fedora.im>
15:13:07
yeah, it has metrics, but the version string has a connection refused, weird
<@gwmngilfen:fedora.im>
15:13:19
nt urgent, but definitely curious
<@gwmngilfen:fedora.im>
15:13:24
not urgent, but definitely curious
<@nirik:matrix.scrye.com>
15:13:35
just the postgresql connection?
<@gwmngilfen:fedora.im>
15:13:56
yeah
<@gwmngilfen:fedora.im>
15:13:59
i'll take a look, np
<@gwmngilfen:fedora.im>
15:14:09
the rest of the recent live issues are accounted for I think
<@nirik:matrix.scrye.com>
15:14:40
I have noticed these from time to time: SSH host bvmhost-x86-05.rdu3.fedoraproject.org not reachable on 22 (and bvmhost-x86-06).
<@nirik:matrix.scrye.com>
15:14:52
It might be that they are under a lot of load (they run buildvm's)
<@gwmngilfen:fedora.im>
15:15:07
true. i wonder if that check should be X fils in a row?
<@gwmngilfen:fedora.im>
15:15:13
true. i wonder if that check should be X fails in a row?
<@gwmngilfen:fedora.im>
15:15:36
its always a tradeoff between how fast do you want to know vs false positives
<@nirik:matrix.scrye.com>
15:15:48
oh, and logserver also had a lot of space taken because the bugzilla toddler was spewing to logs.
<@gwmngilfen:fedora.im>
15:16:14
ok - it just fired again because I'm running the synchtt-logs script atm
<@nirik:matrix.scrye.com>
15:16:18
but it still seems high... so perhaps something else is happening
<@gwmngilfen:fedora.im>
15:16:23
ok - it just fired again because I'm running the sync-http-logs script atm
<@gwmngilfen:fedora.im>
15:16:53
so the dl servers seem under a lot of load
<@sheikhlimon:matrix.org>
15:17:17
I think it was yesterday but i slightly saw openqa giving some kind of how do you say this error 90> or more. i think we increased that to 95% right. am I just seeing it wrong or its a different thing?
<@nirik:matrix.scrye.com>
15:17:49
yeah, the openqa-workers...
<@gwmngilfen:fedora.im>
15:18:01
huh, i thought i rolled that out
<@gwmngilfen:fedora.im>
15:19:32
do you want me to make the ssh check more ribust nirik? say 3 fails in a row?
<@nirik:matrix.scrye.com>
15:20:41
we could, sure...
<@nirik:matrix.scrye.com>
15:20:44
or 2?
<@gwmngilfen:fedora.im>
15:20:58
sure, 2 is probably enough
<@gwmngilfen:fedora.im>
15:21:36
the item runs every 5m, maybe we could run it faster with N fails, then we get a similar window of alert (although with more network traffic,)
<@gwmngilfen:fedora.im>
15:21:53
i'll make a ticket, it's a pretty easy one for someone to pickup
<@gwmngilfen:fedora.im>
15:22:20
i guess the crashloop alerts were the toddlers too?
<@sheikhlimon:matrix.org>
15:22:35
i can do it. I need to do the free disk thing too
<@gwmngilfen:fedora.im>
15:23:20
i'm not seeing anything else scary in there, I don't think
<@gwmngilfen:fedora.im>
15:23:41
oh, `Apache: Service is down` on proxy09 is back :/
<@nirik:matrix.scrye.com>
15:23:42
I need to readd ipa servers to dns/ansible (but that will need a freeze break)
<@nirik:matrix.scrye.com>
15:24:23
humf
<@nirik:matrix.scrye.com>
15:24:55
in top tiggers?
<@nirik:matrix.scrye.com>
15:24:58
I am not seeing it.
<@gwmngilfen:fedora.im>
15:25:31
hmm, looks like a cluster of them on 21st Aug
<@gwmngilfen:fedora.im>
15:25:44
might just be a blip
<@gwmngilfen:fedora.im>
15:26:01
[horrible long zabbix url](https://zabbix.fedoraproject.org/zabbix.php?show=2&name=&acknowledgement_status=0&inventory%5B0%5D%5Bfield%5D=type&inventory%5B0%5D%5Bvalue%5D=&evaltype=0&tags%5B0%5D%5Btag%5D=&tags%5B0%5D%5Boperator%5D=0&tags%5B0%5D%5Bvalue%5D=&show_tags=3&tag_name_format=0&tag_priority=&show_opdata=0&show_timeline=1&filter_name=&filter_show_counter=0&filter_custom_time=0&sort=clock&sortorder=DESC&age_state=0&show_symptoms=0&show_suppressed=0&acknowledged_by_me=0&compact_view=0&details=0&highlight_row=0&action=problem.view&triggerids%5B%5D=56123)
<@nirik:matrix.scrye.com>
15:26:03
there was a large chud of traffic then I think that hit the max connections on proxies.
<@gwmngilfen:fedora.im>
15:26:09
aha
<@nirik:matrix.scrye.com>
15:26:48
or I could be misremembering.
<@gwmngilfen:fedora.im>
15:26:51
oh, the stg ocp cert needs doing if anyone wants to deal with it :)
<@gwmngilfen:fedora.im>
15:27:10
zabbix01.stg.rdu3.fedoraproject.org: api.ocp.stg.fedoraproject.org SSL Certificate expires in 29 days: 18 days left
<@gwmngilfen:fedora.im>
15:27:42
i can sort that next week if no one beats me to it
<@nirik:matrix.scrye.com>
15:27:54
ah, so that 'friday' was last thursday, so it was at the tail end of the mass update/reboot outage...
<@nirik:matrix.scrye.com>
15:29:07
I did a bunch of tweaking on ipsilon01/ipa yesterday. I hope it might mean our auth woes are in the rear view mirror... but time will tell.
<@zlopez:fedora.im>
15:29:47
I opened this to track the issues with IPA https://hackmd.io/_8K_2VCcT_afoTsrUY8a6A
<@nirik:matrix.scrye.com>
15:29:48
It's hard to watch for failed logins... because I often see a failure, then a success and can only guess someone mistyped their password or something.
<@nirik:matrix.scrye.com>
15:30:03
oh good. I can read it now. ;)
<@gwmngilfen:fedora.im>
15:30:06
oh, so i found this ...
<@gwmngilfen:fedora.im>
15:30:20
`openqa-x86-worker03.rdu3.fedoraproject.org` isn't in the openqa_workers group in inventory
<@gwmngilfen:fedora.im>
15:30:26
it jumps from 03 to 05
<@gwmngilfen:fedora.im>
15:30:33
so all the others got that macro set, but not 04
<@nirik:matrix.scrye.com>
15:30:53
some of them are 'staging' (connected to openqa-lab01)...
<@nirik:matrix.scrye.com>
15:31:00
(but not really staging)
<@zlopez:fedora.im>
15:31:19
Do we still want to reach to IPA folks for help? I wanted to first collect everything in the hackmd document before contacting them
<@gwmngilfen:fedora.im>
15:31:23
aha, so I guess we can set that macro in the staging group too. sheikhlimon want to send a PR?
<@nirik:matrix.scrye.com>
15:31:33
openqa_lab_workers should also be in the same group I guess?
<@gwmngilfen:fedora.im>
15:31:45
yeah
<@sheikhlimon:matrix.org>
15:31:54
sure
<@nirik:matrix.scrye.com>
15:31:58
Michal Konečný: yeah. I think so... will look / add to that...
<@gwmngilfen:fedora.im>
15:32:45
there's also a bunch of forgejo alerts that I need to go talk with the forge folks about. the monitoring works, but I;m not sure what action they want us to take (if any yet)
<@gwmngilfen:fedora.im>
15:33:02
otherwise, I think we're done for zabbix?
<@nirik:matrix.scrye.com>
15:33:34
I also see there's a user that tries to login every 15m and fails, since... forever. Perhaps I will mail them
<@nirik:matrix.scrye.com>
15:34:07
yeah. I think so.
<@nirik:matrix.scrye.com>
15:34:43
So, what shall we do next (vote on your choice):
<@nirik:matrix.scrye.com>
15:34:49
1. backlog refinement
<@nirik:matrix.scrye.com>
15:34:58
2. some discussion topic
<@nirik:matrix.scrye.com>
15:35:06
3. go to open floor/end early
<@gwmngilfen:fedora.im>
15:36:04
no strong preference - do we have any discussion topics?
<@nirik:matrix.scrye.com>
15:36:36
nothing planned...
<@nirik:matrix.scrye.com>
15:36:46
so, backlog it is. ;)
<@gwmngilfen:fedora.im>
15:37:54
best escuse
<@nirik:matrix.scrye.com>
15:37:55
!topic Fedora Infra backlog refinement
<@nirik:matrix.scrye.com>
15:37:55
!info Refine oldest tickets on Fedora Infra tracker
<@nirik:matrix.scrye.com>
15:37:55
<@gwmngilfen:fedora.im>
15:37:57
best excuse
<@nirik:matrix.scrye.com>
15:38:23
!ticket 12415
<@zodbot:fedora.im>
15:38:24
**infra/tickets #12415** (https://forge.fedoraproject.org/infra/tickets/issues/12415):**Update the “Common Bugs” footer URI to list the current version, rather than F38.**
<@zodbot:fedora.im>
15:38:24
<@zodbot:fedora.im>
15:38:24
● **Opened:** a year ago by rokejulianlockhart
<@zodbot:fedora.im>
15:38:24
● **Last Updated:** 4 months ago
<@zodbot:fedora.im>
15:38:24
● **Assignee:** Not Assigned
<@nirik:matrix.scrye.com>
15:39:11
I guess I could do this one...
<@nirik:matrix.scrye.com>
15:39:19
or someone else could.
<@nirik:matrix.scrye.com>
15:40:14
It can be done in the package and build in f44-infra...
<@nirik:matrix.scrye.com>
15:40:16
https://koji.fedoraproject.org/koji/buildinfo?buildID=2992828
<@nirik:matrix.scrye.com>
15:40:42
or it could be done upstream and then the package updated.
<@nirik:matrix.scrye.com>
15:41:07
although I am not sure where upstream is
<@gwmngilfen:fedora.im>
15:41:55
i could use the packaging practice but i am wary of saying yes when I have enough on my todo
<@nirik:matrix.scrye.com>
15:42:17
https://gitlab.com/fedora/websites-apps/themes/koji-theme-fedora
<@nirik:matrix.scrye.com>
15:42:39
but that seems out of date.
<@nirik:matrix.scrye.com>
15:42:53
I could just ping ryan again in the ticket?
<@nirik:matrix.scrye.com>
15:43:52
lets do that.
<@nirik:matrix.scrye.com>
15:44:35
!ticket 13288
<@zodbot:fedora.im>
15:44:36
**infra/tickets #13288** (https://forge.fedoraproject.org/infra/tickets/issues/13288):**Review / refresh the "request for resources" process**
<@zodbot:fedora.im>
15:44:36
<@zodbot:fedora.im>
15:44:36
● **Opened:** 4 months ago by adamwill
<@zodbot:fedora.im>
15:44:36
● **Last Updated:** 4 months ago
<@zodbot:fedora.im>
15:44:36
● **Assignee:** Not Assigned
<@nirik:matrix.scrye.com>
15:45:04
so, this is updating docs/process... I think this is good to go... just have not had time to do it.
<@nirik:matrix.scrye.com>
15:45:45
we should move them into docs from the wiki and update them...
<@nirik:matrix.scrye.com>
15:46:08
but I guess we just need to do it? I can try... but not sure if I will get to it.
<@zlopez:fedora.im>
15:46:18
That seems to be a good thing to work on during freeze
<@nirik:matrix.scrye.com>
15:46:28
yeah.
<@gwmngilfen:fedora.im>
15:46:55
yeah
<@nirik:matrix.scrye.com>
15:47:06
well, I can try... no promises tho. ;)
<@gwmngilfen:fedora.im>
15:47:09
i haven't opened the links, is that openshift centric, or more general?
<@nirik:matrix.scrye.com>
15:47:22
it's a more general process from 2009.
<@gwmngilfen:fedora.im>
15:47:27
heh
<@nirik:matrix.scrye.com>
15:47:33
we should definitely make it more openshift centric...
<@nirik:matrix.scrye.com>
15:47:42
and discoverable
<@gwmngilfen:fedora.im>
15:48:07
yeah, for sure. i'm just thinking that I have to do the stg matrix server sometime soon, which might a chance to walk through any docs we write up
<@zlopez:fedora.im>
15:49:13
We should probably update the PDR guidelines as well as we now use the tool Ryan created
<@nirik:matrix.scrye.com>
15:50:33
yeah, and Vít Smolík is currently looking at adding that public-inbox instance in staging too.
<@nirik:matrix.scrye.com>
15:50:34
so perhaps both of you could update it ?
<@nirik:matrix.scrye.com>
15:50:39
as part of that...
<@nirik:matrix.scrye.com>
15:50:50
yes, we definitely need to do that...
<@gwmngilfen:fedora.im>
15:51:07
ok, let me comment/assign
<@gwmngilfen:fedora.im>
15:51:48
done
<@nirik:matrix.scrye.com>
15:51:49
Michal Konečný: Our sop needs some work too now... along with any places that sent people to pagure.io
<@nirik:matrix.scrye.com>
15:52:00
one more?
<@nirik:matrix.scrye.com>
15:52:05
!ticket 13069
<@zodbot:fedora.im>
15:52:07
**infra/tickets #13069** (https://forge.fedoraproject.org/infra/tickets/issues/13069):**Reserving a machine for koji `ci` channel**
<@zodbot:fedora.im>
15:52:07
<@zodbot:fedora.im>
15:52:07
● **Opened:** 7 months ago by lecris
<@zodbot:fedora.im>
15:52:07
● **Last Updated:** 4 months ago
<@zodbot:fedora.im>
15:52:07
● **Assignee:** Not Assigned
<@nirik:matrix.scrye.com>
15:52:35
I'm inclined to just close this one now. We cannot afford dedicated ci machines on some arches...
<@nirik:matrix.scrye.com>
15:53:57
and if we have more it would be less of a problem in general.
<@gwmngilfen:fedora.im>
15:54:14
thats fair, for my limited understanding
<@nirik:matrix.scrye.com>
15:55:03
!topic open floor
<@nirik:matrix.scrye.com>
15:55:09
anything anyone has?
<@sheikhlimon:matrix.org>
15:55:27
i have a question about deployment regarding oraculum. currently needs to be **manually** deployed both to stg and prod. whats the standard for automatic deployment? i looked around other fedora apps and they're mostly automated i believe and 3 branch setup with main branch, stg and prod
<@nirik:matrix.scrye.com>
15:55:31
Michal Konečný: oh, I was wondering if you might be willing to look at the mailman01 upgrade...
<@nirik:matrix.scrye.com>
15:56:18
https://forge.fedoraproject.org/infra/tickets/issues/13519
<@nirik:matrix.scrye.com>
15:56:34
we have a number of apps using s2i...
<@nirik:matrix.scrye.com>
15:56:54
ie, when commits happen to their upstream repos, we build/deploy automatically
<@zlopez:fedora.im>
15:56:57
Oh, mailman01 is not RHEL10 yet?
<@nirik:matrix.scrye.com>
15:57:02
no
<@nirik:matrix.scrye.com>
15:57:15
There's missing packages, not sure how many
<@zlopez:fedora.im>
15:57:24
I can look at that
<@nirik:matrix.scrye.com>
15:57:43
anything else?
<@nirik:matrix.scrye.com>
15:58:12
!endmeeting