백업: DB수집 전체 스냅샷 (공공기관2 정리 전)
공공기관2 작업 중. _temp 몽타주(재생성가능)는 제외. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
commit
df16c98366
1
.claude/scheduled_tasks.lock
Normal file
1
.claude/scheduled_tasks.lock
Normal file
@ -0,0 +1 @@
|
|||||||
|
{"sessionId":"3c6000fc-2299-4a1f-993b-223dcca20883","pid":20156,"acquiredAt":1779975159075}
|
||||||
12
.gitignore
vendored
Normal file
12
.gitignore
vendored
Normal file
@ -0,0 +1,12 @@
|
|||||||
|
# 재생성 가능 임시/대용량 (push 제외)
|
||||||
|
/_temp/
|
||||||
|
작업파일/_temp/
|
||||||
|
**/__pycache__/
|
||||||
|
**/.playwright-mcp/
|
||||||
|
**/_nshots/
|
||||||
|
**/_naudit/
|
||||||
|
**/nshot_*/
|
||||||
|
~$*.xlsx
|
||||||
|
*.pyc
|
||||||
|
*.tmp
|
||||||
|
.DS_Store
|
||||||
35441
.owncloudsync.log
Normal file
35441
.owncloudsync.log
Normal file
File diff suppressed because it is too large
Load Diff
141793
.owncloudsync.log.1
Normal file
141793
.owncloudsync.log.1
Normal file
File diff suppressed because it is too large
Load Diff
BIN
.sync_journal.db
Normal file
BIN
.sync_journal.db
Normal file
Binary file not shown.
BIN
.sync_journal.db-shm
Normal file
BIN
.sync_journal.db-shm
Normal file
Binary file not shown.
BIN
.sync_journal.db-wal
Normal file
BIN
.sync_journal.db-wal
Normal file
Binary file not shown.
38
2단계 안내사항.txt
Normal file
38
2단계 안내사항.txt
Normal file
@ -0,0 +1,38 @@
|
|||||||
|
※ 각자 시트에 게시판명/카테고리(소-세부사항2) 열 추가해주세요 ※
|
||||||
|
|
||||||
|
--------------------------------------------------------------------------------------
|
||||||
|
|
||||||
|
[분배 완료 - 17개 지방자치단체]
|
||||||
|
강원특별자치도/경기도/경상남도/경상북도/광주광역시/대구광역시/대전광역시/부산광역시/서울특별시/세종특별자치시/울산광역시/인천광역시/전라남도/제주특별자치도/충청북도/전북특별자치도/충청남도
|
||||||
|
|
||||||
|
▶ 김재민: 경상남도, 울산광역시, 충청남도
|
||||||
|
(강릉시, 속초시)
|
||||||
|
|
||||||
|
▶ 구소영: 경상북도,서울특별시
|
||||||
|
(고성군,강원도경찰청)
|
||||||
|
|
||||||
|
▶ 심예진: 강원특별자치도, 대전광역시, 세종특별자치시, 제주특별자치도
|
||||||
|
(동해시, 강원특별자치도 교육청)
|
||||||
|
|
||||||
|
▶ 방유정: 경기도, 부산광역시, 전북특별자치도
|
||||||
|
(삼척시, 강원영동병무지청, 강원지방기상청, 강원지방병무청, 강원특별자치도 양구군)
|
||||||
|
|
||||||
|
▶ 황의정: 광주광역시
|
||||||
|
|
||||||
|
▶ 정대원: 대구광역시,충청북도
|
||||||
|
|
||||||
|
▶ 프리랜서1: 인천광역시, 전라남도, 고용노동부, 과학기술정보통신부, 교육부
|
||||||
|
|
||||||
|
--------------------------------------------------------------------------------------
|
||||||
|
✓ 대분류:지방자치단체 - 17개 우선적으로 작성
|
||||||
|
✓ 중분류:부/처/청 - 상단기관만 작성(단, 중분류에 '_기타' 들어가는건 제외)
|
||||||
|
|
||||||
|
① '게시물형태' == '사이트' : '수량'~'마크 하이퍼링크' 공백
|
||||||
|
② '저작물 유형' & '공공누리 연계' 작성 X
|
||||||
|
③ '공공누리 부착' 유형 많을 경우에는 모두 작성
|
||||||
|
ex) 1유형, 4유형
|
||||||
|
④ 게시판에 게시물이 0개인 경우/자물쇠가 걸려있는 경우 '공공누리 부착' == '미부착'
|
||||||
|
⑤ 게시판 및 게시물 동시에 공공누리 마크가 부착되어 있다면 '마크 부착위치' == '게시판'
|
||||||
|
⑥ 페이지에 공공누리 마크가 부착되어 있다면 '마크 부착위치' == '게시물'
|
||||||
|
⑦ 로그인/본인인증 해야 보이는 페이지들은 모두 삭제
|
||||||
|
⑧
|
||||||
5
Desktop.ini
Normal file
5
Desktop.ini
Normal file
@ -0,0 +1,5 @@
|
|||||||
|
[.ShellClassInfo]
|
||||||
|
IconResource=D:\\ownCloud\\owncloud.exe
|
||||||
|
|
||||||
|
[ownCloud]
|
||||||
|
UpdateIcon=true
|
||||||
119
README.md
Normal file
119
README.md
Normal file
@ -0,0 +1,119 @@
|
|||||||
|
# 공공저작물 실태조사 DB수집
|
||||||
|
|
||||||
|
신유형개방지원사업 대상기관(1,160개)의 홈페이지를 사이트맵 단위로 점검하여
|
||||||
|
공공누리 부착 여부·게시물 형태·저작물 유형을 엑셀로 정리하는 데이터 수집 프로젝트.
|
||||||
|
|
||||||
|
루트 디렉터리: `D:\01.프로젝트\DB수집`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. 프로젝트 구성
|
||||||
|
|
||||||
|
```
|
||||||
|
DB수집/
|
||||||
|
├─ ★[아이티앤] 공공저작물 자료조사 매뉴얼자료Ver1.2_260525.hwp # 조사 매뉴얼(v1.2)
|
||||||
|
├─ ★공공저작물 실태조사 조사원 자료 취합본(관리자)_2605212137.xlsx # 관리자 취합본
|
||||||
|
├─ 붙임1_...개방 대상기관(1160개) 실태조사 목록...양식_260518.xlsx # 1,160개 기관 마스터 양식
|
||||||
|
├─ 자료_취합_예시.xlsx # 작성 예시
|
||||||
|
├─ 2단계 안내사항.txt # 2단계 분배·작성 규칙
|
||||||
|
├─ 하이퍼링크활성화.py # K열 URL을 엑셀 하이퍼링크로 변환
|
||||||
|
│
|
||||||
|
├─ 1주차(5.19~5.25)/ # 지자체 2개 + 부처 3개 완료
|
||||||
|
│ ├─ 인천광역시(완료)/ # 분야별 메뉴 .txt + 작업본 xlsx + 크롤러
|
||||||
|
│ ├─ 전라남도(완료)/
|
||||||
|
│ ├─ 고용노동부(완료)/
|
||||||
|
│ ├─ 과학기술정보통신부(완료)/
|
||||||
|
│ ├─ 교육부(완료)/
|
||||||
|
│ └─ 수정/ # 샘플 .xlsm + 심예진 수정본
|
||||||
|
│
|
||||||
|
└─ 2주차/ # 부처 3개 진행
|
||||||
|
├─ 보건복지부/ # test.py = mohw.go.kr 크롤러
|
||||||
|
├─ 성평등가족부/ # test.py = 저작물 8유형 자동 분류기
|
||||||
|
└─ 외교부/ # test.py = /list.do 게시판 건수 수집기
|
||||||
|
```
|
||||||
|
|
||||||
|
각 부처/지자체 폴더의 전형적 구성:
|
||||||
|
- `<기관명>_홈페이지_사이트_*.xlsx` — 단계별 작업본(r1, r2, ... 결과, 완료)
|
||||||
|
- `<기관명>_메뉴.txt` / `메뉴구조.txt` — 사이트맵 메뉴 텍스트 덤프
|
||||||
|
- `*.py` — 해당 기관 사이트 구조에 맞춘 크롤링 스크립트
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. 엑셀 데이터 스키마 (작업 컬럼)
|
||||||
|
|
||||||
|
크롤러 코드 기준 핵심 열:
|
||||||
|
|
||||||
|
| 열 | 번호 | 의미 | 채워지는 방법 |
|
||||||
|
|----|----|------|-------------|
|
||||||
|
| K | 11 | URL 주소 | 수기 입력 (1차 사이트맵 수집) |
|
||||||
|
| L | 12 | 게시판/페이지 구분 | 크롤러가 URL 패턴·DOM으로 자동 판별 |
|
||||||
|
| M | 13 | 게시물 수량 | 게시판일 때 `total`·`totalCount` 등에서 추출 |
|
||||||
|
| N | 14 | 저작물 유형 | 본문 분석 → 어문/이미지/영상/오디오/글꼴/3D/기타/없음 |
|
||||||
|
| O | 15 | 공공누리 부착 유형 | `kogl.or.kr/info/licenseType` 링크에서 1~4유형 추출 |
|
||||||
|
| P | 16 | 마크 부착위치 | 게시판/게시물 |
|
||||||
|
| Q | 17 | 부착 여부 보조 | (스크립트별 상이) |
|
||||||
|
|
||||||
|
데이터는 보통 3행 또는 4행부터 시작(헤더 2행).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. 작성 규칙 (2단계 안내사항 발췌)
|
||||||
|
|
||||||
|
- 대분류는 **지방자치단체 17개 우선**, 중분류는 **부/처/청 상단기관**만 (단, `_기타` 포함 제외).
|
||||||
|
- `게시물형태 == 사이트`: `수량`~`마크 하이퍼링크` 공백 처리.
|
||||||
|
- `저작물 유형` 및 `공공누리 연계`는 작성 X.
|
||||||
|
- `공공누리 부착` 유형이 여러 개이면 모두 기재(예: `1유형, 4유형`).
|
||||||
|
- 게시물 0건 또는 자물쇠 페이지 → `공공누리 부착 = 미부착`.
|
||||||
|
- 게시판·게시물 동시 부착 → `마크 부착위치 = 게시판`.
|
||||||
|
- 페이지에만 부착 → `마크 부착위치 = 게시물`.
|
||||||
|
- 로그인·본인인증 페이지는 전부 삭제.
|
||||||
|
|
||||||
|
분배 현황(2단계 17개 지자체):
|
||||||
|
김재민(경남/울산/충남) · 구소영(경북/서울) · 심예진(강원/대전/세종/제주) ·
|
||||||
|
방유정(경기/부산/전북) · 황의정(광주) · 정대원(대구/충북) ·
|
||||||
|
프리랜서1(인천/전남/고용노동부/과기정통부/교육부).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. 스크립트 카탈로그
|
||||||
|
|
||||||
|
루트 공용:
|
||||||
|
- `하이퍼링크활성화.py` — K열의 http(s) URL을 엑셀 하이퍼링크로 변환 + 파란색 밑줄 스타일 적용. 입력/출력 경로는 스크립트 내부 상수.
|
||||||
|
|
||||||
|
기관별 크롤러(공통 패턴: `requests` + `BeautifulSoup` + `openpyxl`로 기존 서식을 보존하며 셀만 갱신):
|
||||||
|
|
||||||
|
| 위치 | 용도 |
|
||||||
|
|------|------|
|
||||||
|
| `2주차/외교부/test.py` | URL이 `/list.do`로 끝나면 L열에 "게시판" 표시, `<div class="total"><span>` 건수를 M열 기입 |
|
||||||
|
| `2주차/성평등가족부/test.py` | 게시판 목록에서 `fn_selectView(n)` 패턴으로 상세 URL 최대 5개 수집 → 본문(`table.brdView01` 또는 `div#contents`) 분석 → 저작물 8유형 판별 |
|
||||||
|
| `2주차/보건복지부/test.py` | `mohw.go.kr` 페이지 판별, `#totalCount` 또는 `span.total b`로 건수 수집, 상세글 `article.board_view` 본문 영역 분석 |
|
||||||
|
| `1주차/인천광역시(완료)/페이지_마크찾기.py` | `<main class="content">` 내 `kogl.or.kr/info/licenseType{n}` 링크에서 공공누리 유형 추출 → O열 기입 |
|
||||||
|
| `1주차/인천광역시(완료)/저작물유형구분.py` | 저작물 유형 자동 분류 |
|
||||||
|
| `1주차/전라남도(완료)/전라남도_페이지_마크찾기_게시판.py` | 전남 게시판형 마크 탐지 |
|
||||||
|
| `1주차/전라남도(완료)/페이지_마크찾기_페이지.py` | 전남 페이지형 마크 탐지 |
|
||||||
|
| `1주차/고용노동부(완료)/페이지_마크찾기_allinone.py` | 고용부 전 영역 통합 마크 탐지 |
|
||||||
|
|
||||||
|
공통 동작 원칙:
|
||||||
|
- `User-Agent`를 Chrome으로 위장, 0.5~1초 슬리프로 차단 회피.
|
||||||
|
- `openpyxl.load_workbook` → 셀만 수정 → 새 파일명(`_r2`, `_결과`, `_완료` 등)으로 저장.
|
||||||
|
- 본문 영역 셀렉터는 사이트별로 다름(상세 코드 주석 참조).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. 산출물 네이밍 컨벤션
|
||||||
|
|
||||||
|
같은 데이터를 단계적으로 가공하면서 접미사로 버전 표시:
|
||||||
|
|
||||||
|
`{기관명}_홈페이지_사이트_{r1|r2|r3|결과|완료}.xlsx`
|
||||||
|
|
||||||
|
예) `외교부.xlsx → 외교부_업데이트.xlsx → ..._활성화_r4.xlsx → ..._활성화_r5.xlsx → 외교부_완료.xlsx`
|
||||||
|
|
||||||
|
`_완료` 또는 `_완료_ai` 접미사가 붙은 파일이 최종본.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. 실행 환경
|
||||||
|
|
||||||
|
- Python 3 + `openpyxl`, `requests`, `beautifulsoup4`, `urllib3`.
|
||||||
|
- Windows 경로(`D:\01.프로젝트\DB수집`)와 `ownCloud\알바\...` 경로가 스크립트에 혼재 — 실행 전 입출력 경로 확인 필요.
|
||||||
|
- 일부 사이트(성평등가족부 등) SSL 검증 우회(`verify=False`) 사용.
|
||||||
143
Superpowers_사용법.md
Normal file
143
Superpowers_사용법.md
Normal file
@ -0,0 +1,143 @@
|
|||||||
|
# Superpowers 사용법
|
||||||
|
|
||||||
|
> Claude Code용 플러그인. 버전 **5.1.0** (claude-plugins-official)
|
||||||
|
> 제작: Jesse Vincent · https://github.com/obra/superpowers
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. Superpowers가 뭔가
|
||||||
|
|
||||||
|
Superpowers는 **코딩 에이전트에게 "개발 방법론"을 입히는 스킬 모음집**이다.
|
||||||
|
설치하면 Claude가 코드를 짤 때 곧장 코드부터 쓰지 않고,
|
||||||
|
|
||||||
|
1. **무엇을 만들려는지 먼저 캐묻고(brainstorming)** →
|
||||||
|
2. **설계 문서를 보여주고 승인을 받고** →
|
||||||
|
3. **실행 계획서(plan)를 만들고** →
|
||||||
|
4. **TDD(빨강→초록) 기반으로 한 단계씩 구현하고** →
|
||||||
|
5. **검증하고 코드 리뷰까지** 하는
|
||||||
|
|
||||||
|
체계적인 흐름을 자동으로 따른다.
|
||||||
|
|
||||||
|
핵심은 **"스킬이 자동으로 발동된다"**는 것. 사용자가 특별히 명령하지 않아도, Claude가 상황을 보고 알아서 맞는 스킬을 꺼내 쓴다.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. 가장 중요한 사실 — 따로 외울 명령어가 없다
|
||||||
|
|
||||||
|
Superpowers는 슬래시 명령(`/superpowers ...` 같은 것)이 **없다.**
|
||||||
|
대신 **14개의 스킬**로 구성되어 있고, Claude가 대화 맥락에 맞춰 **알아서 호출**한다.
|
||||||
|
|
||||||
|
| 사용자가 하는 일 | Claude가 자동으로 하는 일 |
|
||||||
|
|---|---|
|
||||||
|
| "○○ 기능 만들어줘" | `brainstorming` 발동 → 질문으로 요구사항 정리 |
|
||||||
|
| 설계 승인함 | `writing-plans` 발동 → 실행 계획서 작성 |
|
||||||
|
| "이제 진행해" | `test-driven-development` + `executing-plans` 발동 |
|
||||||
|
| 버그/에러 발생 | `systematic-debugging` 발동 |
|
||||||
|
| "다 됐어/끝났어" | `verification-before-completion` → `requesting-code-review` |
|
||||||
|
|
||||||
|
즉 **평소처럼 한국어로 시키기만 하면 된다.** Claude가 "Using [스킬명] to [목적]" 이라고 알리고 해당 절차를 따른다.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. 14개 스킬 한눈에 보기
|
||||||
|
|
||||||
|
### 🧭 시작 & 설계
|
||||||
|
| 스킬 | 언제 발동 |
|
||||||
|
|---|---|
|
||||||
|
| **using-superpowers** | 모든 대화 시작 시 — 어떤 스킬을 쓸지 판단하는 진입점 |
|
||||||
|
| **brainstorming** | 새 기능·컴포넌트·동작 변경 등 **창작 작업 전 필수**. 의도·요구사항·설계를 먼저 탐색 |
|
||||||
|
| **writing-plans** | 스펙이 정해진 다단계 작업의 **구현 계획서** 작성 |
|
||||||
|
|
||||||
|
### 🔨 구현
|
||||||
|
| 스킬 | 언제 발동 |
|
||||||
|
|---|---|
|
||||||
|
| **test-driven-development** | 모든 기능/버그픽스 구현 시 — 구현 코드 전에 테스트 먼저 |
|
||||||
|
| **executing-plans** | 작성된 계획서를 **별도 세션에서** 체크포인트 단위로 실행 |
|
||||||
|
| **subagent-driven-development** | 계획서의 독립 작업들을 **현재 세션에서** 서브에이전트로 실행 |
|
||||||
|
| **dispatching-parallel-agents** | 의존성 없는 2개 이상 작업을 **병렬** 처리 |
|
||||||
|
| **using-git-worktrees** | 현재 작업공간과 격리된 별도 워크스페이스가 필요할 때 |
|
||||||
|
|
||||||
|
### 🐞 디버깅 & 검증
|
||||||
|
| 스킬 | 언제 발동 |
|
||||||
|
|---|---|
|
||||||
|
| **systematic-debugging** | 버그·테스트 실패·예상 밖 동작 — **고치기 전에** 원인부터 체계적으로 |
|
||||||
|
| **verification-before-completion** | "완료/수정됨/통과" 라고 말하기 전 — 실제로 명령 돌려 **증거** 확인 |
|
||||||
|
|
||||||
|
### 👀 코드 리뷰 & 마무리
|
||||||
|
| 스킬 | 언제 발동 |
|
||||||
|
|---|---|
|
||||||
|
| **requesting-code-review** | 작업 완료·주요 기능 구현·머지 전 검증 |
|
||||||
|
| **receiving-code-review** | 리뷰 피드백 받았을 때 — 맹목적 수용 말고 기술적 검증 |
|
||||||
|
| **finishing-a-development-branch** | 구현 끝 + 테스트 통과 후 머지/PR/정리 결정 |
|
||||||
|
|
||||||
|
### 🛠 메타
|
||||||
|
| 스킬 | 언제 발동 |
|
||||||
|
|---|---|
|
||||||
|
| **writing-skills** | 새 스킬을 만들거나 기존 스킬을 수정할 때 |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. 전형적인 작업 흐름 (예시)
|
||||||
|
|
||||||
|
```
|
||||||
|
나: "사이트맵 수집 결과를 검증하는 작은 스크립트 만들어줘"
|
||||||
|
│
|
||||||
|
▼
|
||||||
|
[brainstorming] Claude가 질문을 하나씩 던짐
|
||||||
|
- 입력 형식? 어떤 검증? 실패 시 동작?
|
||||||
|
- 2~3가지 접근법 + 추천안 제시
|
||||||
|
- 설계를 짧게 정리해 보여주고 승인 요청
|
||||||
|
- docs/superpowers/specs/YYYY-MM-DD-검증스크립트-design.md 로 저장
|
||||||
|
│
|
||||||
|
▼ (내가 "좋아" 승인)
|
||||||
|
[writing-plans] 구현 계획서 작성 → docs/superpowers/plans/ 에 저장
|
||||||
|
│
|
||||||
|
▼ (내가 "진행해")
|
||||||
|
[test-driven-development] 테스트 먼저 작성(빨강) → 구현(초록) → 리팩터
|
||||||
|
[subagent-driven-development] 작업 단위로 서브에이전트가 진행/검토
|
||||||
|
│
|
||||||
|
▼
|
||||||
|
[verification-before-completion] 실제로 돌려서 통과 증거 확보
|
||||||
|
[requesting-code-review] 변경분 리뷰
|
||||||
|
│
|
||||||
|
▼
|
||||||
|
[finishing-a-development-branch] 머지/PR/정리 옵션 제시
|
||||||
|
```
|
||||||
|
|
||||||
|
산출물 저장 위치:
|
||||||
|
- 설계 문서 → `docs/superpowers/specs/`
|
||||||
|
- 계획서 → `docs/superpowers/plans/`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. 실전 팁
|
||||||
|
|
||||||
|
- **그냥 평소대로 시키면 된다.** "○○ 만들어줘"라고 하면 brainstorming부터 시작한다.
|
||||||
|
- **간단한 작업이라도 설계 단계를 거친다.** "이건 너무 간단한데" 싶어도 Claude가 짧게라도 설계를 보여주고 승인을 받는다 — 이게 의도된 동작이다.
|
||||||
|
- **특정 스킬을 직접 부르고 싶으면** 이름을 말하면 된다. 예: "systematic-debugging 스킬 써서 이 에러 봐줘"
|
||||||
|
- **건너뛰고 싶으면 말하면 된다.** 사용자 지시가 스킬보다 우선이다. 예: "설계 단계 생략하고 바로 짜줘", "TDD 말고 그냥 구현해줘" → Claude가 따른다.
|
||||||
|
- **CLAUDE.md 규칙이 항상 최우선.** 우선순위: ① 사용자 지시(CLAUDE.md·직접 요청) → ② Superpowers 스킬 → ③ 기본 동작.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. 설치 / 관리 명령 (Claude Code 터미널)
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# 설치 (이미 설치됨)
|
||||||
|
/plugin install superpowers@claude-plugins-official
|
||||||
|
|
||||||
|
# 적용 (설치 직후 1회)
|
||||||
|
/reload-plugins
|
||||||
|
|
||||||
|
# 플러그인 관리 UI
|
||||||
|
/plugin
|
||||||
|
```
|
||||||
|
|
||||||
|
설치 경로:
|
||||||
|
`~/.claude/plugins/cache/claude-plugins-official/superpowers/5.1.0/`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. 한 줄 요약
|
||||||
|
|
||||||
|
> **"평소처럼 한국어로 작업을 시켜라. Superpowers가 알아서 설계 → 계획 → TDD 구현 → 검증 → 리뷰 흐름을 태운다. 건너뛰고 싶으면 그렇게 말하면 된다."**
|
||||||
10
_avg_result.txt
Normal file
10
_avg_result.txt
Normal file
@ -0,0 +1,10 @@
|
|||||||
|
01_김재민 4463 5 892.6
|
||||||
|
02_심예진 4433 5 886.6
|
||||||
|
03_방유정 4499 7 642.7
|
||||||
|
04_구소영 3234 9 359.3
|
||||||
|
05_정대원 2639 8 329.9
|
||||||
|
06_황의정 3512 6 585.3
|
||||||
|
07_프리랜서 2224 3 741.3
|
||||||
|
08_이희재 1372 5 274.4
|
||||||
|
09_신성범 3677 4 919.2
|
||||||
|
10_신규 0 0 0.0
|
||||||
BIN
_test_buyeo.xlsx
Normal file
BIN
_test_buyeo.xlsx
Normal file
Binary file not shown.
55
_김제_E셸삭제.py
Normal file
55
_김제_E셸삭제.py
Normal file
@ -0,0 +1,55 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""김제시: E가 병합(같은 D,E 연속 ≥2행)인데 그 블록 '첫 행'의 F가 비어있으면 그 행 삭제(중분류 랜딩 셸).
|
||||||
|
F 카테고리 라벨은 자식행이 병합 승계. 사용: python -X utf8 _김제_E셸삭제.py plan | run
|
||||||
|
"""
|
||||||
|
import os, sys, re, io, shutil, importlib.util, warnings
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding='utf-8', line_buffering=True)
|
||||||
|
import openpyxl
|
||||||
|
HERE = os.path.dirname(os.path.abspath(__file__))
|
||||||
|
def _imp(n, p):
|
||||||
|
s = importlib.util.spec_from_file_location(n, p); m = importlib.util.module_from_spec(s); s.loader.exec_module(m); return m
|
||||||
|
dd = _imp('dd', os.path.join(HERE, '_스크립트', '_dedup_all.py'))
|
||||||
|
XP = os.path.join(HERE, '작업파일', '광역_사이트맵', '전북특별자치도', '3.김제시', '전북특별자치도_김제시.xlsx')
|
||||||
|
|
||||||
|
def lab(v):
|
||||||
|
return ' > '.join(str(v.get(c)) for c in range(4, 11) if v.get(c) not in (None, ''))
|
||||||
|
|
||||||
|
def main():
|
||||||
|
mode = sys.argv[1] if len(sys.argv) > 1 else 'plan'
|
||||||
|
wb = openpyxl.load_workbook(XP); ws = wb.active
|
||||||
|
rows = dd.load_flat(ws)
|
||||||
|
# E-block = 같은 (D,E) 연속 run, E 비어있지 않음, run≥2
|
||||||
|
delete = set()
|
||||||
|
n = len(rows); i = 0
|
||||||
|
while i < n:
|
||||||
|
v = rows[i]['vals']
|
||||||
|
E = v.get(5)
|
||||||
|
if E in (None, ''):
|
||||||
|
i += 1; continue
|
||||||
|
j = i
|
||||||
|
while j + 1 < n and rows[j+1]['vals'].get(5) == E and rows[j+1]['vals'].get(4) == v.get(4):
|
||||||
|
j += 1
|
||||||
|
runlen = j - i + 1
|
||||||
|
if runlen >= 2 and rows[i]['vals'].get(6) in (None, ''):
|
||||||
|
delete.add(id(rows[i]))
|
||||||
|
i = j + 1
|
||||||
|
del_rows = [r for r in rows if id(r) in delete]
|
||||||
|
keep = [r for r in rows if id(r) not in delete]
|
||||||
|
print('=== 김제 E셸삭제 plan ===')
|
||||||
|
print(f'기존행 {len(rows)} → {len(keep)} (삭제 {len(del_rows)})')
|
||||||
|
for r in del_rows:
|
||||||
|
v = r['vals']
|
||||||
|
print(f' r{r["src"]}: {lab(v)} | L={v.get(12)} M={v.get(13)} | K={str(v.get(11))[-38:]}')
|
||||||
|
if mode != 'run':
|
||||||
|
return
|
||||||
|
bak = XP.replace('.xlsx', '_backup_E셸전.xlsx')
|
||||||
|
shutil.copy(XP, bak)
|
||||||
|
dd.write_back(ws, keep)
|
||||||
|
try:
|
||||||
|
wb.save(XP); print(f'\n저장완료 {len(rows)}→{len(keep)}행. 백업 {os.path.basename(bak)}')
|
||||||
|
except PermissionError:
|
||||||
|
wb.save(XP.replace('.xlsx', '_LP.xlsx')); print('\n!! 잠김 → _LP')
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
116
_김제_인페이지탭수.py
Normal file
116
_김제_인페이지탭수.py
Normal file
@ -0,0 +1,116 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""김제시 인페이지 탭(#anchor) 수량 → M 기입.
|
||||||
|
페이지 본문의 div[class*=basic_tab] 중 탭 링크가 전부 '#anchor'인 그룹(예 basic_tab2>ul.col4,
|
||||||
|
#tab1~#tabN)을 인페이지 스크롤탭으로 보고 M=탭수 기입(공주/금산 전례, 매뉴얼 1-5b: #anchor는 행 분리 안함).
|
||||||
|
menuCd 형제 nav(basic_tab depth4 등, href=실제 .gimje)는 제외. L은 페이지 유지.
|
||||||
|
대상: 17행~끝(요청). 3~16행은 참고 보고만.
|
||||||
|
사용: python -X utf8 _김제_인페이지탭수.py dry | run
|
||||||
|
"""
|
||||||
|
import os, sys, re, io, time, shutil
|
||||||
|
from concurrent.futures import ThreadPoolExecutor, as_completed
|
||||||
|
import warnings; warnings.filterwarnings('ignore')
|
||||||
|
sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding='utf-8', line_buffering=True)
|
||||||
|
import requests
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
import openpyxl
|
||||||
|
HERE = os.path.dirname(os.path.abspath(__file__))
|
||||||
|
XP = os.path.join(HERE, '작업파일', '광역_사이트맵', '전북특별자치도', '3.김제시', '전북특별자치도_김제시.xlsx')
|
||||||
|
UA = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) Chrome/120 Safari/537.36'}
|
||||||
|
|
||||||
|
|
||||||
|
def fetch(session, url):
|
||||||
|
r = session.get(url, headers=UA, verify=False, timeout=10, allow_redirects=True)
|
||||||
|
meta = re.search(rb'charset=["\']?\s*([\w-]+)', r.content[:3000], re.I)
|
||||||
|
r.encoding = meta.group(1).decode('ascii', 'ignore') if meta else r.apparent_encoding
|
||||||
|
return BeautifulSoup(r.text, 'html.parser')
|
||||||
|
|
||||||
|
|
||||||
|
BASIC_TAB = re.compile('basic_tab')
|
||||||
|
|
||||||
|
def inpage_tabs(soup):
|
||||||
|
"""가장 큰 '전부 #anchor' basic_tab 그룹의 (탭수, 라벨)."""
|
||||||
|
best, labels = 0, None
|
||||||
|
for d in soup.find_all('div', class_=BASIC_TAB):
|
||||||
|
ul = d.find('ul')
|
||||||
|
if not ul:
|
||||||
|
continue
|
||||||
|
links = ul.select('li > a')
|
||||||
|
if len(links) < 2:
|
||||||
|
continue
|
||||||
|
hrefs = [(a.get('href') or '').strip() for a in links]
|
||||||
|
if all(h.startswith('#') for h in hrefs):
|
||||||
|
if len(links) > best:
|
||||||
|
best = len(links)
|
||||||
|
labels = [a.get_text(strip=True) for a in links]
|
||||||
|
return best, labels
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
mode = sys.argv[1] if len(sys.argv) > 1 else 'dry'
|
||||||
|
wb = openpyxl.load_workbook(XP); ws = wb.active
|
||||||
|
last = max(r for r in range(3, ws.max_row + 1) if ws.cell(r, 2).value not in (None, ''))
|
||||||
|
|
||||||
|
def lab(r):
|
||||||
|
return ' > '.join(str(ws.cell(r, c).value) for c in range(4, 11) if ws.cell(r, c).value not in (None, ''))
|
||||||
|
|
||||||
|
# 대상 수집: L=페이지 & http URL
|
||||||
|
def targets(lo, hi):
|
||||||
|
out = []
|
||||||
|
for r in range(lo, hi + 1):
|
||||||
|
L = ws.cell(r, 12).value
|
||||||
|
url = ws.cell(r, 11).value
|
||||||
|
if L == '페이지' and isinstance(url, str) and url.startswith('http'):
|
||||||
|
out.append((r, url))
|
||||||
|
return out
|
||||||
|
|
||||||
|
main_t = targets(17, last)
|
||||||
|
pre_t = targets(3, 16)
|
||||||
|
session = requests.Session()
|
||||||
|
|
||||||
|
def scan(rows_urls):
|
||||||
|
res = {}
|
||||||
|
def w(ru):
|
||||||
|
r, u = ru
|
||||||
|
try:
|
||||||
|
n, ls = inpage_tabs(fetch(session, u))
|
||||||
|
return r, n, ls
|
||||||
|
except Exception:
|
||||||
|
return r, 0, None
|
||||||
|
with ThreadPoolExecutor(max_workers=10) as ex:
|
||||||
|
for f in as_completed([ex.submit(w, ru) for ru in rows_urls]):
|
||||||
|
r, n, ls = f.result()
|
||||||
|
if n >= 2:
|
||||||
|
res[r] = (n, ls)
|
||||||
|
return res
|
||||||
|
|
||||||
|
print(f'스캔: 17~{last} ({len(main_t)}개 페이지), 참고 3~16 ({len(pre_t)}개)')
|
||||||
|
t0 = time.time()
|
||||||
|
main_hits = scan(main_t)
|
||||||
|
pre_hits = scan(pre_t)
|
||||||
|
print(f'스캔완료 {time.time()-t0:.0f}s')
|
||||||
|
|
||||||
|
print(f'\n[대상 17~끝] 인페이지탭 발견 {len(main_hits)}행:')
|
||||||
|
for r in sorted(main_hits):
|
||||||
|
n, ls = main_hits[r]
|
||||||
|
old = ws.cell(r, 13).value
|
||||||
|
print(f' r{r}: M {old}→{n} [{lab(r)}] 탭={ls}')
|
||||||
|
if pre_hits:
|
||||||
|
print(f'\n[참고 3~16] 인페이지탭 {len(pre_hits)}행 (요청범위 밖, 미적용):')
|
||||||
|
for r in sorted(pre_hits):
|
||||||
|
n, ls = pre_hits[r]
|
||||||
|
print(f' r{r}: M {ws.cell(r,13).value}→{n}? [{lab(r)}] 탭={ls}')
|
||||||
|
|
||||||
|
if mode != 'run':
|
||||||
|
return
|
||||||
|
bak = XP.replace('.xlsx', '_backup_인페이지탭전.xlsx')
|
||||||
|
shutil.copy(XP, bak)
|
||||||
|
for r, (n, ls) in main_hits.items():
|
||||||
|
ws.cell(r, 13).value = n
|
||||||
|
try:
|
||||||
|
wb.save(XP); print(f'\n저장완료. {len(main_hits)}행 M 갱신. 백업 {os.path.basename(bak)}')
|
||||||
|
except PermissionError:
|
||||||
|
wb.save(XP.replace('.xlsx', '_LP.xlsx')); print('\n!! 잠김 → _LP')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
170
_김제_재분류.py
Normal file
170
_김제_재분류.py
Normal file
@ -0,0 +1,170 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""김제시 재검토: 게시판↔페이지 재분류 + 페이지 인페이지탭(basic_tab) 부모 병합.
|
||||||
|
- 판별: 본문에 게시판 리스트 클래스(bbs_list/news_list/photo_list/video_list/magazine_list/gallery_list/board_list) → 게시판, 없으면 페이지.
|
||||||
|
- 외부 리다이렉트/ERR 행(새창열림 외부시스템)은 손대지 않음.
|
||||||
|
- basic_tab 탭그룹: 페이지 탭만 부모(페이지)로 접고 M=페이지탭수. 게시판 탭은 별도 행 유지.
|
||||||
|
사용: python -X utf8 _김제_재분류.py dry | run
|
||||||
|
"""
|
||||||
|
import os, sys, re, io, hashlib, shutil, importlib.util, warnings, pickle
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding='utf-8')
|
||||||
|
import openpyxl
|
||||||
|
HERE=os.path.dirname(os.path.abspath(__file__))
|
||||||
|
spec=importlib.util.spec_from_file_location('dd', os.path.join(HERE,'_스크립트','_dedup_all.py'))
|
||||||
|
dd=importlib.util.module_from_spec(spec); spec.loader.exec_module(dd)
|
||||||
|
XP=r"작업파일\광역_사이트맵\전북특별자치도\3.김제시\전북특별자치도_김제시.xlsx"
|
||||||
|
CACHE=r"_cache_gimje"
|
||||||
|
LIST=re.compile(r'class="[^"]*(bbs_list|news_list|photo_list|video_list|magazine_list|gallery_list|board_list)[^"]*"')
|
||||||
|
TAB_DIV=re.compile(r'<div class="basic_tab[^"]*">(.*?)</div>', re.S)
|
||||||
|
EXTERNAL={69,249,440,441,148,195,277,318,390,400,134,368} # 외부 리다이렉트/ERR/외부새창 (현행 유지)
|
||||||
|
TYPE_ORDER=['어문','이미지','영상','음악','소프트웨어','데이터','3D','기타']
|
||||||
|
def union_types(vals_list):
|
||||||
|
seen=[]
|
||||||
|
for s in vals_list:
|
||||||
|
if not s: continue
|
||||||
|
for t in str(s).split(','):
|
||||||
|
t=t.strip()
|
||||||
|
if t and t not in seen: seen.append(t)
|
||||||
|
pri=[t for t in TYPE_ORDER if t in seen]+[t for t in seen if t not in TYPE_ORDER]
|
||||||
|
return ','.join(pri) if pri else None
|
||||||
|
|
||||||
|
def mc_of(K):
|
||||||
|
m=re.search(r'menuCd=(DOM_\w+)', str(K)); return m.group(1) if m else None
|
||||||
|
def cpath(url):
|
||||||
|
return os.path.join(CACHE, hashlib.md5(str(url).encode()).hexdigest()+".html")
|
||||||
|
def html_of(K):
|
||||||
|
p=cpath(K); return open(p,encoding='utf-8').read() if os.path.exists(p) else None
|
||||||
|
def tail(mc):
|
||||||
|
m=re.search(r'(\d{18})$', mc); return m.group(1) if m else None
|
||||||
|
def parent_mc(mc):
|
||||||
|
t=tail(mc)
|
||||||
|
if not t: return None
|
||||||
|
g=[t[i:i+3] for i in range(0,18,3)]
|
||||||
|
for i in range(5,-1,-1):
|
||||||
|
if g[i]!='000':
|
||||||
|
g2=g[:]; g2[i]='000'; return mc[:mc.rindex(t)]+''.join(g2)
|
||||||
|
return None
|
||||||
|
|
||||||
|
def main():
|
||||||
|
mode=sys.argv[1] if len(sys.argv)>1 else 'dry'
|
||||||
|
wb=openpyxl.load_workbook(XP); ws=wb.active
|
||||||
|
rows=dd.load_flat(ws) # src 기준
|
||||||
|
by_src={r['src']:r for r in rows}
|
||||||
|
# menuCd 매핑 (src 기준)
|
||||||
|
src_mc={}; mc_src={}
|
||||||
|
for r in rows:
|
||||||
|
mc=mc_of(r['vals'].get(11))
|
||||||
|
if mc: src_mc[r['src']]=mc; mc_src.setdefault(mc,r['src'])
|
||||||
|
# 판별
|
||||||
|
det={}
|
||||||
|
for r in rows:
|
||||||
|
s=r['src']; K=r['vals'].get(11)
|
||||||
|
if s in EXTERNAL or s not in src_mc:
|
||||||
|
det[s]=None; continue
|
||||||
|
h=html_of(K)
|
||||||
|
det[s]='게시판' if (h and LIST.search(h)) else '페이지'
|
||||||
|
# 탭그룹 발굴 → parent별 distinct 탭 menuCd
|
||||||
|
plan={}
|
||||||
|
for r in rows:
|
||||||
|
mc=src_mc.get(r['src']);
|
||||||
|
if not mc: continue
|
||||||
|
h=html_of(r['vals'].get(11))
|
||||||
|
if not h: continue
|
||||||
|
m=TAB_DIV.search(h)
|
||||||
|
if not m: continue
|
||||||
|
tabs=re.findall(r'menuCd=(DOM_\w+)', m.group(1))
|
||||||
|
seen=set(); tabs=[t for t in tabs if not (t in seen or seen.add(t))]
|
||||||
|
if len(tabs)<2: continue
|
||||||
|
par=parent_mc(tabs[0])
|
||||||
|
plan.setdefault(par,set()).update(tabs)
|
||||||
|
# 병합 액션 산출
|
||||||
|
fold_src=set() # 흡수(삭제)될 src
|
||||||
|
survivor_M={} # src -> M(페이지탭수)
|
||||||
|
survivor_N={} # src -> 저작물유형 합집합
|
||||||
|
survivor_clear_leaf=set() # 라벨 leaf 컬럼 비울 survivor
|
||||||
|
actions=[]
|
||||||
|
warn_data=[]; warn_new=[]
|
||||||
|
for par, tabset in plan.items():
|
||||||
|
tab_srcs=[mc_src[t] for t in tabset if t in mc_src]
|
||||||
|
if len(tab_srcs)<2: continue
|
||||||
|
# 폴드 가능한 페이지 탭 = det 페이지 & 외부아님
|
||||||
|
page_tabs=[s for s in tab_srcs if det.get(s)=='페이지' and s not in EXTERNAL]
|
||||||
|
board_tabs=[s for s in tab_srcs if not (det.get(s)=='페이지' and s not in EXTERNAL)]
|
||||||
|
if len(page_tabs)<1:
|
||||||
|
continue
|
||||||
|
par_src=mc_src.get(par)
|
||||||
|
if par_src and det.get(par_src)=='페이지' and par_src not in tab_srcs:
|
||||||
|
survivor=par_src; members=page_tabs
|
||||||
|
else:
|
||||||
|
survivor=min(page_tabs); members=[s for s in page_tabs if s!=survivor]
|
||||||
|
survivor_clear_leaf.add(survivor)
|
||||||
|
if len(members)<1:
|
||||||
|
continue
|
||||||
|
for s in members: fold_src.add(s)
|
||||||
|
survivor_M[survivor]=len(page_tabs)
|
||||||
|
survivor_N[survivor]=union_types([by_src[survivor]['vals'].get(14)]+[by_src[s]['vals'].get(14) for s in members])
|
||||||
|
actions.append((survivor, members, board_tabs, len(page_tabs)))
|
||||||
|
# 경고: 흡수행에 KOGL/저작물 데이터(N=14,O=15,..S=19) 있나
|
||||||
|
for s in members:
|
||||||
|
v=by_src[s]['vals']
|
||||||
|
if any(v.get(c) not in (None,'') for c in (15,16,17,18)): # O~R 공공누리
|
||||||
|
warn_data.append((s, v.get(7) or v.get(6), {c:v.get(c) for c in (14,15,19)}))
|
||||||
|
lab=' > '.join(str(v.get(c)) for c in range(4,11) if v.get(c))
|
||||||
|
if '새창열림' in lab:
|
||||||
|
warn_new.append((s,lab))
|
||||||
|
# 새 행 구성
|
||||||
|
out=[]
|
||||||
|
reclass={'게시판→페이지':0,'페이지→게시판(미적용/외부)':0,'유지':0}
|
||||||
|
for r in rows:
|
||||||
|
s=r['src']
|
||||||
|
if s in fold_src:
|
||||||
|
continue
|
||||||
|
v=r['vals']
|
||||||
|
oldL=v.get(12)
|
||||||
|
d=det.get(s)
|
||||||
|
# 재분류 (외부/비menuCd 제외)
|
||||||
|
if d in ('게시판','페이지'):
|
||||||
|
if oldL!=d and not (oldL=='페이지' and d=='게시판'):
|
||||||
|
# 페이지→게시판은 외부 제외했으므로 여기 오면 진짜 게시판
|
||||||
|
pass
|
||||||
|
newL=d
|
||||||
|
if oldL=='페이지' and d=='게시판' and s in EXTERNAL:
|
||||||
|
newL=oldL
|
||||||
|
# 통계
|
||||||
|
if oldL=='게시판' and newL=='페이지': reclass['게시판→페이지']+=1
|
||||||
|
v[12]=newL
|
||||||
|
# M 설정
|
||||||
|
if newL=='페이지':
|
||||||
|
v[13]=survivor_M.get(s, 1)
|
||||||
|
# 게시판이면 기존 M 유지
|
||||||
|
# survivor: 저작물유형 합집합 갱신
|
||||||
|
if s in survivor_N and survivor_N[s]:
|
||||||
|
v[14]=survivor_N[s]
|
||||||
|
# survivor 라벨 정리(탭이 survivor가 된 경우 leaf 비움)
|
||||||
|
if s in survivor_clear_leaf:
|
||||||
|
lc=dd.leaf_depth(v)
|
||||||
|
v[lc]=None
|
||||||
|
out.append(r)
|
||||||
|
print(f"=== 김제시 재분류 dry ===")
|
||||||
|
print(f"현재행 {len(rows)} → 병합후 {len(out)} (흡수 {len(fold_src)}행)")
|
||||||
|
print(f"게시판→페이지 재분류: {reclass['게시판→페이지']}건")
|
||||||
|
print(f"병합 그룹 수: {len(actions)}")
|
||||||
|
print(f"\n[경고] 흡수행 중 공공누리(O~R) 데이터 보유: {len(warn_data)}건")
|
||||||
|
for s,nm,d2 in warn_data: print(f" r{s} {nm}: {d2}")
|
||||||
|
print(f"\n[확인] 흡수 대상 '새창열림' 페이지: {len(warn_new)}건")
|
||||||
|
for s,lab in warn_new[:20]: print(f" r{s}: {lab}")
|
||||||
|
# 병합 그룹 요약
|
||||||
|
print("\n=== 병합 그룹 (survivor M=페이지탭수, 게시판탭은 유지) ===")
|
||||||
|
def lab(s):
|
||||||
|
v=by_src[s]['vals']; return ' > '.join(str(v.get(c)) for c in range(4,11) if v.get(c))
|
||||||
|
for survivor,members,boards,m in sorted(actions):
|
||||||
|
print(f" survivor r{survivor} M={m} [{lab(survivor)}] 흡수 {len(members)}행" + (f", 게시판유지 {len(boards)}행" if boards else ""))
|
||||||
|
if mode=='run':
|
||||||
|
bak=XP.replace('.xlsx','_backup_재분류전.xlsx')
|
||||||
|
shutil.copy(XP,bak)
|
||||||
|
dd.write_back(ws,out)
|
||||||
|
try: wb.save(XP); print(f"\n저장완료. 백업:{os.path.basename(bak)}")
|
||||||
|
except PermissionError: wb.save(XP.replace('.xlsx','_LP.xlsx')); print("\n!! 잠김→_LP")
|
||||||
|
|
||||||
|
if __name__=='__main__':
|
||||||
|
main()
|
||||||
93
_김제_정리.py
Normal file
93
_김제_정리.py
Normal file
@ -0,0 +1,93 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""김제시 후처리 2규칙:
|
||||||
|
R1) F가 병합(같은 D,E,F 연속 ≥2행)인데 그 블록 '첫 행'의 G가 비어있으면 그 행 삭제(부모 랜딩 셸).
|
||||||
|
R2) F 또는 G 텍스트 끝이 '새창열림'이면 → L=사이트, M·N·O·P·Q 삭제, 텍스트에서 '새창열림' 제거.
|
||||||
|
사용: python -X utf8 _김제_정리.py plan | run
|
||||||
|
"""
|
||||||
|
import os, sys, re, io, shutil, importlib.util, warnings
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding='utf-8', line_buffering=True)
|
||||||
|
import openpyxl
|
||||||
|
HERE = os.path.dirname(os.path.abspath(__file__))
|
||||||
|
def _imp(n, p):
|
||||||
|
s = importlib.util.spec_from_file_location(n, p); m = importlib.util.module_from_spec(s); s.loader.exec_module(m); return m
|
||||||
|
dd = _imp('dd', os.path.join(HERE, '_스크립트', '_dedup_all.py'))
|
||||||
|
XP = os.path.join(HERE, '작업파일', '광역_사이트맵', '전북특별자치도', '3.김제시', '전북특별자치도_김제시.xlsx')
|
||||||
|
NEWWIN = re.compile(r'\s*새\s*창\s*열림\s*$')
|
||||||
|
|
||||||
|
def lab(v):
|
||||||
|
return ' > '.join(str(v.get(c)) for c in range(4, 11) if v.get(c) not in (None, ''))
|
||||||
|
|
||||||
|
def main():
|
||||||
|
mode = sys.argv[1] if len(sys.argv) > 1 else 'plan'
|
||||||
|
wb = openpyxl.load_workbook(XP); ws = wb.active
|
||||||
|
rows = dd.load_flat(ws)
|
||||||
|
|
||||||
|
# --- R2: 새창열림 — 행의 leaf(가장 깊은 카테고리) 라벨 끝이 새창열림이면 사이트화 ---
|
||||||
|
r2 = []
|
||||||
|
for r in rows:
|
||||||
|
v = r['vals']
|
||||||
|
lc = dd.leaf_depth(v) # 가장 깊은 카테고리 컬럼(E/F/G…)
|
||||||
|
t = v.get(lc)
|
||||||
|
if isinstance(t, str) and NEWWIN.search(t):
|
||||||
|
r2.append(r)
|
||||||
|
# 적용: 사이트화 + M/N/O/P/Q 삭제. 텍스트 strip 은 전체 행·전 카테고리 컬럼에서.
|
||||||
|
for r in r2:
|
||||||
|
v = r['vals']
|
||||||
|
r['_had_pq'] = any(v.get(c) not in (None, '') for c in (16, 17))
|
||||||
|
v[12] = '사이트' # L
|
||||||
|
for c in (13, 14, 15, 16, 17): # M N O P Q
|
||||||
|
v[c] = None
|
||||||
|
# 새창열림 텍스트 제거(모든 행, D~J)
|
||||||
|
for r in rows:
|
||||||
|
v = r['vals']
|
||||||
|
for c in range(4, 11):
|
||||||
|
t = v.get(c)
|
||||||
|
if isinstance(t, str) and NEWWIN.search(t):
|
||||||
|
v[c] = NEWWIN.sub('', t).strip()
|
||||||
|
|
||||||
|
# --- R1: F-block 첫행 G 비면 삭제 ---
|
||||||
|
# F-block = 같은 (D,E,F) 연속 run, F 비어있지 않음, run길이>=2
|
||||||
|
delete = set() # id(row)
|
||||||
|
n = len(rows)
|
||||||
|
i = 0
|
||||||
|
while i < n:
|
||||||
|
v = rows[i]['vals']
|
||||||
|
F = v.get(6)
|
||||||
|
if F in (None, ''):
|
||||||
|
i += 1; continue
|
||||||
|
j = i
|
||||||
|
while j + 1 < n and rows[j+1]['vals'].get(6) == F \
|
||||||
|
and rows[j+1]['vals'].get(4) == v.get(4) \
|
||||||
|
and rows[j+1]['vals'].get(5) == v.get(5):
|
||||||
|
j += 1
|
||||||
|
runlen = j - i + 1
|
||||||
|
if runlen >= 2 and rows[i]['vals'].get(7) in (None, ''):
|
||||||
|
delete.add(id(rows[i]))
|
||||||
|
i = j + 1
|
||||||
|
|
||||||
|
del_rows = [r for r in rows if id(r) in delete]
|
||||||
|
keep = [r for r in rows if id(r) not in delete]
|
||||||
|
|
||||||
|
print('=== 김제 정리 plan ===')
|
||||||
|
print(f'기존행 {len(rows)}')
|
||||||
|
print(f'\n[R2] 새창열림 → 사이트: {len(r2)}건 (그중 P/Q 보유 {sum(1 for r in r2 if r.get("_had_pq"))}건)')
|
||||||
|
for r in r2[:40]:
|
||||||
|
print(f' r{r["src"]}: {lab(r["vals"])}')
|
||||||
|
print(f'\n[R1] F병합 첫행 G빈칸 → 삭제: {len(del_rows)}건')
|
||||||
|
for r in del_rows[:60]:
|
||||||
|
print(f' r{r["src"]}: {lab(r["vals"])} (K={str(r["vals"].get(11))[-40:]})')
|
||||||
|
print(f'\n결과행: {len(rows)} → {len(keep)}')
|
||||||
|
|
||||||
|
if mode != 'run':
|
||||||
|
return
|
||||||
|
bak = XP.replace('.xlsx', '_backup_정리전.xlsx')
|
||||||
|
shutil.copy(XP, bak)
|
||||||
|
dd.write_back(ws, keep)
|
||||||
|
try:
|
||||||
|
wb.save(XP); print(f'\n저장완료 {len(rows)}→{len(keep)}행. 백업 {os.path.basename(bak)}')
|
||||||
|
except PermissionError:
|
||||||
|
alt = XP.replace('.xlsx', '_LP.xlsx'); wb.save(alt); print(f'\n!! 잠김 → {os.path.basename(alt)}')
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
299
_김제_확장.py
Normal file
299
_김제_확장.py
Normal file
@ -0,0 +1,299 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""김제시 누락 하위메뉴 확장 — 공식 사이트맵(전체메뉴) 기준.
|
||||||
|
|
||||||
|
basic_tab 부모병합 등으로 빠진 말단 하위메뉴(각자 menuCd 보유)를 사이트맵 트리에서
|
||||||
|
복원해 부모 아래 올바른 컬럼(D=4+tree_depth)·위치로 삽입한다. 부모행(병합 survivor,
|
||||||
|
M=탭수)은 자식이 분리되므로 자기 페이지로 Phase2~4 재수집(M=1 등). 기존 행은 보존.
|
||||||
|
|
||||||
|
L 판별은 김제 방식(본문 리스트클래스=게시판) + 미디어/KOGL은 전북 phase234 로직 재사용.
|
||||||
|
사용: python -X utf8 _김제_확장.py plan | run
|
||||||
|
"""
|
||||||
|
import os, sys, re, io, time, importlib.util, warnings
|
||||||
|
from urllib.parse import urljoin
|
||||||
|
from copy import copy
|
||||||
|
from concurrent.futures import ThreadPoolExecutor, as_completed
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding='utf-8', line_buffering=True)
|
||||||
|
import requests
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
import openpyxl
|
||||||
|
|
||||||
|
HERE = os.path.dirname(os.path.abspath(__file__))
|
||||||
|
def _imp(name, path):
|
||||||
|
spec = importlib.util.spec_from_file_location(name, path)
|
||||||
|
m = importlib.util.module_from_spec(spec); spec.loader.exec_module(m); return m
|
||||||
|
dd = _imp('dd', os.path.join(HERE, '_스크립트', '_dedup_all.py'))
|
||||||
|
ph = _imp('ph', os.path.join(HERE, '_스크립트', '_jeonbuk_phase234_all.py'))
|
||||||
|
|
||||||
|
XP = os.path.join(HERE, '작업파일', '광역_사이트맵', '전북특별자치도', '3.김제시', '전북특별자치도_김제시.xlsx')
|
||||||
|
BASE = 'https://www.gimje.go.kr'
|
||||||
|
ALLMENU = BASE + '/index.gimje?menuCd=DOM_000000107002000000'
|
||||||
|
LIST = re.compile(r'class="[^"]*(bbs_list|news_list|photo_list|video_list|magazine_list|gallery_list|board_list)[^"]*"')
|
||||||
|
NEWWIN = re.compile(r'\s*새\s*창\s*열림\s*$')
|
||||||
|
UA = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) Chrome/120 Safari/537.36'}
|
||||||
|
|
||||||
|
|
||||||
|
def mc_of(s):
|
||||||
|
m = re.search(r'menuCd=(DOM_\w+)', str(s) or ''); return m.group(1) if m else None
|
||||||
|
|
||||||
|
|
||||||
|
def parse_tree():
|
||||||
|
"""반환: nodes(dict mc->{label,url,depth,parent,children[]}), order(list of mc in DFS)."""
|
||||||
|
r = requests.get(ALLMENU, headers=UA, verify=False, timeout=20)
|
||||||
|
soup = BeautifulSoup(r.content, 'html.parser')
|
||||||
|
smap = soup.select_one('div.sitemap')
|
||||||
|
nodes = {}; order = []
|
||||||
|
|
||||||
|
def add(mc, label, url, depth, parent):
|
||||||
|
label = NEWWIN.sub('', label).strip()
|
||||||
|
if mc in nodes:
|
||||||
|
return
|
||||||
|
nodes[mc] = {'label': label, 'url': url, 'depth': depth,
|
||||||
|
'parent': parent, 'children': []}
|
||||||
|
order.append(mc)
|
||||||
|
if parent and parent in nodes:
|
||||||
|
nodes[parent]['children'].append(mc)
|
||||||
|
|
||||||
|
def walk(ul, depth, parent):
|
||||||
|
for li in ul.find_all('li', recursive=False):
|
||||||
|
a = li.find('a', recursive=False) or li.find('a')
|
||||||
|
if not a:
|
||||||
|
continue
|
||||||
|
mc = mc_of(a.get('href'))
|
||||||
|
if not mc:
|
||||||
|
continue
|
||||||
|
label = a.get_text(' ', strip=True)
|
||||||
|
url = urljoin(BASE, a.get('href'))
|
||||||
|
add(mc, label, url, depth, parent)
|
||||||
|
sub = li.find('ul', recursive=False)
|
||||||
|
if sub:
|
||||||
|
walk(sub, depth + 1, mc)
|
||||||
|
|
||||||
|
for mdiv in smap.find_all('div', recursive=False):
|
||||||
|
head = mdiv.find(['h2', 'h3', 'strong', 'a'])
|
||||||
|
if not head:
|
||||||
|
continue
|
||||||
|
# 대분류 head 의 menuCd (a면) 아니면 라벨만 (컨테이너) — 트리 노드로 등록(depth0)
|
||||||
|
hmc = mc_of(head.get('href')) if head.name == 'a' else None
|
||||||
|
htxt = head.get_text(' ', strip=True)
|
||||||
|
if not hmc:
|
||||||
|
# 컨테이너 라벨용 가짜 mc
|
||||||
|
hmc = 'CAT_' + str(len(nodes))
|
||||||
|
add(hmc, htxt, urljoin(BASE, head.get('href')) if head.name == 'a' and head.get('href') else '', 0, None)
|
||||||
|
topul = mdiv.find('ul')
|
||||||
|
if topul:
|
||||||
|
walk(topul, 1, hmc)
|
||||||
|
return nodes, order
|
||||||
|
|
||||||
|
|
||||||
|
def subtree(nodes, mc):
|
||||||
|
out = set()
|
||||||
|
stack = [mc]
|
||||||
|
while stack:
|
||||||
|
x = stack.pop(); out.add(x)
|
||||||
|
stack.extend(nodes[x]['children'])
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def path_labels(nodes, mc):
|
||||||
|
"""mc 의 조상→자기 라벨 리스트 (depth 순)."""
|
||||||
|
chain = []
|
||||||
|
cur = mc
|
||||||
|
while cur is not None:
|
||||||
|
chain.append(cur)
|
||||||
|
cur = nodes[cur]['parent']
|
||||||
|
chain.reverse()
|
||||||
|
return chain # list of mc by depth
|
||||||
|
|
||||||
|
|
||||||
|
# ---- Phase 2~4 (김제 list-class L + 전북 미디어/KOGL) ----
|
||||||
|
def collect(session, url):
|
||||||
|
out = {'L': '', 'M': '', 'N': '', 'O': '', 'P': '', 'Q': '', 'note': ''}
|
||||||
|
try:
|
||||||
|
r = session.get(url, headers=UA, verify=False, timeout=8, allow_redirects=True)
|
||||||
|
meta = re.search(rb'<meta[^>]*charset=["\']?\s*([\w-]+)', r.content[:4096], re.I)
|
||||||
|
r.encoding = meta.group(1).decode('ascii', 'ignore') if meta else r.apparent_encoding
|
||||||
|
if r.status_code != 200:
|
||||||
|
out['note'] = '접근 실패'; return out
|
||||||
|
html = r.text
|
||||||
|
except Exception:
|
||||||
|
out['note'] = '접근 실패'; return out
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
body = ph.get_body(soup, ph.BODY_SEL)
|
||||||
|
is_board = bool(LIST.search(html))
|
||||||
|
if is_board:
|
||||||
|
out['L'] = '게시판'
|
||||||
|
_, cnt = ph.detect_form(body)
|
||||||
|
out['M'] = cnt
|
||||||
|
else:
|
||||||
|
out['L'] = '페이지'; out['M'] = 1
|
||||||
|
has_img, has_vid, has_txt = ph.detect_media(body)
|
||||||
|
types, q = ph.detect_kogl(body)
|
||||||
|
if is_board:
|
||||||
|
for du in ph.extract_detail_urls(body, url, limit=2):
|
||||||
|
try:
|
||||||
|
dr = session.get(du, headers=UA, verify=False, timeout=8)
|
||||||
|
ds = BeautifulSoup(dr.content, 'html.parser')
|
||||||
|
db = ph.get_body(ds, ph.BODY_SEL)
|
||||||
|
di, dv, dt = ph.detect_media(db)
|
||||||
|
has_img |= di; has_vid |= dv; has_txt |= dt
|
||||||
|
dt_types, dt_q = ph.detect_kogl(db)
|
||||||
|
if dt_types and not types:
|
||||||
|
out['P'] = '게시물'
|
||||||
|
types |= dt_types
|
||||||
|
if dt_q == 'Y':
|
||||||
|
q = 'Y'
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
out['N'] = ph.n_string(has_txt, has_img, has_vid)
|
||||||
|
if types:
|
||||||
|
out['O'] = ','.join(f'{n}유형' for n in sorted(types))
|
||||||
|
out['P'] = out['P'] or '게시판'
|
||||||
|
out['Q'] = q or 'N'
|
||||||
|
else:
|
||||||
|
out['O'] = '미부착'
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
mode = sys.argv[1] if len(sys.argv) > 1 else 'plan'
|
||||||
|
nodes, order = parse_tree()
|
||||||
|
real = {mc: n for mc, n in nodes.items() if not mc.startswith('CAT_')}
|
||||||
|
leaves = [mc for mc in order if not mc.startswith('CAT_') and not nodes[mc]['children']]
|
||||||
|
|
||||||
|
wb = openpyxl.load_workbook(XP); ws = wb.active
|
||||||
|
rows = dd.load_flat(ws)
|
||||||
|
pos = {}; row_by_mc = {}
|
||||||
|
for i, r in enumerate(rows):
|
||||||
|
mc = mc_of(r['vals'].get(11))
|
||||||
|
if mc:
|
||||||
|
pos[mc] = i; row_by_mc[mc] = r
|
||||||
|
have = set(row_by_mc)
|
||||||
|
|
||||||
|
missing = [mc for mc in leaves if mc not in have]
|
||||||
|
# parents that will gain children
|
||||||
|
gain_parents = {}
|
||||||
|
for mc in missing:
|
||||||
|
p = nodes[mc]['parent']
|
||||||
|
gain_parents.setdefault(p, []).append(mc)
|
||||||
|
|
||||||
|
# anchor index for each parent group: after last existing row in subtree of parent
|
||||||
|
inserts = {} # anchor_idx -> [mc,...] in tree order
|
||||||
|
no_anchor = []
|
||||||
|
for p, kids in gain_parents.items():
|
||||||
|
# nearest existing ancestor (incl parent) to anchor after its subtree
|
||||||
|
anc = p
|
||||||
|
anchor_idx = None
|
||||||
|
while anc is not None:
|
||||||
|
st = subtree(nodes, anc)
|
||||||
|
present = [pos[m] for m in st if m in pos]
|
||||||
|
if present:
|
||||||
|
anchor_idx = max(present); break
|
||||||
|
anc = nodes[anc]['parent']
|
||||||
|
if anchor_idx is None:
|
||||||
|
no_anchor.append((p, kids)); continue
|
||||||
|
# order kids by their order in tree (children list of p)
|
||||||
|
ordered = [m for m in nodes[p]['children'] if m in kids]
|
||||||
|
inserts.setdefault(anchor_idx, []).extend(ordered)
|
||||||
|
|
||||||
|
survivors = [p for p in gain_parents if p in have]
|
||||||
|
print('=== 김제 확장 plan ===')
|
||||||
|
print(f'트리 노드(실): {len(real)} 말단leaf: {len(leaves)} 기존행: {len(rows)}')
|
||||||
|
print(f'누락 말단메뉴(추가대상): {len(missing)}')
|
||||||
|
print(f'자식 얻는 부모: {len(gain_parents)} (그중 기존행=재수집대상 survivor: {len(survivors)})')
|
||||||
|
print(f'앵커 못찾음: {len(no_anchor)}')
|
||||||
|
# sample
|
||||||
|
def lab(mc):
|
||||||
|
return ' > '.join(nodes[m]['label'] for m in path_labels(nodes, mc) if not m.startswith('CAT_') or nodes[m]['label'])
|
||||||
|
print('\n[샘플 추가 행 20]')
|
||||||
|
for mc in missing[:20]:
|
||||||
|
d = nodes[mc]['depth']; col = chr(ord('A') + 3 + d)
|
||||||
|
print(f' +{col}({d}) {lab(mc)} ({mc})')
|
||||||
|
print('\n[survivor 부모 재수집 대상]')
|
||||||
|
for p in survivors:
|
||||||
|
r = row_by_mc[p]; v = r['vals']
|
||||||
|
print(f' r{r["src"]} M={v.get(13)} L={v.get(12)} [{nodes[p]["label"]}] 자식 {len(gain_parents[p])}개')
|
||||||
|
if no_anchor:
|
||||||
|
print('\n[!] 앵커 못찾은 그룹:')
|
||||||
|
for p, kids in no_anchor:
|
||||||
|
print(f' parent {p} kids {len(kids)}')
|
||||||
|
|
||||||
|
if mode != 'run':
|
||||||
|
return
|
||||||
|
|
||||||
|
# ---- build new rows + fetch ----
|
||||||
|
session = requests.Session()
|
||||||
|
template = row_by_mc.get(list(survivors)[0]) if survivors else rows[0]
|
||||||
|
tmpl_styles = template['styles']
|
||||||
|
|
||||||
|
def make_row(mc):
|
||||||
|
v = {c: None for c in range(1, dd.MAXCOL + 1)}
|
||||||
|
v[3] = '김제시'
|
||||||
|
chain = path_labels(nodes, mc)
|
||||||
|
for m in chain:
|
||||||
|
d = nodes[m]['depth']
|
||||||
|
col = 4 + d
|
||||||
|
if col <= 10 and nodes[m]['label']:
|
||||||
|
v[col] = nodes[m]['label']
|
||||||
|
v[11] = nodes[mc]['url']
|
||||||
|
return {'src': None, 'vals': v,
|
||||||
|
'styles': {c: tuple(copy(x) if hasattr(x, 'copy') or True else x for x in tmpl_styles[c]) for c in tmpl_styles},
|
||||||
|
'hyperlink': nodes[mc]['url']}
|
||||||
|
|
||||||
|
# fetch all missing + survivor parents
|
||||||
|
fetch_targets = list(missing) + survivors
|
||||||
|
print(f'\nPhase2~4 수집 {len(fetch_targets)}건 ...')
|
||||||
|
res = {}
|
||||||
|
t0 = time.time()
|
||||||
|
def work(mc):
|
||||||
|
return mc, collect(session, nodes[mc]['url'])
|
||||||
|
with ThreadPoolExecutor(max_workers=10) as ex:
|
||||||
|
futs = [ex.submit(work, mc) for mc in fetch_targets]
|
||||||
|
done = 0
|
||||||
|
for f in as_completed(futs):
|
||||||
|
mc, o = f.result(); res[mc] = o; done += 1
|
||||||
|
if done % 30 == 0 or done == len(fetch_targets):
|
||||||
|
print(f' {done}/{len(fetch_targets)} ({time.time()-t0:.0f}s)')
|
||||||
|
|
||||||
|
# apply to survivors (overwrite own page data; M back to own)
|
||||||
|
for p in survivors:
|
||||||
|
v = row_by_mc[p]['vals']; o = res.get(p, {})
|
||||||
|
for col, key in [(12, 'L'), (13, 'M'), (14, 'N'), (15, 'O'), (16, 'P'), (17, 'Q')]:
|
||||||
|
if o.get(key) != '':
|
||||||
|
v[col] = o[key]
|
||||||
|
|
||||||
|
# build new row objects with phase data
|
||||||
|
new_obj = {}
|
||||||
|
for mc in missing:
|
||||||
|
ro = make_row(mc); o = res.get(mc, {})
|
||||||
|
v = ro['vals']
|
||||||
|
for col, key in [(12, 'L'), (13, 'M'), (14, 'N'), (15, 'O'), (16, 'P'), (17, 'Q')]:
|
||||||
|
if o.get(key) != '':
|
||||||
|
v[col] = o[key]
|
||||||
|
if o.get('note'):
|
||||||
|
v[19] = o['note']
|
||||||
|
new_obj[mc] = ro
|
||||||
|
|
||||||
|
# assemble final ordered list
|
||||||
|
out = []
|
||||||
|
for i, r in enumerate(rows):
|
||||||
|
out.append(r)
|
||||||
|
if i in inserts:
|
||||||
|
for mc in inserts[i]:
|
||||||
|
out.append(new_obj[mc])
|
||||||
|
|
||||||
|
# backup + write
|
||||||
|
import shutil
|
||||||
|
bak = XP.replace('.xlsx', '_backup_확장전.xlsx')
|
||||||
|
shutil.copy(XP, bak)
|
||||||
|
dd.write_back(ws, out)
|
||||||
|
try:
|
||||||
|
wb.save(XP)
|
||||||
|
print(f'\n저장완료 {len(rows)}→{len(out)}행. 백업 {os.path.basename(bak)}')
|
||||||
|
except PermissionError:
|
||||||
|
alt = XP.replace('.xlsx', '_LP.xlsx'); wb.save(alt)
|
||||||
|
print(f'\n!! 원본 잠김 → {os.path.basename(alt)} 로 저장')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
53
_스크립트/_br_scan.py
Normal file
53
_스크립트/_br_scan.py
Normal file
@ -0,0 +1,53 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""보령시 210행~끝: E병합 그룹별 F자식 형태/수량 스캔. 합치기 대상 판정."""
|
||||||
|
import openpyxl
|
||||||
|
XLSX = r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\6.보령시\충청남도_보령시.xlsx'
|
||||||
|
ws = openpyxl.load_workbook(XLSX).active
|
||||||
|
|
||||||
|
last = 3
|
||||||
|
for r in range(3, ws.max_row + 1):
|
||||||
|
if ws.cell(r, 11).value:
|
||||||
|
last = r
|
||||||
|
print('마지막 데이터행', last)
|
||||||
|
|
||||||
|
# E 병합 범위(>=210)
|
||||||
|
emerges = {}
|
||||||
|
for mr in ws.merged_cells.ranges:
|
||||||
|
if mr.min_col == 5 and mr.max_row >= 210:
|
||||||
|
emerges[mr.min_row] = (mr.min_row, mr.max_row, ws.cell(mr.min_row, 5).value)
|
||||||
|
|
||||||
|
# 210부터 행 훑기
|
||||||
|
r = 210
|
||||||
|
elig = noelig = 0
|
||||||
|
while r <= last:
|
||||||
|
e = ws.cell(r, 5).value
|
||||||
|
# E 병합?
|
||||||
|
rng = None
|
||||||
|
for mr in ws.merged_cells.ranges:
|
||||||
|
if mr.min_col == 5 and mr.min_row <= r <= mr.max_row:
|
||||||
|
rng = (mr.min_row, mr.max_row); break
|
||||||
|
if rng and rng[1] > rng[0]:
|
||||||
|
r1, r2 = rng
|
||||||
|
elabel = ws.cell(r1, 5).value
|
||||||
|
dlabel = ws.cell(r1, 4).value
|
||||||
|
ls = [ws.cell(x, 12).value for x in range(r1, r2 + 1)]
|
||||||
|
ms = [ws.cell(x, 13).value for x in range(r1, r2 + 1)]
|
||||||
|
fs = [ws.cell(x, 6).value for x in range(r1, r2 + 1)]
|
||||||
|
allpage = all((l in ('페이지', '사이트')) for l in ls)
|
||||||
|
kinds = set(ls)
|
||||||
|
tag = 'O합치기' if allpage else 'X(게시판섞임)'
|
||||||
|
if allpage: elig += 1
|
||||||
|
else: noelig += 1
|
||||||
|
try:
|
||||||
|
ssum = sum(int(m) for m in ms if m is not None)
|
||||||
|
except Exception:
|
||||||
|
ssum = '?'
|
||||||
|
print('E[%d~%d] %s > %s (%d행) L=%s M합=%s %s' % (r1, r2, dlabel, elabel, r2 - r1 + 1, kinds, ssum, tag))
|
||||||
|
if not allpage:
|
||||||
|
for x in range(r1, r2 + 1):
|
||||||
|
print(' - %s L=%s M=%s' % (ws.cell(x, 6).value, ws.cell(x, 12).value, ws.cell(x, 13).value))
|
||||||
|
r = r2 + 1
|
||||||
|
else:
|
||||||
|
# 단일행(E 비병합) 또는 E없음
|
||||||
|
r += 1
|
||||||
|
print('\n합치기대상 E그룹:', elig, '/ 제외(게시판섞임):', noelig)
|
||||||
140
_스크립트/_buyeo_rejudge_N.py
Normal file
140
_스크립트/_buyeo_rejudge_N.py
Normal file
@ -0,0 +1,140 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""부여군 N(저작물 유형) 재판정 — 장식/공통UI 이미지 false '이미지' 제거.
|
||||||
|
- phase234 detect_media 는 KOGL 마크만 제외 → file_icon.gif·see_btn.gif·btn_page_*.gif 등
|
||||||
|
게시판 스킨/공통버튼 아이콘을 '이미지'로 오탐.
|
||||||
|
- [[feedback_N_image_rule]] 적용: 장식/공통UI·페이징버튼·첨부아이콘 제외, 실제 사진/업로드·지도/PDF임베드만 인정.
|
||||||
|
- 대상: 현재 N 에 '이미지' 포함된 같은도메인 행만(false 양성 제거 전용 → 이미지 추가 안 함).
|
||||||
|
- 게시판은 상세글 상위5 추적. body=#txt|#contents|main.
|
||||||
|
사용: python -X utf8 _buyeo_rejudge_N.py [--write]
|
||||||
|
"""
|
||||||
|
import sys, warnings, re
|
||||||
|
from urllib.parse import urljoin
|
||||||
|
import requests, openpyxl
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
S = requests.Session(); S.headers.update({'User-Agent': 'Mozilla/5.0'})
|
||||||
|
XLSX = r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\7.부여군\충청남도_부여군.xlsx'
|
||||||
|
DOMAIN = 'buyeo.go.kr'
|
||||||
|
|
||||||
|
KOGL = re.compile(r'(?:new_)?img_open(?:type|code)\d', re.I)
|
||||||
|
# 장식/공통UI: 기존 논산 DECO + 부여 게시판 스킨/공통버튼/첨부아이콘 패턴 보강
|
||||||
|
DECO = re.compile(
|
||||||
|
r'/site/common/img/|move\.png|no[-_]?img|blank\.|spacer\.|'
|
||||||
|
r'/img/(?:icon|ico|bul|bullet|arrow|btn|bg|tit|h\d)|'
|
||||||
|
r'ico_|btn_|bul_|bg_|_bg\.|icon_|'
|
||||||
|
r'file_icon|_icon\.|icon\.gif|_btn\.|see_btn|' # 첨부아이콘·바로보기버튼
|
||||||
|
r'/skin/|/images/(?:kr/)?common/|/common/img/|' # 게시판 스킨·공통 UI 디렉터리
|
||||||
|
r'btn_page|page_(?:next|prev|first|last)', re.I) # 페이징 버튼
|
||||||
|
MAPPDF = re.compile(r'pdf|viewer\.html|map|kakao|daum|naver.*map|google.*map|/map', re.I)
|
||||||
|
VIDEO = re.compile(r'youtube\.com|youtu\.be|/embed/|\.mp4|\.webm|vimeo', re.I)
|
||||||
|
DETAIL = re.compile(r'mode=V|view\.do|/view|seq=|idx=|mng_no=|bbsSeq|nttId|no=', re.I)
|
||||||
|
|
||||||
|
|
||||||
|
def get(url):
|
||||||
|
r = S.get(url, verify=False, timeout=20)
|
||||||
|
return BeautifulSoup(r.content, 'html.parser')
|
||||||
|
|
||||||
|
|
||||||
|
def body(soup):
|
||||||
|
return soup.select_one('#txt') or soup.select_one('#contents') or soup.select_one('main') or soup
|
||||||
|
|
||||||
|
|
||||||
|
def real_imgs(node):
|
||||||
|
n = 0
|
||||||
|
for im in node.find_all('img'):
|
||||||
|
src = im.get('src') or ''
|
||||||
|
if not src or KOGL.search(src) or DECO.search(src):
|
||||||
|
continue
|
||||||
|
n += 1
|
||||||
|
return n
|
||||||
|
|
||||||
|
|
||||||
|
def media_flags(node):
|
||||||
|
img = real_imgs(node) > 0
|
||||||
|
vid = False
|
||||||
|
for ifr in node.find_all('iframe'):
|
||||||
|
s = ifr.get('src') or ''
|
||||||
|
if VIDEO.search(s):
|
||||||
|
vid = True
|
||||||
|
elif MAPPDF.search(s):
|
||||||
|
img = True
|
||||||
|
for a in node.find_all('a'):
|
||||||
|
if VIDEO.search(a.get('href') or ''):
|
||||||
|
vid = True
|
||||||
|
if node.find('video'):
|
||||||
|
vid = True
|
||||||
|
return img, vid
|
||||||
|
|
||||||
|
|
||||||
|
def detail_links(soup, base):
|
||||||
|
out = []
|
||||||
|
for a in body(soup).find_all('a'):
|
||||||
|
h = a.get('href') or ''
|
||||||
|
if h.startswith('#'):
|
||||||
|
continue
|
||||||
|
if DETAIL.search(h):
|
||||||
|
out.append(urljoin(base, h))
|
||||||
|
seen = set(); res = []
|
||||||
|
for u in out:
|
||||||
|
if u not in seen:
|
||||||
|
seen.add(u); res.append(u)
|
||||||
|
return res[:5]
|
||||||
|
|
||||||
|
|
||||||
|
def judge(k, L, M, N):
|
||||||
|
if not isinstance(k, str) or DOMAIN not in k:
|
||||||
|
return N, 'skip(외부)'
|
||||||
|
try:
|
||||||
|
soup = get(k)
|
||||||
|
except Exception as e:
|
||||||
|
return N, f'fetch실패:{str(e)[:25]}'
|
||||||
|
bd = body(soup)
|
||||||
|
txtlen = len(bd.get_text(strip=True))
|
||||||
|
if L == '게시판' and (M in (0, '0', None)):
|
||||||
|
return '없음', 'board0'
|
||||||
|
has_img, has_vid = media_flags(bd)
|
||||||
|
detail = False
|
||||||
|
if L == '게시판' and not has_img:
|
||||||
|
for du in detail_links(soup, k):
|
||||||
|
try:
|
||||||
|
di, dv = media_flags(body(get(du)))
|
||||||
|
if di: has_img = True
|
||||||
|
if dv: has_vid = True
|
||||||
|
if di or dv: detail = True
|
||||||
|
if has_img and has_vid: break
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
parts = []
|
||||||
|
if txtlen >= 1: parts.append('어문')
|
||||||
|
if has_img: parts.append('이미지')
|
||||||
|
if has_vid: parts.append('영상')
|
||||||
|
return (','.join(parts) if parts else '없음'), ('detail' if detail else '')
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
write = '--write' in sys.argv
|
||||||
|
wb = openpyxl.load_workbook(XLSX); ws = wb.active
|
||||||
|
changes = []
|
||||||
|
for r in range(3, ws.max_row + 1):
|
||||||
|
N = ws.cell(r, 14).value
|
||||||
|
if not (isinstance(N, str) and '이미지' in N):
|
||||||
|
continue
|
||||||
|
k = ws.cell(r, 11).value; L = ws.cell(r, 12).value; M = ws.cell(r, 13).value
|
||||||
|
new, tag = judge(k, L, M, N)
|
||||||
|
if new != N:
|
||||||
|
changes.append((r, N, new, tag, ws.cell(r, 7).value or ws.cell(r, 6).value, k))
|
||||||
|
print(f'현재 N에 이미지 포함 행 재판정 → 변경 {len(changes)}건')
|
||||||
|
for r, old, new, tag, label, k in changes:
|
||||||
|
print(f' 행{r} [{old} → {new}] {tag} | {label} | {str(k)[-45:]}')
|
||||||
|
if write and changes:
|
||||||
|
for r, old, new, tag, label, k in changes:
|
||||||
|
ws.cell(r, 14).value = new
|
||||||
|
wb.save(XLSX)
|
||||||
|
print(f'\n저장: {XLSX} ({len(changes)}건 수정)')
|
||||||
|
elif not write:
|
||||||
|
print('\n(계획만. --write 로 반영)')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
266
_스크립트/_buyeo_rejudge_NO.py
Normal file
266
_스크립트/_buyeo_rejudge_NO.py
Normal file
@ -0,0 +1,266 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""부여군 N(저작물 유형) + O/P/Q(공공누리) 전수 재판정.
|
||||||
|
사용자 검수(3~37행)를 정답으로 검증 후 38행~ 적용.
|
||||||
|
|
||||||
|
핵심(부여 buyeo 템플릿 특성):
|
||||||
|
- KOGL 마크는 게시판 '상세글'의 `div.gnuri_layer > .codeView01 > img[src=/_module/gnuri/images/img_opentypeNN.png]`
|
||||||
|
+ kogl.or.kr/info/licenseTypeN.do 링크. **본문(#txt) 바깥**이라 기존 phase234가 놓침.
|
||||||
|
- 상세 URL = `?mode=V&no=...`.
|
||||||
|
- 게시판은 글마다 유형 다를 수 있음 → O = **첫 마크 글의 유형**(사용자 컨벤션: 01030502 글1=1유형 → O=1유형).
|
||||||
|
- N 이미지: 장식/공통UI(file_icon·see_btn·btn_page·/skin/·/common/·ico/btn/bg) + KOGL 마크 제외.
|
||||||
|
지도/PDF iframe = 이미지. 게시판 M=0 = 없음. 실제 업로드 사진만 이미지.
|
||||||
|
|
||||||
|
사용: python -X utf8 _buyeo_rejudge_NO.py --validate # 3~37 사용자값과 비교
|
||||||
|
python -X utf8 _buyeo_rejudge_NO.py --write # 38행~ 기입(백업)
|
||||||
|
"""
|
||||||
|
import sys, warnings, re, time
|
||||||
|
from urllib.parse import urljoin
|
||||||
|
from concurrent.futures import ThreadPoolExecutor, as_completed
|
||||||
|
import requests, openpyxl
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
S = requests.Session(); S.headers.update({'User-Agent': 'Mozilla/5.0'})
|
||||||
|
XLSX = r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\7.부여군\충청남도_부여군.xlsx'
|
||||||
|
DOMAIN = 'buyeo.go.kr'
|
||||||
|
REF_MAX = 37 # 3~37 = 사용자 검수 정답
|
||||||
|
|
||||||
|
KOGL_IMG = re.compile(r'(?:new_)?img_open(?:type|code)(\d{1,2})\.(?:png|jpe?g|gif)', re.I)
|
||||||
|
KOGL_LINK = re.compile(r'licenseType(\d)', re.I)
|
||||||
|
DECO = re.compile(
|
||||||
|
r'/site/common/img/|move\.png|no[-_]?img|blank\.|spacer\.|'
|
||||||
|
r'/img/(?:icon|ico|bul|bullet|arrow|btn|bg|tit|h\d)|'
|
||||||
|
r'ico_|btn_|bul_|bg_|_bg\.|icon_|'
|
||||||
|
r'file_icon|_icon\.|icon\.gif|_btn\.|see_btn|flag\.|webaccess|'
|
||||||
|
r'/skin/|/images/(?:kr/)?common(?:_new)?/|/common/img/|/_module/gnuri/|'
|
||||||
|
r'btn_page|page_(?:next|prev|first|last)', re.I)
|
||||||
|
MAPPDF = re.compile(r'pdf|viewer\.html|/map|kakao|daum|=map|naver.*map|google.*map', re.I)
|
||||||
|
VIDEO = re.compile(r'youtube\.com|youtu\.be|/embed/|\.mp4|\.webm|vimeo', re.I)
|
||||||
|
DETAIL = re.compile(r'mode=V&no=', re.I)
|
||||||
|
|
||||||
|
|
||||||
|
def get(url):
|
||||||
|
last = None
|
||||||
|
for _ in range(3):
|
||||||
|
try:
|
||||||
|
r = S.get(url, verify=False, timeout=20)
|
||||||
|
return BeautifulSoup(r.content, 'html.parser')
|
||||||
|
except Exception as e:
|
||||||
|
last = e
|
||||||
|
time.sleep(1.2)
|
||||||
|
raise last
|
||||||
|
|
||||||
|
|
||||||
|
def body(soup):
|
||||||
|
return soup.select_one('#txt') or soup.select_one('#contents') or soup.select_one('main') or soup
|
||||||
|
|
||||||
|
|
||||||
|
def real_img(node):
|
||||||
|
for im in node.find_all('img'):
|
||||||
|
src = im.get('src') or ''
|
||||||
|
if not src or KOGL_IMG.search(src) or DECO.search(src):
|
||||||
|
continue
|
||||||
|
return True
|
||||||
|
return False
|
||||||
|
|
||||||
|
|
||||||
|
def media_flags(node):
|
||||||
|
img = real_img(node)
|
||||||
|
vid = bool(node.find('video'))
|
||||||
|
for ifr in node.find_all('iframe'):
|
||||||
|
s = ifr.get('src') or ''
|
||||||
|
if VIDEO.search(s):
|
||||||
|
vid = True
|
||||||
|
elif MAPPDF.search(s):
|
||||||
|
img = True
|
||||||
|
for a in node.find_all('a'):
|
||||||
|
if VIDEO.search(a.get('href') or ''):
|
||||||
|
vid = True
|
||||||
|
return img, vid
|
||||||
|
|
||||||
|
|
||||||
|
def kogl_from(html):
|
||||||
|
"""전체 HTML에서 KOGL 이미지유형·링크유형 추출."""
|
||||||
|
imgs = [int(x) for x in KOGL_IMG.findall(html) if 1 <= int(x) <= 4]
|
||||||
|
links = [int(x) for x in KOGL_LINK.findall(html) if 1 <= int(x) <= 4]
|
||||||
|
return imgs, links
|
||||||
|
|
||||||
|
|
||||||
|
def _menu_cd(u):
|
||||||
|
m = re.search(r'menu_dvs_cd=([^&]+)', u or '')
|
||||||
|
return m.group(1) if m else None
|
||||||
|
|
||||||
|
|
||||||
|
def detail_links(soup, base):
|
||||||
|
"""같은 게시판(_prog/_board + 동일 menu_dvs_cd) 상세글만. 타 게시판 stray 링크 배제."""
|
||||||
|
if '/_prog/_board/' not in base:
|
||||||
|
return [] # 게시판 템플릿이 아니면 상세 없음(scate 등)
|
||||||
|
base_menu = _menu_cd(base)
|
||||||
|
out = []
|
||||||
|
for a in body(soup).find_all('a', href=True):
|
||||||
|
h = a['href']
|
||||||
|
if not DETAIL.search(h):
|
||||||
|
continue
|
||||||
|
full = urljoin(base, h)
|
||||||
|
if base_menu and _menu_cd(full) and _menu_cd(full) != base_menu:
|
||||||
|
continue # 다른 게시판 글 → 제외
|
||||||
|
out.append(full)
|
||||||
|
seen = set(); res = []
|
||||||
|
for u in out:
|
||||||
|
if u not in seen:
|
||||||
|
seen.add(u); res.append(u)
|
||||||
|
return res[:5]
|
||||||
|
|
||||||
|
|
||||||
|
def decide_O(img_types, link_types):
|
||||||
|
"""KOGL 규칙: 이미지유형 우선. 링크만이면 licenseType, 1·2·3·4 전부=범례→미부착."""
|
||||||
|
if img_types:
|
||||||
|
t = img_types[0]
|
||||||
|
mism = bool(link_types) and (link_types[0] != t)
|
||||||
|
return f'{t}유형', mism
|
||||||
|
if link_types:
|
||||||
|
if {1, 2, 3, 4}.issubset(set(link_types)):
|
||||||
|
return '미부착', False
|
||||||
|
return f'{link_types[0]}유형', False
|
||||||
|
return '미부착', False
|
||||||
|
|
||||||
|
|
||||||
|
def judge(k, L, M):
|
||||||
|
"""반환 dict: N, O, P, Q, S(mismatch)."""
|
||||||
|
try:
|
||||||
|
soup = get(k)
|
||||||
|
except Exception as e:
|
||||||
|
return {'err': f'fetch:{str(e)[:25]}'}
|
||||||
|
bd = body(soup); full = str(soup)
|
||||||
|
txt = len(bd.get_text(strip=True)) >= 1
|
||||||
|
is_board_url = '/_prog/_board/' in k # 실제 게시판 URL만(.html·서브사이트는 게시판 아님)
|
||||||
|
|
||||||
|
has_img, has_vid = media_flags(bd)
|
||||||
|
# O: 페이지/리스트 자체 마크 우선
|
||||||
|
img_t, link_t = kogl_from(full)
|
||||||
|
o_loc = '게시판' if L == '게시판' else '페이지'
|
||||||
|
o_imgs, o_links = list(img_t), list(link_t)
|
||||||
|
|
||||||
|
# '빈 게시판→없음'은 진짜 게시판(_prog/_board)에 글이 0개일 때만.
|
||||||
|
# L=게시판 M=0 이라도 .html 페이지·서브사이트는 본문내용으로 판정(원본 L/M 오기 보정).
|
||||||
|
dls = detail_links(soup, k) if (L == '게시판') else []
|
||||||
|
board0 = (L == '게시판' and is_board_url and (M in (0, '0', None)) and not dls)
|
||||||
|
|
||||||
|
if L == '게시판' and not board0:
|
||||||
|
# N: 상위 5글 본문에서 실제 이미지/영상 누적(이미지 콘텐츠 회수율↑)
|
||||||
|
for idx, du in enumerate(dls):
|
||||||
|
try:
|
||||||
|
ds = get(du); db = body(ds); df = str(ds)
|
||||||
|
except Exception:
|
||||||
|
continue
|
||||||
|
di, dv = media_flags(db)
|
||||||
|
if di: has_img = True
|
||||||
|
if dv: has_vid = True
|
||||||
|
# O: 리스트 마크 없으면 '첫(대표) 글'만으로 판정
|
||||||
|
# (사용자 컨벤션: 01030502 글1=1유형→1유형 / 업무추진비 글1무·글5만3유형→미부착)
|
||||||
|
if idx == 0 and not o_imgs and not o_links:
|
||||||
|
dimg, dlink = kogl_from(df)
|
||||||
|
if dimg or dlink:
|
||||||
|
o_imgs, o_links, o_loc = dimg, dlink, '게시물'
|
||||||
|
if has_img and has_vid:
|
||||||
|
break
|
||||||
|
|
||||||
|
# N
|
||||||
|
if board0:
|
||||||
|
N = '없음'
|
||||||
|
else:
|
||||||
|
parts = []
|
||||||
|
if txt: parts.append('어문')
|
||||||
|
if has_img: parts.append('이미지')
|
||||||
|
if has_vid: parts.append('영상')
|
||||||
|
N = ','.join(parts) if parts else '없음'
|
||||||
|
|
||||||
|
O, mism = decide_O(o_imgs, o_links)
|
||||||
|
P = (o_loc if O != '미부착' else None)
|
||||||
|
Q = ('Y' if (O != '미부착' and o_links) else None)
|
||||||
|
return {'N': N, 'O': O, 'P': P, 'Q': Q, 'mismatch': mism}
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
mode = '--validate' if '--validate' in sys.argv else ('--write' if '--write' in sys.argv else '--dry')
|
||||||
|
wb = openpyxl.load_workbook(XLSX); ws = wb.active
|
||||||
|
|
||||||
|
rows = []
|
||||||
|
for r in range(3, ws.max_row + 1):
|
||||||
|
k = ws.cell(r, 11).value; L = ws.cell(r, 12).value; M = ws.cell(r, 13).value
|
||||||
|
if not (isinstance(k, str) and DOMAIN in k):
|
||||||
|
continue
|
||||||
|
if L == '사이트':
|
||||||
|
continue
|
||||||
|
if mode == '--validate' and r > REF_MAX:
|
||||||
|
continue
|
||||||
|
if mode == '--write' and r <= REF_MAX:
|
||||||
|
continue
|
||||||
|
rows.append((r, k, L, M))
|
||||||
|
print(f'mode={mode} | 대상 {len(rows)}행')
|
||||||
|
|
||||||
|
res = {}
|
||||||
|
t0 = time.time()
|
||||||
|
with ThreadPoolExecutor(max_workers=3) as ex:
|
||||||
|
futs = {ex.submit(judge, k, L, M): (r, k, L, M) for r, k, L, M in rows}
|
||||||
|
done = 0
|
||||||
|
for fut in as_completed(futs):
|
||||||
|
r, k, L, M = futs[fut]
|
||||||
|
res[r] = fut.result()
|
||||||
|
done += 1
|
||||||
|
if done % 40 == 0:
|
||||||
|
print(f' {done}/{len(rows)} ({time.time()-t0:.0f}s)')
|
||||||
|
print(f'크롤링 완료 ({time.time()-t0:.0f}s)')
|
||||||
|
|
||||||
|
if mode == '--validate':
|
||||||
|
miss = 0
|
||||||
|
for r, k, L, M in rows:
|
||||||
|
d = res[r]
|
||||||
|
if d.get('err'):
|
||||||
|
print(f' 행{r} ERR {d["err"]}'); continue
|
||||||
|
curN = ws.cell(r, 14).value; curO = ws.cell(r, 15).value
|
||||||
|
curP = ws.cell(r, 16).value; curQ = ws.cell(r, 17).value
|
||||||
|
dn = (d['N'] or '') != (curN or '')
|
||||||
|
do = (d['O'] or '') != (curO or '')
|
||||||
|
dp = (d['P'] or '') != (curP or '')
|
||||||
|
dq = (d['Q'] or '') != (curQ or '')
|
||||||
|
if dn or do or dp or dq:
|
||||||
|
miss += 1
|
||||||
|
g = ws.cell(r, 7).value or ws.cell(r, 6).value
|
||||||
|
print(f' ✗행{r} {str(g)[:16]:16}| N:{curN}→{d["N"]} {"≠" if dn else "="} | '
|
||||||
|
f'O:{curO}→{d["O"]}{"≠" if do else "="} P:{curP}→{d["P"]}{"≠" if dp else ""} Q:{curQ}→{d["Q"]}{"≠" if dq else ""}')
|
||||||
|
print(f'\n불일치 {miss}/{len(rows)}행 (0이면 로직이 사용자 검수와 100% 일치)')
|
||||||
|
return
|
||||||
|
|
||||||
|
# write (38행~)
|
||||||
|
if mode == '--write':
|
||||||
|
import shutil
|
||||||
|
bk = XLSX.replace('.xlsx', '_backup_NO재판정전.xlsx'); shutil.copy(XLSX, bk)
|
||||||
|
print('백업:', bk)
|
||||||
|
chg = 0
|
||||||
|
for r, k, L, M in rows:
|
||||||
|
d = res[r]
|
||||||
|
if d.get('err'):
|
||||||
|
print(f' 행{r} ERR {d["err"]}'); continue
|
||||||
|
old = (ws.cell(r,14).value, ws.cell(r,15).value, ws.cell(r,16).value, ws.cell(r,17).value)
|
||||||
|
ws.cell(r, 14).value = d['N']
|
||||||
|
ws.cell(r, 15).value = d['O']
|
||||||
|
ws.cell(r, 16).value = d['P']
|
||||||
|
ws.cell(r, 17).value = d['Q']
|
||||||
|
if d['mismatch']:
|
||||||
|
cur = (ws.cell(r,19).value or '').strip()
|
||||||
|
ws.cell(r,19).value = '링크주소 오기' if not cur else cur+' / 링크주소 오기'
|
||||||
|
new = (d['N'], d['O'], d['P'], d['Q'])
|
||||||
|
if old != new:
|
||||||
|
chg += 1
|
||||||
|
g = ws.cell(r,7).value or ws.cell(r,6).value
|
||||||
|
print(f' 행{r} {str(g)[:16]:16}| N:{old[0]}→{new[0]} O:{old[1]}→{new[1]} P:{old[2]}→{new[2]} Q:{old[3]}→{new[3]}')
|
||||||
|
wb.save(XLSX)
|
||||||
|
print(f'\n저장: {XLSX} ({chg}행 변경)')
|
||||||
|
return
|
||||||
|
|
||||||
|
print('(--validate 또는 --write 지정)')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
166
_스크립트/_buyeo_scate.py
Normal file
166
_스크립트/_buyeo_scate.py
Normal file
@ -0,0 +1,166 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""부여군 사전정보공개(정보목록공개) fiexd_tab1 scate=1~9 분야탭 전개.
|
||||||
|
- _tab_expand 의 일반 필터탭 제외 규칙(같은 경로·쿼리만 다름)에 걸려 누락된 케이스.
|
||||||
|
사용자 지시로 9개 분야탭을 각각 별도 게시판 행으로 전개(보령시 사전정보공표 컨벤션).
|
||||||
|
- 각 분야 = 게시판 / M=목록 항목수(테이블 데이터행) / N=어문 / O=미부착.
|
||||||
|
- 평탄화→치환→재구성(D/E/F 재병합·K하이퍼링크·B순번)은 _tab_expand 재사용.
|
||||||
|
"""
|
||||||
|
import sys, shutil, warnings, re
|
||||||
|
from copy import copy
|
||||||
|
import requests, openpyxl
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
from openpyxl.styles import Alignment, Font
|
||||||
|
from openpyxl.utils import get_column_letter
|
||||||
|
|
||||||
|
sys.path.insert(0, r'D:\01.프로젝트\DB수집\_스크립트')
|
||||||
|
from _tab_expand import load_flat, MAXCOL, CAT_COLS, HEADER_MERGES
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
H = {'User-Agent': 'Mozilla/5.0'}
|
||||||
|
XLSX = r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\7.부여군\충청남도_부여군.xlsx'
|
||||||
|
BASE = 'https://www.buyeo.go.kr/_prog/service/index.php?scate='
|
||||||
|
|
||||||
|
|
||||||
|
def collect():
|
||||||
|
"""9개 scate 라벨+항목수 수집."""
|
||||||
|
r = requests.get(BASE + '1', headers=H, timeout=15, verify=False)
|
||||||
|
soup = BeautifulSoup(r.content, 'html.parser')
|
||||||
|
labels = {}
|
||||||
|
for li in soup.select('.fiexd_tab1 li'):
|
||||||
|
a = li.find('a'); m = re.search(r'scate=(\d+)', a.get('href') or '')
|
||||||
|
if m:
|
||||||
|
labels[int(m.group(1))] = a.get_text(strip=True)
|
||||||
|
out = []
|
||||||
|
for sc in sorted(labels):
|
||||||
|
rr = requests.get(BASE + str(sc), headers=H, timeout=15, verify=False)
|
||||||
|
s = BeautifulSoup(rr.content, 'html.parser')
|
||||||
|
t = s.find('table')
|
||||||
|
cnt = sum(1 for tr in t.find_all('tr') if tr.find_all('td'))
|
||||||
|
out.append((sc, labels[sc], cnt))
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
write = '--write' in sys.argv
|
||||||
|
scates = collect()
|
||||||
|
print('=== scate 분야탭 ===')
|
||||||
|
for sc, lab, cnt in scates:
|
||||||
|
print(f' scate={sc} | 항목={cnt} | {lab}')
|
||||||
|
print(f'합계 항목 {sum(c for _,_,c in scates)} / 9행 전개\n')
|
||||||
|
|
||||||
|
wb = openpyxl.load_workbook(XLSX)
|
||||||
|
ws = wb.active
|
||||||
|
rows = load_flat(ws)
|
||||||
|
|
||||||
|
# 대상 행 식별: K가 scate=1
|
||||||
|
tgt = None
|
||||||
|
for row in rows:
|
||||||
|
k = row['vals'].get(11)
|
||||||
|
if isinstance(k, str) and 'service/index.php?scate=1' in k:
|
||||||
|
tgt = row; break
|
||||||
|
if tgt is None:
|
||||||
|
print('!! 대상(scate=1) 행 없음'); return
|
||||||
|
lc = 6 # leaf=F (사전정보공개)
|
||||||
|
childc = lc + 1 # G
|
||||||
|
print(f'대상 sheet행 {tgt["src"]} (F={tgt["vals"].get(6)}) → 자식 {get_column_letter(childc)} 에 9분야 전개')
|
||||||
|
|
||||||
|
if not write:
|
||||||
|
print('\n(계획만. --write 로 기입)')
|
||||||
|
return
|
||||||
|
|
||||||
|
backup = XLSX.replace('.xlsx', '_backup_scate전.xlsx')
|
||||||
|
shutil.copy(XLSX, backup)
|
||||||
|
print('백업:', backup)
|
||||||
|
|
||||||
|
out_rows = []
|
||||||
|
for row in rows:
|
||||||
|
if row is tgt:
|
||||||
|
for sc, lab, cnt in scates:
|
||||||
|
nv = dict(tgt['vals'])
|
||||||
|
for c in CAT_COLS:
|
||||||
|
if c > lc:
|
||||||
|
nv[c] = None
|
||||||
|
nv[childc] = lab
|
||||||
|
nv[11] = BASE + str(sc)
|
||||||
|
nv[12] = '게시판'
|
||||||
|
nv[13] = cnt
|
||||||
|
nv[14] = '어문'
|
||||||
|
nv[15] = '미부착'
|
||||||
|
for c in (16, 17, 18): # P,Q,R 비움
|
||||||
|
nv[c] = None
|
||||||
|
# 신규 행(scate>1)은 비고~ 비움, scate=1은 기존 보존
|
||||||
|
if sc != 1:
|
||||||
|
for c in range(19, MAXCOL + 1):
|
||||||
|
nv[c] = None
|
||||||
|
out_rows.append({'vals': nv, 'style': tgt, 'url': nv[11]})
|
||||||
|
else:
|
||||||
|
out_rows.append({'vals': dict(row['vals']), 'style': row, 'url': row['vals'].get(11)})
|
||||||
|
|
||||||
|
# 데이터 클리어
|
||||||
|
for r in range(3, ws.max_row + 1):
|
||||||
|
for c in range(1, MAXCOL + 1):
|
||||||
|
ws.cell(r, c).value = None
|
||||||
|
ws.cell(r, c).hyperlink = None
|
||||||
|
|
||||||
|
START = 3
|
||||||
|
for i, orow in enumerate(out_rows):
|
||||||
|
r = START + i
|
||||||
|
sty = orow['style']['styles']
|
||||||
|
for c in range(1, MAXCOL + 1):
|
||||||
|
cell = ws.cell(r, c)
|
||||||
|
f, fl, bd, al, nf, pr = sty[c]
|
||||||
|
cell.font = copy(f); cell.fill = copy(fl); cell.border = copy(bd)
|
||||||
|
cell.alignment = copy(al); cell.number_format = nf; cell.protection = copy(pr)
|
||||||
|
v = orow['vals']
|
||||||
|
for c in range(1, MAXCOL + 1):
|
||||||
|
ws.cell(r, c).value = v.get(c)
|
||||||
|
ws.cell(r, 2).value = i + 1
|
||||||
|
END = START + len(out_rows) - 1
|
||||||
|
|
||||||
|
center = Alignment(horizontal='center', vertical='center', wrap_text=True)
|
||||||
|
left = Alignment(horizontal='left', vertical='center', wrap_text=False)
|
||||||
|
|
||||||
|
def merge_runs(col_letter, col_idx, group_cols=()):
|
||||||
|
cur_val = ws.cell(START, col_idx).value
|
||||||
|
cur_grp = tuple(ws.cell(START, g).value for g in group_cols)
|
||||||
|
run_start = START; runs = []
|
||||||
|
for r in range(START + 1, END + 1):
|
||||||
|
val = ws.cell(r, col_idx).value
|
||||||
|
grp = tuple(ws.cell(r, gg).value for gg in group_cols)
|
||||||
|
if val == cur_val and grp == cur_grp:
|
||||||
|
continue
|
||||||
|
if cur_val not in (None, '') and r - 1 > run_start:
|
||||||
|
runs.append((run_start, r - 1))
|
||||||
|
cur_val, cur_grp, run_start = val, grp, r
|
||||||
|
if cur_val not in (None, '') and END > run_start:
|
||||||
|
runs.append((run_start, END))
|
||||||
|
for s, e in runs:
|
||||||
|
ws.merge_cells(f'{col_letter}{s}:{col_letter}{e}')
|
||||||
|
ws.cell(s, col_idx).alignment = center
|
||||||
|
|
||||||
|
merge_runs('F', 6, group_cols=(4, 5))
|
||||||
|
merge_runs('E', 5, group_cols=(4,))
|
||||||
|
merge_runs('D', 4)
|
||||||
|
for r in range(START, END + 1):
|
||||||
|
for c in (4, 5, 6, 7):
|
||||||
|
if ws.cell(r, c).value is not None:
|
||||||
|
ws.cell(r, c).alignment = center
|
||||||
|
for r in range(1, END + 1):
|
||||||
|
ws.row_dimensions[r].height = 15
|
||||||
|
|
||||||
|
for r in range(START, END + 1):
|
||||||
|
cell = ws.cell(r, 11)
|
||||||
|
u = cell.value
|
||||||
|
if isinstance(u, str) and u.startswith(('http://', 'https://')):
|
||||||
|
cell.hyperlink = u
|
||||||
|
old = cell.font
|
||||||
|
cell.font = Font(name=old.name or '맑은 고딕', size=old.size or 11,
|
||||||
|
bold=old.bold, italic=old.italic, color='0000FF', underline='single')
|
||||||
|
cell.alignment = left
|
||||||
|
|
||||||
|
wb.save(XLSX)
|
||||||
|
print(f'저장 완료: {XLSX} (총 {len(out_rows)}행)')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
660
_스크립트/_chungbuk_phase1_all.py
Normal file
660
_스크립트/_chungbuk_phase1_all.py
Normal file
@ -0,0 +1,660 @@
|
|||||||
|
"""충청북도 11개 시·군 Phase 1 일괄 처리.
|
||||||
|
|
||||||
|
매뉴얼: D:\\01.프로젝트\\DB수집\\사이트맵_수집_매뉴얼.md
|
||||||
|
출력: 각 폴더의 {기관명}.xlsx
|
||||||
|
"""
|
||||||
|
import os
|
||||||
|
import re
|
||||||
|
import shutil
|
||||||
|
import ssl
|
||||||
|
import sys
|
||||||
|
import warnings
|
||||||
|
from copy import copy
|
||||||
|
from urllib.parse import urljoin
|
||||||
|
|
||||||
|
import openpyxl
|
||||||
|
import requests
|
||||||
|
from requests.adapters import HTTPAdapter
|
||||||
|
from urllib3.util.ssl_ import create_urllib3_context
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
from openpyxl.styles import Alignment, Font
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
TEMPLATE = r'D:\01.프로젝트\DB수집\자료_취합_예시.xlsx'
|
||||||
|
|
||||||
|
|
||||||
|
class WeakSSLAdapter(HTTPAdapter):
|
||||||
|
"""레거시 SSL 핸드셰이크 허용 (영동군 등)."""
|
||||||
|
def init_poolmanager(self, *args, **kwargs):
|
||||||
|
ctx = create_urllib3_context()
|
||||||
|
ctx.set_ciphers('DEFAULT@SECLEVEL=0')
|
||||||
|
ctx.options |= 0x4 # OP_LEGACY_SERVER_CONNECT
|
||||||
|
ctx.check_hostname = False
|
||||||
|
ctx.verify_mode = ssl.CERT_NONE
|
||||||
|
kwargs['ssl_context'] = ctx
|
||||||
|
return super().init_poolmanager(*args, **kwargs)
|
||||||
|
|
||||||
|
|
||||||
|
def make_session(weak_ssl=False):
|
||||||
|
s = requests.Session()
|
||||||
|
s.headers.update(H)
|
||||||
|
if weak_ssl:
|
||||||
|
s.mount('https://', WeakSSLAdapter())
|
||||||
|
return s
|
||||||
|
|
||||||
|
|
||||||
|
def fetch_html(url, session=None, timeout=20, force_encoding=None):
|
||||||
|
s = session or requests.Session()
|
||||||
|
if not session:
|
||||||
|
s.headers.update(H)
|
||||||
|
r = s.get(url, timeout=timeout, verify=False, allow_redirects=True)
|
||||||
|
if force_encoding:
|
||||||
|
r.encoding = force_encoding
|
||||||
|
else:
|
||||||
|
# 메타 태그에서 charset 시도 → 없으면 apparent_encoding
|
||||||
|
meta_charset = re.search(rb'<meta[^>]*charset=["\']?\s*([\w-]+)', r.content[:4096], re.I)
|
||||||
|
if meta_charset:
|
||||||
|
r.encoding = meta_charset.group(1).decode('ascii', errors='ignore')
|
||||||
|
else:
|
||||||
|
r.encoding = r.apparent_encoding
|
||||||
|
return r.text
|
||||||
|
|
||||||
|
|
||||||
|
def clean_text(s):
|
||||||
|
return re.sub(r'\s+', ' ', s or '').strip().replace('\xa0', '')
|
||||||
|
|
||||||
|
|
||||||
|
def extract_href(a):
|
||||||
|
if a is None:
|
||||||
|
return ''
|
||||||
|
href = (a.get('href') or '').strip()
|
||||||
|
if not href or href.startswith('#') or href.lower().startswith('javascript:'):
|
||||||
|
return ''
|
||||||
|
return href
|
||||||
|
|
||||||
|
|
||||||
|
ENCODE_URI_PAT = re.compile(r"""(?:location\.href|window\.location(?:\.href)?)\s*=\s*encodeURI\(\s*['"]([^'"]+)['"]\s*\)""")
|
||||||
|
ONCLICK_HREF_PAT = re.compile(r"""(?:location\.href|window\.open)\s*\(?\s*['"]([^'"]+)['"]""")
|
||||||
|
|
||||||
|
|
||||||
|
def extract_href_with_onclick(a):
|
||||||
|
if a is None:
|
||||||
|
return ''
|
||||||
|
href = (a.get('href') or '').strip()
|
||||||
|
if href and not href.startswith('#') and not href.lower().startswith('javascript:'):
|
||||||
|
return href
|
||||||
|
onclick = a.get('onclick', '')
|
||||||
|
if onclick:
|
||||||
|
m = ENCODE_URI_PAT.search(onclick)
|
||||||
|
if m:
|
||||||
|
return m.group(1)
|
||||||
|
m = ONCLICK_HREF_PAT.search(onclick)
|
||||||
|
if m:
|
||||||
|
return m.group(1)
|
||||||
|
return ''
|
||||||
|
|
||||||
|
|
||||||
|
# ================================================================
|
||||||
|
# 파서들
|
||||||
|
# ================================================================
|
||||||
|
|
||||||
|
|
||||||
|
def parse_depth1_chungbuk(soup, base):
|
||||||
|
"""충북 e-Gov depth1 형: div.depth1 > ul.depth1_list > li.depth1_item > a.depth1_text + div.depth2 > [div.depth2_content|depth2_wrap] > ul.depth2_list."""
|
||||||
|
rows = []
|
||||||
|
container = soup.select_one('div.depth.depth1, div.depth1')
|
||||||
|
if not container:
|
||||||
|
return rows
|
||||||
|
top_ul = container.find('ul', class_=re.compile(r'depth1?_list'), recursive=False)
|
||||||
|
if not top_ul:
|
||||||
|
return rows
|
||||||
|
|
||||||
|
def walk_depth(ul, level, base_path, out):
|
||||||
|
for li in ul.find_all('li', recursive=False):
|
||||||
|
a = li.find('a', class_=re.compile(rf'depth{level}_text'), recursive=False) or li.find('a', recursive=False)
|
||||||
|
if not a:
|
||||||
|
continue
|
||||||
|
text = clean_text(a.get_text())
|
||||||
|
href = extract_href(a)
|
||||||
|
path = base_path + [(text, href)]
|
||||||
|
# Find next depth container — 다양한 wrapper 클래스 지원: _content / _wrap / 없음
|
||||||
|
next_div = li.find('div', class_=re.compile(rf'depth\s+depth{level+1}\b'), recursive=False) or \
|
||||||
|
li.find('div', class_=re.compile(rf'\bdepth{level+1}\b'), recursive=False)
|
||||||
|
if next_div:
|
||||||
|
# wrapper 가 있을 수도, 없을 수도
|
||||||
|
next_ul = next_div.find('ul', class_=re.compile(rf'depth{level+1}_list'), recursive=True)
|
||||||
|
if next_ul:
|
||||||
|
# 같은 깊이의 첫번째 ul.depth_list (자손 검색이지만 보통 1번 wrapper 안)
|
||||||
|
out.append({'path': list(path), 'href': href})
|
||||||
|
walk_depth(next_ul, level + 1, path, out)
|
||||||
|
continue
|
||||||
|
out.append({'path': list(path), 'href': href})
|
||||||
|
|
||||||
|
tmp = []
|
||||||
|
walk_depth(top_ul, 1, [], tmp)
|
||||||
|
cols = 'DEFGHIJ'
|
||||||
|
for item in tmp:
|
||||||
|
p = item['path']
|
||||||
|
row = {c: '' for c in 'DEFGHIJ'}
|
||||||
|
row['href'] = item['href']
|
||||||
|
for i, (t, _) in enumerate(p):
|
||||||
|
col = cols[i] if i < len(cols) else 'J'
|
||||||
|
row[col] = t
|
||||||
|
rows.append(row)
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def parse_danyang_ld(soup, base):
|
||||||
|
"""단양군: ul#menu_sitemap.ld1 > li.cd1 > a.l1 + div.lb1 > ul.ld2 > li.cd2 > a.l2 + div.lb2 > ul.ld3 > ...
|
||||||
|
|
||||||
|
'menutype_empty' (메인 등) 은 스킵.
|
||||||
|
"""
|
||||||
|
rows = []
|
||||||
|
container = soup.select_one('ul#menu_sitemap')
|
||||||
|
if not container:
|
||||||
|
return rows
|
||||||
|
|
||||||
|
def walk(ul, level, base_path, out):
|
||||||
|
for li in ul.find_all('li', recursive=False):
|
||||||
|
a = li.find('a', class_=re.compile(rf'\bl{level}\b'), recursive=False) or li.find('a', recursive=False)
|
||||||
|
if not a:
|
||||||
|
continue
|
||||||
|
a_cls = ' '.join(a.get('class', []))
|
||||||
|
if 'menutype_empty' in a_cls:
|
||||||
|
continue # 메인 같은 빈 항목
|
||||||
|
text = clean_text(a.get_text())
|
||||||
|
href = extract_href(a)
|
||||||
|
path = base_path + [(text, href)]
|
||||||
|
next_div = li.find('div', class_=re.compile(rf'\blb{level}\b'), recursive=False)
|
||||||
|
if next_div:
|
||||||
|
next_ul = next_div.find('ul', class_=re.compile(rf'\bld{level+1}\b'), recursive=False)
|
||||||
|
if next_ul:
|
||||||
|
out.append({'path': list(path), 'href': href})
|
||||||
|
walk(next_ul, level + 1, path, out)
|
||||||
|
continue
|
||||||
|
out.append({'path': list(path), 'href': href})
|
||||||
|
|
||||||
|
tmp = []
|
||||||
|
walk(container, 1, [], tmp)
|
||||||
|
cols = 'DEFGHIJ'
|
||||||
|
for item in tmp:
|
||||||
|
p = item['path']
|
||||||
|
row = {c: '' for c in cols}
|
||||||
|
row['href'] = item['href']
|
||||||
|
for i, (t, _) in enumerate(p):
|
||||||
|
col = cols[i] if i < len(cols) else 'J'
|
||||||
|
row[col] = t
|
||||||
|
rows.append(row)
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def parse_dl_dt_dd(soup, base, use_onclick=False):
|
||||||
|
"""계룡시·홍성군·영동군형: div.sitemap > dl > dt + dd > b > a + ul > li > a (+ ul > li > a)."""
|
||||||
|
rows = []
|
||||||
|
sm = soup.select_one('div.sitemap[class*=type2]') or soup.select_one('div.sitemap[class*=type1]') or soup.select_one('div.sitemap')
|
||||||
|
if not sm:
|
||||||
|
return rows
|
||||||
|
href_fn = extract_href_with_onclick if use_onclick else extract_href
|
||||||
|
|
||||||
|
def walk(ul, depth, base_path, out):
|
||||||
|
for li in ul.find_all('li', recursive=False):
|
||||||
|
a = li.find('a', recursive=False)
|
||||||
|
if not a:
|
||||||
|
continue
|
||||||
|
text = clean_text(a.get_text())
|
||||||
|
href = href_fn(a)
|
||||||
|
path = base_path[:depth] + [(text, href)]
|
||||||
|
nested = li.find('ul', recursive=False)
|
||||||
|
if nested:
|
||||||
|
out.append({'path': list(path), 'href': href})
|
||||||
|
walk(nested, depth + 1, path, out)
|
||||||
|
else:
|
||||||
|
out.append({'path': list(path), 'href': href})
|
||||||
|
|
||||||
|
for dl in sm.find_all('dl', recursive=False):
|
||||||
|
dt = dl.find('dt')
|
||||||
|
dt_a = dt.find('a') if dt else None
|
||||||
|
D = clean_text(dt_a.get_text() if dt_a else (dt.get_text() if dt else ''))
|
||||||
|
for dd in dl.find_all('dd', recursive=False):
|
||||||
|
b = dd.find('b')
|
||||||
|
b_a = b.find('a') if b else None
|
||||||
|
E = clean_text(b_a.get_text()) if b_a else ''
|
||||||
|
E_href = href_fn(b_a) if b_a else ''
|
||||||
|
nested = dd.find('ul', recursive=False)
|
||||||
|
if nested:
|
||||||
|
tmp = []
|
||||||
|
walk(nested, 0, [], tmp)
|
||||||
|
for item in tmp:
|
||||||
|
p = item['path']
|
||||||
|
row = {'D': D, 'E': E, 'href': item['href'],
|
||||||
|
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''}
|
||||||
|
for di, (t, _) in enumerate(p):
|
||||||
|
col = 'FGHIJ'[di] if di < 5 else 'J'
|
||||||
|
row[col] = t
|
||||||
|
rows.append(row)
|
||||||
|
else:
|
||||||
|
rows.append({'D': D, 'E': E, 'href': E_href,
|
||||||
|
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def parse_yesan_depth(soup, base):
|
||||||
|
"""예산·청양·증평형: ul.depth1_ul > li > a (D, .th_1st/.th1_lnk) + [div.item >]? ul.depth2_ul > li > a (E) + ul.depth3_ul > li > a (F).
|
||||||
|
|
||||||
|
div.item wrapper 있을수도 없을수도 — 둘 다 지원.
|
||||||
|
a 클래스: .th_1st (예산), .th1_lnk (증평) — 첫 번째 a 가져옴.
|
||||||
|
"""
|
||||||
|
rows = []
|
||||||
|
container = soup.select_one('ul.depth1_ul')
|
||||||
|
if not container:
|
||||||
|
return rows
|
||||||
|
for top_li in container.find_all('li', recursive=False):
|
||||||
|
a1 = top_li.find('a', class_=re.compile(r'th[_]?1?(?:_1st|_lnk)?'), recursive=False) or \
|
||||||
|
top_li.find('a', recursive=False)
|
||||||
|
if not a1:
|
||||||
|
continue
|
||||||
|
D = clean_text(a1.get_text())
|
||||||
|
# div.item wrapper 있을 수도 없을 수도
|
||||||
|
item = top_li.find('div', class_='item', recursive=False)
|
||||||
|
d2_ul = (item.find('ul', class_='depth2_ul') if item else None) or \
|
||||||
|
top_li.find('ul', class_='depth2_ul', recursive=False)
|
||||||
|
if not d2_ul:
|
||||||
|
continue
|
||||||
|
for d2_li in d2_ul.find_all('li', recursive=False):
|
||||||
|
d2_a = d2_li.find('a', recursive=False)
|
||||||
|
if not d2_a:
|
||||||
|
continue
|
||||||
|
E = clean_text(d2_a.get_text())
|
||||||
|
E_href = extract_href(d2_a)
|
||||||
|
d3_ul = d2_li.find('ul', class_='depth3_ul', recursive=False)
|
||||||
|
if not d3_ul:
|
||||||
|
rows.append({'D': D, 'E': E, 'href': E_href,
|
||||||
|
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
continue
|
||||||
|
rows.append({'D': D, 'E': E, 'href': E_href,
|
||||||
|
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
for d3_li in d3_ul.find_all('li', recursive=False):
|
||||||
|
d3_a = d3_li.find('a', recursive=False)
|
||||||
|
if not d3_a:
|
||||||
|
continue
|
||||||
|
F = clean_text(d3_a.get_text())
|
||||||
|
rows.append({'D': D, 'E': E, 'F': F, 'href': extract_href(d3_a),
|
||||||
|
'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def parse_jincheon_recursive(soup, base):
|
||||||
|
"""진천군: div.sitemap > ul > li > a + div > ul > li > a + div > ul > ... 재귀."""
|
||||||
|
rows = []
|
||||||
|
container = soup.select_one('div.sitemap')
|
||||||
|
if not container:
|
||||||
|
return rows
|
||||||
|
top_ul = container.find('ul', recursive=False)
|
||||||
|
if not top_ul:
|
||||||
|
return rows
|
||||||
|
|
||||||
|
def walk(ul, base_path, out):
|
||||||
|
for li in ul.find_all('li', recursive=False):
|
||||||
|
a = li.find('a', recursive=False)
|
||||||
|
if not a:
|
||||||
|
continue
|
||||||
|
text = clean_text(a.get_text())
|
||||||
|
if not text or text == '메뉴명이 없습니다.':
|
||||||
|
continue
|
||||||
|
href = extract_href(a)
|
||||||
|
path = base_path + [(text, href)]
|
||||||
|
next_div = li.find('div', recursive=False)
|
||||||
|
if next_div:
|
||||||
|
next_ul = next_div.find('ul', recursive=False)
|
||||||
|
if next_ul:
|
||||||
|
out.append({'path': list(path), 'href': href})
|
||||||
|
walk(next_ul, path, out)
|
||||||
|
continue
|
||||||
|
out.append({'path': list(path), 'href': href})
|
||||||
|
|
||||||
|
tmp = []
|
||||||
|
walk(top_ul, [], tmp)
|
||||||
|
cols = 'DEFGHIJ'
|
||||||
|
for item in tmp:
|
||||||
|
p = item['path']
|
||||||
|
row = {c: '' for c in cols}
|
||||||
|
row['href'] = item['href']
|
||||||
|
for i, (t, _) in enumerate(p):
|
||||||
|
col = cols[i] if i < len(cols) else 'J'
|
||||||
|
row[col] = t
|
||||||
|
rows.append(row)
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def parse_cheongju_sitemap(soup, base):
|
||||||
|
"""청주시: div#sitemap > div.site_map_col > div.sitemap_box > h3 > a (D) + ul.sm2depth > li > a (E) + ul.sm3depth > li > a (F).
|
||||||
|
|
||||||
|
별도의 '인트로' sitemap_box 는 D만 있고 ul.sm2depth 단순한 외부 링크 묶음 — 그대로 둠.
|
||||||
|
"""
|
||||||
|
rows = []
|
||||||
|
container = soup.select_one('div#sitemap')
|
||||||
|
if not container:
|
||||||
|
return rows
|
||||||
|
for box in container.select('div.sitemap_box'):
|
||||||
|
h3 = box.find('h3')
|
||||||
|
h3_a = h3.find('a') if h3 else None
|
||||||
|
D = clean_text(h3_a.get_text() if h3_a else (h3.get_text() if h3 else ''))
|
||||||
|
D_href = extract_href(h3_a) if h3_a else ''
|
||||||
|
sm2 = box.find('ul', class_='sm2depth')
|
||||||
|
if not sm2:
|
||||||
|
rows.append({'D': D, 'href': D_href, 'E': '', 'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
continue
|
||||||
|
for li2 in sm2.find_all('li', recursive=False):
|
||||||
|
a2 = li2.find('a', recursive=False)
|
||||||
|
if not a2:
|
||||||
|
continue
|
||||||
|
E = clean_text(a2.get_text())
|
||||||
|
E_href = extract_href(a2)
|
||||||
|
sm3 = li2.find('ul', class_='sm3depth', recursive=False)
|
||||||
|
if not sm3:
|
||||||
|
rows.append({'D': D, 'E': E, 'href': E_href,
|
||||||
|
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
continue
|
||||||
|
rows.append({'D': D, 'E': E, 'href': E_href,
|
||||||
|
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
for li3 in sm3.find_all('li', recursive=False):
|
||||||
|
a3 = li3.find('a', recursive=False)
|
||||||
|
if not a3:
|
||||||
|
continue
|
||||||
|
F = clean_text(a3.get_text())
|
||||||
|
rows.append({'D': D, 'E': E, 'F': F, 'href': extract_href(a3),
|
||||||
|
'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def parse_chungju_sitemap(soup, base):
|
||||||
|
"""충주시: div#sitemap > div.site_map_col > div.sitemap_box > h3.h0 > a (D) + ul > li > a.h4 (E) + ul.bu > li > a (F)."""
|
||||||
|
rows = []
|
||||||
|
container = soup.select_one('div#sitemap')
|
||||||
|
if not container:
|
||||||
|
return rows
|
||||||
|
for box in container.select('div.sitemap_box'):
|
||||||
|
h3 = box.find('h3')
|
||||||
|
h3_a = h3.find('a') if h3 else None
|
||||||
|
D = clean_text(h3_a.get_text() if h3_a else (h3.get_text() if h3 else ''))
|
||||||
|
D_href = extract_href(h3_a) if h3_a else ''
|
||||||
|
e_ul = box.find('ul', recursive=False)
|
||||||
|
if not e_ul:
|
||||||
|
rows.append({'D': D, 'href': D_href, 'E': '', 'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
continue
|
||||||
|
for e_li in e_ul.find_all('li', recursive=False):
|
||||||
|
a_E = e_li.find('a', class_='h4', recursive=False) or e_li.find('a', recursive=False)
|
||||||
|
if not a_E:
|
||||||
|
continue
|
||||||
|
E = clean_text(a_E.get_text())
|
||||||
|
E_href = extract_href(a_E)
|
||||||
|
f_ul = e_li.find('ul', class_='bu', recursive=False)
|
||||||
|
if not f_ul:
|
||||||
|
rows.append({'D': D, 'E': E, 'href': E_href,
|
||||||
|
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
continue
|
||||||
|
rows.append({'D': D, 'E': E, 'href': E_href,
|
||||||
|
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
for f_li in f_ul.find_all('li', recursive=False):
|
||||||
|
f_a = f_li.find('a', recursive=False)
|
||||||
|
if not f_a:
|
||||||
|
continue
|
||||||
|
F = clean_text(f_a.get_text())
|
||||||
|
rows.append({'D': D, 'E': E, 'F': F, 'href': extract_href(f_a),
|
||||||
|
'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
# ================================================================
|
||||||
|
# 사이트 설정
|
||||||
|
# ================================================================
|
||||||
|
|
||||||
|
SITES = [
|
||||||
|
{
|
||||||
|
'idx': 1, 'name': '괴산군', 'base': 'https://www.goesan.go.kr',
|
||||||
|
'sitemap': 'https://www.goesan.go.kr/www/sitemap.do?key=28',
|
||||||
|
'sheet': '01_괴산군', 'parser': parse_depth1_chungbuk,
|
||||||
|
'domain': 'goesan.go.kr', 'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\1.괴산군',
|
||||||
|
'weak_ssl': False,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
'idx': 2, 'name': '단양군', 'base': 'https://www.danyang.go.kr',
|
||||||
|
'sitemap': 'https://www.danyang.go.kr/dy21/98',
|
||||||
|
'sheet': '02_단양군', 'parser': parse_danyang_ld,
|
||||||
|
'domain': 'danyang.go.kr', 'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\2.단양군',
|
||||||
|
'weak_ssl': False,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
# 보은군 — www.boeun.go.kr DNS 차단, apex boeun.go.kr 는 정상(2026-05-30 재확인)
|
||||||
|
'idx': 3, 'name': '보은군', 'base': 'https://boeun.go.kr',
|
||||||
|
'sitemap': 'https://boeun.go.kr/www/sitemap.do?key=1323',
|
||||||
|
'sheet': '03_보은군', 'parser': parse_depth1_chungbuk,
|
||||||
|
'domain': 'boeun.go.kr', 'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\3.보은군',
|
||||||
|
'weak_ssl': False,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
'idx': 4, 'name': '영동군', 'base': 'https://www.yd21.go.kr',
|
||||||
|
'sitemap': 'https://www.yd21.go.kr/kr/html/guide/0701.html',
|
||||||
|
'sheet': '04_영동군', 'parser': parse_dl_dt_dd,
|
||||||
|
'domain': 'yd21.go.kr', 'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\4.영동군',
|
||||||
|
'weak_ssl': True,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
'idx': 5, 'name': '옥천군', 'base': 'https://www.oc.go.kr',
|
||||||
|
'sitemap': 'https://www.oc.go.kr/www/sub.do?key=121',
|
||||||
|
'sheet': '05_옥천군', 'parser': parse_depth1_chungbuk,
|
||||||
|
'domain': 'oc.go.kr', 'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\5.옥천군',
|
||||||
|
'weak_ssl': False,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
'idx': 6, 'name': '음성군', 'base': 'https://www.eumseong.go.kr',
|
||||||
|
'sitemap': 'https://www.eumseong.go.kr/www/sub.do?key=722',
|
||||||
|
'sheet': '06_음성군', 'parser': parse_depth1_chungbuk,
|
||||||
|
'domain': 'eumseong.go.kr', 'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\6.음성군',
|
||||||
|
'weak_ssl': False,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
'idx': 7, 'name': '제천시', 'base': 'https://www.jecheon.go.kr',
|
||||||
|
'sitemap': 'https://www.jecheon.go.kr/www/sitemap.do?key=553',
|
||||||
|
'sheet': '07_제천시', 'parser': parse_depth1_chungbuk,
|
||||||
|
'domain': 'jecheon.go.kr', 'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\7.제천시',
|
||||||
|
'weak_ssl': False,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
'idx': 8, 'name': '증평군', 'base': 'https://www.jp.go.kr',
|
||||||
|
'sitemap': 'https://www.jp.go.kr/kor/sitemap_11.do',
|
||||||
|
'sheet': '08_증평군', 'parser': parse_yesan_depth,
|
||||||
|
'domain': 'jp.go.kr', 'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\8.증평군',
|
||||||
|
'weak_ssl': False,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
'idx': 9, 'name': '진천군', 'base': 'https://www.jincheon.go.kr',
|
||||||
|
'sitemap': 'https://www.jincheon.go.kr/home/sub.do?menukey=445',
|
||||||
|
'sheet': '09_진천군', 'parser': parse_jincheon_recursive,
|
||||||
|
'domain': 'jincheon.go.kr', 'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\9.진천군',
|
||||||
|
'weak_ssl': False,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
'idx': 10, 'name': '청주시', 'base': 'https://www.cheongju.go.kr',
|
||||||
|
'sitemap': 'https://www.cheongju.go.kr/www/sitemap.do?key=589',
|
||||||
|
'sheet': '10_청주시', 'parser': parse_cheongju_sitemap,
|
||||||
|
'domain': 'cheongju.go.kr', 'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\10.청주시',
|
||||||
|
'weak_ssl': False,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
'idx': 11, 'name': '충주시', 'base': 'https://www.chungju.go.kr',
|
||||||
|
'sitemap': 'https://www.chungju.go.kr/www/sub.do?key=692',
|
||||||
|
'sheet': '11_충주시', 'parser': parse_chungju_sitemap,
|
||||||
|
'domain': 'chungju.go.kr', 'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\11.충주시',
|
||||||
|
'weak_ssl': False,
|
||||||
|
},
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
# ================================================================
|
||||||
|
# 엑셀 생성 (공통)
|
||||||
|
# ================================================================
|
||||||
|
|
||||||
|
|
||||||
|
def write_excel(site, raw_rows):
|
||||||
|
name = site['name']
|
||||||
|
base = site['base']
|
||||||
|
domain = site['domain']
|
||||||
|
output = f"{site['folder']}\\{os.path.basename(os.path.dirname(site['folder']))}_{name}.xlsx"
|
||||||
|
|
||||||
|
def abs_url(href):
|
||||||
|
if not href:
|
||||||
|
return ''
|
||||||
|
href = href.strip()
|
||||||
|
if href.startswith(('javascript:', '#')):
|
||||||
|
return ''
|
||||||
|
if href.startswith(('http://', 'https://')):
|
||||||
|
return href
|
||||||
|
return urljoin(base + '/', href)
|
||||||
|
|
||||||
|
def is_external(url):
|
||||||
|
return url.startswith(('http://', 'https://')) and domain not in url
|
||||||
|
|
||||||
|
# 부모-자식 URL 중복 제거
|
||||||
|
final_rows = []
|
||||||
|
i = 0
|
||||||
|
removed = 0
|
||||||
|
while i < len(raw_rows):
|
||||||
|
row = raw_rows[i]
|
||||||
|
if (i + 1 < len(raw_rows)
|
||||||
|
and row.get('G', '') == ''
|
||||||
|
and raw_rows[i + 1].get('D') == row.get('D')
|
||||||
|
and raw_rows[i + 1].get('E') == row.get('E')
|
||||||
|
and raw_rows[i + 1].get('F') == row.get('F')
|
||||||
|
and raw_rows[i + 1].get('G', '') != ''
|
||||||
|
and raw_rows[i + 1].get('href') == row.get('href')):
|
||||||
|
removed += 1
|
||||||
|
i += 1
|
||||||
|
continue
|
||||||
|
final_rows.append(row)
|
||||||
|
i += 1
|
||||||
|
|
||||||
|
print(f' [{name}] 원본 {len(raw_rows)} → 중복 {removed} → 최종 {len(final_rows)}')
|
||||||
|
if not final_rows:
|
||||||
|
print(f' [{name}] !! 행 0개 — 파서 점검 필요. 엑셀 생성 스킵.')
|
||||||
|
return False
|
||||||
|
|
||||||
|
shutil.copy(TEMPLATE, output)
|
||||||
|
wb = openpyxl.load_workbook(output)
|
||||||
|
ws = wb.active
|
||||||
|
ws.title = site['sheet']
|
||||||
|
|
||||||
|
HEADER = {'B1:R1', 'S1:W1', 'Y1:AA1'}
|
||||||
|
for rng in [str(m) for m in ws.merged_cells.ranges if str(m) not in HEADER]:
|
||||||
|
ws.unmerge_cells(rng)
|
||||||
|
for row in ws.iter_rows(min_row=3, max_row=ws.max_row, min_col=1, max_col=ws.max_column):
|
||||||
|
for cell in row:
|
||||||
|
cell.value = None
|
||||||
|
|
||||||
|
START = 3
|
||||||
|
template_r = 3
|
||||||
|
cur_max = ws.max_row
|
||||||
|
for idx, item in enumerate(final_rows, start=START):
|
||||||
|
if idx > cur_max:
|
||||||
|
for c in range(1, ws.max_column + 1):
|
||||||
|
srcc = ws.cell(template_r, c)
|
||||||
|
tgt = ws.cell(idx, c)
|
||||||
|
if srcc.has_style:
|
||||||
|
tgt.font = copy(srcc.font)
|
||||||
|
tgt.fill = copy(srcc.fill)
|
||||||
|
tgt.border = copy(srcc.border)
|
||||||
|
tgt.alignment = copy(srcc.alignment)
|
||||||
|
tgt.number_format = srcc.number_format
|
||||||
|
tgt.protection = copy(srcc.protection)
|
||||||
|
url = abs_url(item.get('href', ''))
|
||||||
|
ws.cell(idx, 2).value = idx - 2
|
||||||
|
ws.cell(idx, 3).value = name
|
||||||
|
ws.cell(idx, 4).value = item.get('D', '')
|
||||||
|
ws.cell(idx, 5).value = item.get('E', '')
|
||||||
|
ws.cell(idx, 6).value = item.get('F', '')
|
||||||
|
ws.cell(idx, 7).value = item.get('G', '')
|
||||||
|
ws.cell(idx, 8).value = item.get('H', '')
|
||||||
|
ws.cell(idx, 9).value = item.get('I', '')
|
||||||
|
ws.cell(idx, 10).value = item.get('J', '')
|
||||||
|
ws.cell(idx, 11).value = url
|
||||||
|
if is_external(url):
|
||||||
|
ws.cell(idx, 19).value = '외부링크'
|
||||||
|
|
||||||
|
END = START + len(final_rows) - 1
|
||||||
|
center = Alignment(horizontal='center', vertical='center', wrap_text=True)
|
||||||
|
left = Alignment(horizontal='left', vertical='center', wrap_text=False)
|
||||||
|
|
||||||
|
def merge_runs(col_letter, col_idx, group_cols=()):
|
||||||
|
runs = []
|
||||||
|
cur_val = ws.cell(START, col_idx).value
|
||||||
|
cur_grp = tuple(ws.cell(START, g).value for g in group_cols)
|
||||||
|
run_start = START
|
||||||
|
for r in range(START + 1, END + 1):
|
||||||
|
v = ws.cell(r, col_idx).value
|
||||||
|
g = tuple(ws.cell(r, gg).value for gg in group_cols)
|
||||||
|
if v == cur_val and g == cur_grp:
|
||||||
|
continue
|
||||||
|
if cur_val not in (None, '') and r - 1 > run_start:
|
||||||
|
runs.append((run_start, r - 1))
|
||||||
|
cur_val, cur_grp, run_start = v, g, r
|
||||||
|
if cur_val not in (None, '') and END > run_start:
|
||||||
|
runs.append((run_start, END))
|
||||||
|
for s, e in runs:
|
||||||
|
ws.merge_cells(f'{col_letter}{s}:{col_letter}{e}')
|
||||||
|
ws.cell(s, col_idx).alignment = center
|
||||||
|
return len(runs)
|
||||||
|
|
||||||
|
# 병합 순서 F→E→D (D를 먼저 병합하면 2행부터 D=None이 되어 E 그룹키가 깨짐)
|
||||||
|
n_f = merge_runs('F', 6, group_cols=(4, 5))
|
||||||
|
n_e = merge_runs('E', 5, group_cols=(4,))
|
||||||
|
n_d = merge_runs('D', 4)
|
||||||
|
for r in range(START, END + 1):
|
||||||
|
for c in (4, 5, 6, 7):
|
||||||
|
if ws.cell(r, c).value is not None:
|
||||||
|
ws.cell(r, c).alignment = center
|
||||||
|
|
||||||
|
for r in range(1, END + 1):
|
||||||
|
ws.row_dimensions[r].height = 15
|
||||||
|
|
||||||
|
link_n = 0
|
||||||
|
for r in range(START, END + 1):
|
||||||
|
cell = ws.cell(r, 11)
|
||||||
|
u = cell.value
|
||||||
|
if u and isinstance(u, str) and u.startswith(('http://', 'https://')):
|
||||||
|
cell.hyperlink = u
|
||||||
|
old = cell.font
|
||||||
|
cell.font = Font(name=old.name or '맑은 고딕', size=old.size or 11,
|
||||||
|
bold=old.bold, italic=old.italic,
|
||||||
|
color='0000FF', underline='single')
|
||||||
|
cell.alignment = left
|
||||||
|
link_n += 1
|
||||||
|
|
||||||
|
wb.save(output)
|
||||||
|
ext_n = sum(1 for r in range(START, END + 1) if ws.cell(r, 19).value == '외부링크')
|
||||||
|
print(f' [{name}] 병합 D:{n_d} E:{n_e} F:{n_f} | 외부링크 {ext_n} | K링크 {link_n} → {output}')
|
||||||
|
return True
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
targets = sys.argv[1:] if len(sys.argv) > 1 else [s['name'] for s in SITES]
|
||||||
|
for site in SITES:
|
||||||
|
if site['name'] not in targets:
|
||||||
|
continue
|
||||||
|
print(f"\n{'='*60}\n[{site['idx']}.{site['name']}] {site['sitemap']}\n{'='*60}")
|
||||||
|
try:
|
||||||
|
sess = make_session(weak_ssl=site.get('weak_ssl', False))
|
||||||
|
html = fetch_html(site['sitemap'], session=sess)
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
raw_rows = site['parser'](soup, site['base'])
|
||||||
|
write_excel(site, raw_rows)
|
||||||
|
except Exception as e:
|
||||||
|
print(f' [{site["name"]}] !! 실패: {e}')
|
||||||
|
import traceback
|
||||||
|
traceback.print_exc()
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
357
_스크립트/_chungbuk_phase234_all.py
Normal file
357
_스크립트/_chungbuk_phase234_all.py
Normal file
@ -0,0 +1,357 @@
|
|||||||
|
"""충청북도 10개 시·군 Phase 2~4 일괄 처리 (보은군 제외 — DNS 차단).
|
||||||
|
|
||||||
|
매뉴얼: D:\\01.프로젝트\\DB수집\\사이트맵_수집_매뉴얼.md
|
||||||
|
"""
|
||||||
|
import re
|
||||||
|
import ssl
|
||||||
|
import sys
|
||||||
|
import time
|
||||||
|
import warnings
|
||||||
|
from urllib.parse import urljoin
|
||||||
|
from concurrent.futures import ThreadPoolExecutor, as_completed
|
||||||
|
|
||||||
|
import openpyxl
|
||||||
|
import requests
|
||||||
|
from requests.adapters import HTTPAdapter
|
||||||
|
from urllib3.util.ssl_ import create_urllib3_context
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
|
||||||
|
|
||||||
|
class WeakSSLAdapter(HTTPAdapter):
|
||||||
|
def init_poolmanager(self, *args, **kwargs):
|
||||||
|
ctx = create_urllib3_context()
|
||||||
|
ctx.set_ciphers('DEFAULT@SECLEVEL=0')
|
||||||
|
ctx.options |= 0x4
|
||||||
|
ctx.check_hostname = False
|
||||||
|
ctx.verify_mode = ssl.CERT_NONE
|
||||||
|
kwargs['ssl_context'] = ctx
|
||||||
|
return super().init_poolmanager(*args, **kwargs)
|
||||||
|
|
||||||
|
|
||||||
|
# ================================================================
|
||||||
|
# 정규식
|
||||||
|
# ================================================================
|
||||||
|
TOTAL_PAT = re.compile(r'총\s*(?:게시물\s*)?(\d[\d,]*)\s*(?:건|개|page|페이지)', re.I)
|
||||||
|
TOTAL_PAT_LOOSE = re.compile(r'(?:전체|총)\s*[:\-]?\s*(\d[\d,]*)\s*건', re.I)
|
||||||
|
KOGL_IMG_PAT = re.compile(r'(?:new_)?img_opentype(\d{2})\.png', re.I)
|
||||||
|
KOGL_LINK_PAT = re.compile(r'kogl\.or\.kr/info/licenseType(\d)', re.I)
|
||||||
|
YOUTUBE_PAT = re.compile(r'(?:youtube\.com|youtu\.be)', re.I)
|
||||||
|
VIDEO_EXT = re.compile(r'\.(mp4|webm|mov|avi)(?:\?|$)', re.I)
|
||||||
|
DETAIL_PAT = re.compile(r'(mode=V|view\.do|bbtSn=|seqRepeat=|nttId=|articleNo=|boardSeq=|menukey=)', re.I)
|
||||||
|
|
||||||
|
|
||||||
|
# 사이트별: xlsx, body selectors, weak_ssl 여부
|
||||||
|
SITES = {
|
||||||
|
'괴산군': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\1.괴산군\충청북도_괴산군.xlsx',
|
||||||
|
'body_sel': ['#contents', '#txt', 'main'], 'weak_ssl': False},
|
||||||
|
'단양군': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\2.단양군\충청북도_단양군.xlsx',
|
||||||
|
'body_sel': ['#contents', '#txt', 'main'], 'weak_ssl': False},
|
||||||
|
'보은군': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\3.보은군\충청북도_보은군.xlsx',
|
||||||
|
'body_sel': ['#contents', '#txt', 'main'], 'weak_ssl': False},
|
||||||
|
'영동군': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\4.영동군\충청북도_영동군.xlsx',
|
||||||
|
'body_sel': ['#txt', '#contents', 'main'], 'weak_ssl': True},
|
||||||
|
'옥천군': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\5.옥천군\충청북도_옥천군.xlsx',
|
||||||
|
'body_sel': ['#contents', '#txt', 'main'], 'weak_ssl': False},
|
||||||
|
'음성군': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\6.음성군\충청북도_음성군.xlsx',
|
||||||
|
'body_sel': ['#contents', '#txt', 'main'], 'weak_ssl': False},
|
||||||
|
'제천시': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\7.제천시\충청북도_제천시.xlsx',
|
||||||
|
'body_sel': ['#contents', '#txt', 'main'], 'weak_ssl': False},
|
||||||
|
'증평군': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\8.증평군\충청북도_증평군.xlsx',
|
||||||
|
'body_sel': ['#txt', '#contents', 'main'], 'weak_ssl': False},
|
||||||
|
'진천군': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\9.진천군\충청북도_진천군.xlsx',
|
||||||
|
'body_sel': ['#contents', '#txt', 'main'], 'weak_ssl': False},
|
||||||
|
'청주시': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\10.청주시\충청북도_청주시.xlsx',
|
||||||
|
'body_sel': ['#contents', '#txt', 'main'], 'weak_ssl': False},
|
||||||
|
'충주시': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\11.충주시\충청북도_충주시.xlsx',
|
||||||
|
'body_sel': ['#contents', '#txt', 'main'], 'weak_ssl': False},
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def make_session(weak_ssl=False):
|
||||||
|
s = requests.Session()
|
||||||
|
s.headers.update(H)
|
||||||
|
if weak_ssl:
|
||||||
|
s.mount('https://', WeakSSLAdapter())
|
||||||
|
return s
|
||||||
|
|
||||||
|
|
||||||
|
def fetch(session, url, timeout=12):
|
||||||
|
try:
|
||||||
|
r = session.get(url, timeout=timeout, verify=False, allow_redirects=True)
|
||||||
|
# 메타 charset 우선
|
||||||
|
meta_charset = re.search(rb'<meta[^>]*charset=["\']?\s*([\w-]+)', r.content[:4096], re.I)
|
||||||
|
if meta_charset:
|
||||||
|
r.encoding = meta_charset.group(1).decode('ascii', errors='ignore')
|
||||||
|
else:
|
||||||
|
r.encoding = r.apparent_encoding
|
||||||
|
if r.status_code == 200:
|
||||||
|
return BeautifulSoup(r.text, 'html.parser')
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def get_body(soup, selectors):
|
||||||
|
for sel in selectors:
|
||||||
|
el = soup.select_one(sel)
|
||||||
|
if el:
|
||||||
|
return el
|
||||||
|
return soup
|
||||||
|
|
||||||
|
|
||||||
|
def detect_form(body):
|
||||||
|
has_paging = bool(body.select('.pagination, .paging, nav.paging, .page_nav'))
|
||||||
|
text_inputs = [i for i in body.find_all('input')
|
||||||
|
if (i.get('type') or 'text').lower() in ('text', 'search')]
|
||||||
|
has_search = len(text_inputs) >= 1
|
||||||
|
txt = body.get_text(' ', strip=True)
|
||||||
|
m = TOTAL_PAT.search(txt) or TOTAL_PAT_LOOSE.search(txt)
|
||||||
|
total = None
|
||||||
|
if m:
|
||||||
|
digits = m.group(1).replace(',', '')
|
||||||
|
if digits.isdigit():
|
||||||
|
total = int(digits)
|
||||||
|
is_board = has_paging or has_search or (total is not None)
|
||||||
|
if is_board:
|
||||||
|
return '게시판', total if total is not None else 0
|
||||||
|
return '페이지', 1
|
||||||
|
|
||||||
|
|
||||||
|
def extract_detail_urls(body, base_url, limit=5):
|
||||||
|
urls = []
|
||||||
|
seen = set()
|
||||||
|
for a in body.find_all('a', href=True):
|
||||||
|
h = a['href']
|
||||||
|
if not h or h.startswith('#'):
|
||||||
|
continue
|
||||||
|
if DETAIL_PAT.search(h):
|
||||||
|
full = urljoin(base_url, h)
|
||||||
|
if full not in seen:
|
||||||
|
seen.add(full)
|
||||||
|
urls.append(full)
|
||||||
|
if len(urls) >= limit:
|
||||||
|
break
|
||||||
|
return urls
|
||||||
|
|
||||||
|
|
||||||
|
def detect_media(body):
|
||||||
|
has_text = len(body.get_text(strip=True)) > 30
|
||||||
|
has_image = False
|
||||||
|
for img in body.find_all('img'):
|
||||||
|
src = img.get('src', '')
|
||||||
|
if KOGL_IMG_PAT.search(src):
|
||||||
|
continue
|
||||||
|
if not src:
|
||||||
|
continue
|
||||||
|
has_image = True
|
||||||
|
break
|
||||||
|
has_video = False
|
||||||
|
for iframe in body.find_all('iframe'):
|
||||||
|
if YOUTUBE_PAT.search(iframe.get('src', '')):
|
||||||
|
has_video = True
|
||||||
|
break
|
||||||
|
if not has_video:
|
||||||
|
for a in body.find_all('a', href=True):
|
||||||
|
if YOUTUBE_PAT.search(a['href']):
|
||||||
|
has_video = True
|
||||||
|
break
|
||||||
|
if not has_video and body.find_all('video'):
|
||||||
|
has_video = True
|
||||||
|
if not has_video and VIDEO_EXT.search(str(body)):
|
||||||
|
has_video = True
|
||||||
|
return has_image, has_video, has_text
|
||||||
|
|
||||||
|
|
||||||
|
def n_string(has_text, has_image, has_video):
|
||||||
|
parts = []
|
||||||
|
if has_text: parts.append('어문')
|
||||||
|
if has_image: parts.append('이미지')
|
||||||
|
if has_video: parts.append('영상')
|
||||||
|
return ','.join(parts) if parts else '없음'
|
||||||
|
|
||||||
|
|
||||||
|
def img_has_valid_anchor(img):
|
||||||
|
p = img.parent
|
||||||
|
while p is not None:
|
||||||
|
if p.name == 'a':
|
||||||
|
href = p.get('href', '')
|
||||||
|
if href and not href.startswith('#') and not href.lower().startswith('javascript:'):
|
||||||
|
return True
|
||||||
|
return False
|
||||||
|
p = p.parent
|
||||||
|
return False
|
||||||
|
|
||||||
|
|
||||||
|
def detect_kogl(body):
|
||||||
|
types = set()
|
||||||
|
q_any_y = False
|
||||||
|
q_any_n = False
|
||||||
|
for a in body.find_all('a', href=True):
|
||||||
|
m = KOGL_LINK_PAT.search(a['href'])
|
||||||
|
if m:
|
||||||
|
types.add(int(m.group(1)))
|
||||||
|
q_any_y = True
|
||||||
|
for img in body.find_all('img'):
|
||||||
|
src = img.get('src', '')
|
||||||
|
m = KOGL_IMG_PAT.search(src)
|
||||||
|
if m:
|
||||||
|
types.add(int(m.group(1)))
|
||||||
|
if img_has_valid_anchor(img):
|
||||||
|
q_any_y = True
|
||||||
|
else:
|
||||||
|
q_any_n = True
|
||||||
|
for el in body.find_all(style=True):
|
||||||
|
m = KOGL_IMG_PAT.search(el.get('style', ''))
|
||||||
|
if m:
|
||||||
|
types.add(int(m.group(1)))
|
||||||
|
q_any_n = True
|
||||||
|
if not types:
|
||||||
|
return set(), None
|
||||||
|
return types, ('Y' if q_any_y else 'N')
|
||||||
|
|
||||||
|
|
||||||
|
def process_row(session, url, body_selectors):
|
||||||
|
out = {'L': '', 'M': '', 'N': '', 'O': '', 'P': '', 'Q': '', 'note': ''}
|
||||||
|
soup = fetch(session, url)
|
||||||
|
if soup is None:
|
||||||
|
out['note'] = '접근 실패'
|
||||||
|
return out
|
||||||
|
body = get_body(soup, body_selectors)
|
||||||
|
|
||||||
|
form, count = detect_form(body)
|
||||||
|
out['L'] = form
|
||||||
|
out['M'] = count if form == '게시판' else 1
|
||||||
|
|
||||||
|
has_img, has_vid, has_txt = detect_media(body)
|
||||||
|
types_main, q_main = detect_kogl(body)
|
||||||
|
P = '게시판' if types_main else ''
|
||||||
|
types_all = set(types_main)
|
||||||
|
q_flags = []
|
||||||
|
if q_main:
|
||||||
|
q_flags.append(q_main)
|
||||||
|
|
||||||
|
if form == '게시판':
|
||||||
|
detail_urls = extract_detail_urls(body, url, limit=5)
|
||||||
|
for du in detail_urls:
|
||||||
|
d_soup = fetch(session, du, timeout=10)
|
||||||
|
if not d_soup:
|
||||||
|
continue
|
||||||
|
d_body = get_body(d_soup, body_selectors)
|
||||||
|
di, dv, dt = detect_media(d_body)
|
||||||
|
has_img = has_img or di
|
||||||
|
has_vid = has_vid or dv
|
||||||
|
has_txt = has_txt or dt
|
||||||
|
dt_types, dt_q = detect_kogl(d_body)
|
||||||
|
if dt_types and not types_main and not P:
|
||||||
|
P = '게시물'
|
||||||
|
types_all |= dt_types
|
||||||
|
if dt_q:
|
||||||
|
q_flags.append(dt_q)
|
||||||
|
|
||||||
|
out['N'] = n_string(has_txt, has_img, has_vid)
|
||||||
|
|
||||||
|
if not types_all:
|
||||||
|
out['O'] = '미부착'
|
||||||
|
else:
|
||||||
|
sorted_types = sorted(types_all)
|
||||||
|
out['O'] = ','.join(f'{n}유형' for n in sorted_types)
|
||||||
|
out['P'] = P if P else '게시판'
|
||||||
|
out['Q'] = 'Y' if 'Y' in q_flags else 'N'
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def run_site(name, xlsx, body_selectors, weak_ssl=False, workers=10):
|
||||||
|
print(f'\n{"="*60}\n[{name}] {xlsx}\n{"="*60}')
|
||||||
|
wb = openpyxl.load_workbook(xlsx)
|
||||||
|
ws = wb.active
|
||||||
|
START = 3
|
||||||
|
END = START - 1
|
||||||
|
for r in range(START, ws.max_row + 1):
|
||||||
|
if ws.cell(r, 2).value is None:
|
||||||
|
break
|
||||||
|
END = r
|
||||||
|
|
||||||
|
tasks = []
|
||||||
|
for r in range(START, END + 1):
|
||||||
|
url = ws.cell(r, 11).value
|
||||||
|
is_ext = (ws.cell(r, 19).value == '외부링크')
|
||||||
|
tasks.append((r, url, is_ext))
|
||||||
|
n_ext = sum(1 for t in tasks if t[2])
|
||||||
|
print(f' 처리 대상: 총 {len(tasks)}행 (외부링크 {n_ext})')
|
||||||
|
|
||||||
|
t0 = time.time()
|
||||||
|
results = {}
|
||||||
|
session = make_session(weak_ssl=weak_ssl)
|
||||||
|
|
||||||
|
def worker(task):
|
||||||
|
row, url, is_ext = task
|
||||||
|
if is_ext:
|
||||||
|
return row, {'L': '사이트', 'M': '', 'N': '', 'O': '', 'P': '', 'Q': '', 'note': ''}
|
||||||
|
if not url or not isinstance(url, str):
|
||||||
|
return row, {'L': '', 'M': '', 'N': '', 'O': '', 'P': '', 'Q': '', 'note': 'URL 없음'}
|
||||||
|
return row, process_row(session, url, body_selectors)
|
||||||
|
|
||||||
|
done = 0
|
||||||
|
with ThreadPoolExecutor(max_workers=workers) as ex:
|
||||||
|
futs = [ex.submit(worker, t) for t in tasks]
|
||||||
|
for fut in as_completed(futs):
|
||||||
|
row, res = fut.result()
|
||||||
|
results[row] = res
|
||||||
|
done += 1
|
||||||
|
if done % 50 == 0 or done == len(tasks):
|
||||||
|
print(f' 진행 {done}/{len(tasks)} ({time.time()-t0:.0f}s)')
|
||||||
|
|
||||||
|
for r in range(START, END + 1):
|
||||||
|
res = results.get(r, {})
|
||||||
|
if not res:
|
||||||
|
continue
|
||||||
|
if res.get('L'): ws.cell(r, 12).value = res['L']
|
||||||
|
if res.get('M') != '': ws.cell(r, 13).value = res['M']
|
||||||
|
if res.get('N'): ws.cell(r, 14).value = res['N']
|
||||||
|
if res.get('O'): ws.cell(r, 15).value = res['O']
|
||||||
|
if res.get('P'): ws.cell(r, 16).value = res['P']
|
||||||
|
if res.get('Q'): ws.cell(r, 17).value = res['Q']
|
||||||
|
if res.get('note'):
|
||||||
|
existing = ws.cell(r, 19).value
|
||||||
|
if not existing:
|
||||||
|
ws.cell(r, 19).value = res['note']
|
||||||
|
|
||||||
|
wb.save(xlsx)
|
||||||
|
forms = {}
|
||||||
|
attach = {'미부착': 0, '부착': 0, '기타': 0}
|
||||||
|
q_dist = {'Y': 0, 'N': 0, '': 0}
|
||||||
|
for r, res in results.items():
|
||||||
|
forms[res.get('L', '')] = forms.get(res.get('L', ''), 0) + 1
|
||||||
|
o = res.get('O', '')
|
||||||
|
if o == '미부착': attach['미부착'] += 1
|
||||||
|
elif o and '유형' in o: attach['부착'] += 1
|
||||||
|
else: attach['기타'] += 1
|
||||||
|
q = res.get('Q', '')
|
||||||
|
q_dist[q] = q_dist.get(q, 0) + 1
|
||||||
|
print(f' [{name}] L분포 {forms} | 부착 {attach} | Q {q_dist} | 시간 {time.time()-t0:.0f}s')
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
targets = sys.argv[1:] if len(sys.argv) > 1 else list(SITES.keys())
|
||||||
|
total_t0 = time.time()
|
||||||
|
for name in targets:
|
||||||
|
if name not in SITES:
|
||||||
|
print(f' 알 수 없음: {name}')
|
||||||
|
continue
|
||||||
|
cfg = SITES[name]
|
||||||
|
try:
|
||||||
|
run_site(name, cfg['xlsx'], cfg['body_sel'], weak_ssl=cfg.get('weak_ssl', False))
|
||||||
|
except Exception as e:
|
||||||
|
print(f' [{name}] 실패: {e}')
|
||||||
|
import traceback
|
||||||
|
traceback.print_exc()
|
||||||
|
print(f'\n총 소요: {time.time()-total_t0:.0f}s')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
734
_스크립트/_chungnam_phase1_all.py
Normal file
734
_스크립트/_chungnam_phase1_all.py
Normal file
@ -0,0 +1,734 @@
|
|||||||
|
"""충청남도 11개 시·군 Phase 1 일괄 처리.
|
||||||
|
|
||||||
|
매뉴얼: D:\\01.프로젝트\\DB수집\\사이트맵_수집_매뉴얼.md
|
||||||
|
출력: 각 폴더의 {기관명}.xlsx
|
||||||
|
"""
|
||||||
|
import os
|
||||||
|
import re
|
||||||
|
import shutil
|
||||||
|
import sys
|
||||||
|
import warnings
|
||||||
|
from copy import copy
|
||||||
|
from urllib.parse import urljoin
|
||||||
|
|
||||||
|
import openpyxl
|
||||||
|
import requests
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
from openpyxl.styles import Alignment, Font
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
TEMPLATE = r'D:\01.프로젝트\DB수집\자료_취합_예시.xlsx'
|
||||||
|
|
||||||
|
# ================================================================
|
||||||
|
# 공통 유틸
|
||||||
|
# ================================================================
|
||||||
|
|
||||||
|
|
||||||
|
def fetch_html(url, timeout=20):
|
||||||
|
r = requests.get(url, headers=H, timeout=timeout, verify=False, allow_redirects=True)
|
||||||
|
r.encoding = r.apparent_encoding
|
||||||
|
return r.text
|
||||||
|
|
||||||
|
|
||||||
|
def clean_text(s):
|
||||||
|
return re.sub(r'\s+', ' ', s or '').strip().replace('\xa0', '')
|
||||||
|
|
||||||
|
|
||||||
|
def extract_href(a):
|
||||||
|
if a is None:
|
||||||
|
return ''
|
||||||
|
href = (a.get('href') or '').strip()
|
||||||
|
if not href or href.startswith('#') or href.lower().startswith('javascript:'):
|
||||||
|
# Try onclick encodeURI extraction
|
||||||
|
return ''
|
||||||
|
return href
|
||||||
|
|
||||||
|
|
||||||
|
ENCODE_URI_PAT = re.compile(r"""(?:location\.href|window\.location(?:\.href)?)\s*=\s*encodeURI\(\s*['"]([^'"]+)['"]\s*\)""")
|
||||||
|
ONCLICK_HREF_PAT = re.compile(r"""(?:location\.href|window\.open)\s*\(?\s*['"]([^'"]+)['"]""")
|
||||||
|
|
||||||
|
|
||||||
|
def extract_href_with_onclick(a):
|
||||||
|
"""Extract href; if href is dummy (#...), check onclick for encodeURI/location.href."""
|
||||||
|
if a is None:
|
||||||
|
return ''
|
||||||
|
href = (a.get('href') or '').strip()
|
||||||
|
if href and not href.startswith('#') and not href.lower().startswith('javascript:'):
|
||||||
|
return href
|
||||||
|
# Try onclick
|
||||||
|
onclick = a.get('onclick', '')
|
||||||
|
if onclick:
|
||||||
|
m = ENCODE_URI_PAT.search(onclick)
|
||||||
|
if m:
|
||||||
|
return m.group(1)
|
||||||
|
m = ONCLICK_HREF_PAT.search(onclick)
|
||||||
|
if m:
|
||||||
|
return m.group(1)
|
||||||
|
return ''
|
||||||
|
|
||||||
|
|
||||||
|
# ================================================================
|
||||||
|
# 사이트별 파서 (각각 raw_rows 리스트 반환)
|
||||||
|
# raw_rows: [{'D': str, 'E': str, 'F': str, 'G': str, 'H': str, 'I': str, 'J': str, 'href': str}, ...]
|
||||||
|
# ================================================================
|
||||||
|
|
||||||
|
|
||||||
|
def parse_eGov_type1(soup, base):
|
||||||
|
"""논산시·아산시 형: div.sitemap.type1 > [div.s_1th + div.inner > div.s_2th + ul...]."""
|
||||||
|
sitemap = soup.select_one('div.sitemap.type1') or soup.select_one('div.sitemap')
|
||||||
|
rows = []
|
||||||
|
if not sitemap:
|
||||||
|
return rows
|
||||||
|
current_D = ''
|
||||||
|
|
||||||
|
def walk(ul, depth, base_path, out):
|
||||||
|
for li in ul.find_all('li', recursive=False):
|
||||||
|
a = li.find('a', recursive=False)
|
||||||
|
if not a:
|
||||||
|
continue
|
||||||
|
text = clean_text(a.get_text())
|
||||||
|
href = extract_href(a)
|
||||||
|
path = base_path[:depth] + [(text, href)]
|
||||||
|
nested = li.find('ul', recursive=False)
|
||||||
|
if nested:
|
||||||
|
out.append({'path': list(path), 'href': href})
|
||||||
|
walk(nested, depth + 1, path, out)
|
||||||
|
else:
|
||||||
|
out.append({'path': list(path), 'href': href})
|
||||||
|
|
||||||
|
for child in sitemap.find_all('div', recursive=False):
|
||||||
|
cls = child.get('class', [])
|
||||||
|
if 's_1th' in cls:
|
||||||
|
a = child.find('a')
|
||||||
|
current_D = clean_text(a.get_text()) if a else ''
|
||||||
|
elif 'inner' in cls:
|
||||||
|
s2 = child.find('div', class_='s_2th')
|
||||||
|
mid_a = s2.find('a') if s2 else None
|
||||||
|
mid_name = clean_text(mid_a.get_text()) if mid_a else ''
|
||||||
|
mid_href = extract_href(mid_a) if mid_a else ''
|
||||||
|
uls = child.find_all('ul', recursive=False)
|
||||||
|
if not uls:
|
||||||
|
rows.append({'D': current_D, 'E': mid_name, 'href': mid_href,
|
||||||
|
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
continue
|
||||||
|
for ul in uls:
|
||||||
|
tmp = []
|
||||||
|
walk(ul, 0, [], tmp)
|
||||||
|
for item in tmp:
|
||||||
|
p = item['path']
|
||||||
|
row = {'D': current_D, 'E': mid_name, 'href': item['href'],
|
||||||
|
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''}
|
||||||
|
for di, (t, _) in enumerate(p):
|
||||||
|
col = 'FGHIJ'[di] if di < 5 else 'J'
|
||||||
|
row[col] = t
|
||||||
|
rows.append(row)
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def parse_amThum(soup, base):
|
||||||
|
"""당진시·태안군 형: div.amThum > h2.siteNN + div.sitemap_grep > ul.sitemap_list > li > a.first + ul > li > a (+ ul > li > a)."""
|
||||||
|
rows = []
|
||||||
|
for amthum in soup.select('div.amThum'):
|
||||||
|
h2 = amthum.find(['h2', 'h3'], class_=re.compile(r'site\d+'))
|
||||||
|
D = clean_text(h2.find('span').get_text()) if h2 and h2.find('span') else (clean_text(h2.get_text()) if h2 else '')
|
||||||
|
grep = amthum.find('div', class_='sitemap_grep') or amthum
|
||||||
|
for sl in grep.find_all('ul', class_='sitemap_list'):
|
||||||
|
# Each ul.sitemap_list contains li > a.first + ul > li > a
|
||||||
|
for top_li in sl.find_all('li', recursive=False):
|
||||||
|
a_first = top_li.find('a', class_='first', recursive=False)
|
||||||
|
if not a_first:
|
||||||
|
continue
|
||||||
|
E = clean_text(a_first.get_text())
|
||||||
|
E_href = extract_href(a_first)
|
||||||
|
inner_ul = top_li.find('ul', recursive=False)
|
||||||
|
if not inner_ul:
|
||||||
|
rows.append({'D': D, 'E': E, 'href': E_href,
|
||||||
|
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
continue
|
||||||
|
for mid_li in inner_ul.find_all('li', recursive=False):
|
||||||
|
mid_a = mid_li.find('a', recursive=False)
|
||||||
|
if not mid_a:
|
||||||
|
continue
|
||||||
|
F = clean_text(mid_a.get_text())
|
||||||
|
F_href = extract_href(mid_a)
|
||||||
|
deeper = mid_li.find('ul', recursive=False)
|
||||||
|
if not deeper:
|
||||||
|
rows.append({'D': D, 'E': E, 'F': F, 'href': F_href,
|
||||||
|
'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
continue
|
||||||
|
rows.append({'D': D, 'E': E, 'F': F, 'href': F_href,
|
||||||
|
'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
for deep_li in deeper.find_all('li', recursive=False):
|
||||||
|
deep_a = deep_li.find('a', recursive=False)
|
||||||
|
if not deep_a:
|
||||||
|
continue
|
||||||
|
G = clean_text(deep_a.get_text())
|
||||||
|
rows.append({'D': D, 'E': E, 'F': F, 'G': G,
|
||||||
|
'href': extract_href(deep_a),
|
||||||
|
'H': '', 'I': '', 'J': ''})
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def parse_ul_sitemap_h4(soup, base):
|
||||||
|
"""보령시·서천군 형: ul.sitemap > li (대분류) > h4.siteNN > span + ul > li > h5 > a + ul > li.list > a."""
|
||||||
|
rows = []
|
||||||
|
container = soup.select_one('ul.sitemap') or soup.select_one('div#contents ul.sitemap')
|
||||||
|
if not container:
|
||||||
|
return rows
|
||||||
|
for top_li in container.find_all('li', recursive=False):
|
||||||
|
h4 = top_li.find(['h4', 'h3'], recursive=False)
|
||||||
|
D = ''
|
||||||
|
if h4:
|
||||||
|
span = h4.find('span')
|
||||||
|
D = clean_text(span.get_text() if span else h4.get_text())
|
||||||
|
# Each direct ul under top_li is a sub-group
|
||||||
|
for sub_ul in top_li.find_all('ul', recursive=False):
|
||||||
|
for sub_li in sub_ul.find_all('li', recursive=False):
|
||||||
|
h5 = sub_li.find(['h5', 'h6'], recursive=False)
|
||||||
|
if h5:
|
||||||
|
h5_a = h5.find('a')
|
||||||
|
E = clean_text(h5_a.get_text()) if h5_a else clean_text(h5.get_text())
|
||||||
|
E_href = extract_href(h5_a) if h5_a else ''
|
||||||
|
else:
|
||||||
|
a = sub_li.find('a', recursive=False)
|
||||||
|
E = clean_text(a.get_text()) if a else ''
|
||||||
|
E_href = extract_href(a) if a else ''
|
||||||
|
deeper = sub_li.find('ul', recursive=False)
|
||||||
|
if not deeper:
|
||||||
|
rows.append({'D': D, 'E': E, 'href': E_href,
|
||||||
|
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
continue
|
||||||
|
rows.append({'D': D, 'E': E, 'href': E_href,
|
||||||
|
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
for leaf_li in deeper.find_all('li', recursive=False):
|
||||||
|
leaf_a = leaf_li.find('a', recursive=False)
|
||||||
|
if not leaf_a:
|
||||||
|
continue
|
||||||
|
F = clean_text(leaf_a.get_text())
|
||||||
|
rows.append({'D': D, 'E': E, 'F': F, 'href': extract_href(leaf_a),
|
||||||
|
'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
# 4-level (rare)
|
||||||
|
sub_ul2 = leaf_li.find('ul', recursive=False)
|
||||||
|
if sub_ul2:
|
||||||
|
for ll2 in sub_ul2.find_all('li', recursive=False):
|
||||||
|
la2 = ll2.find('a', recursive=False)
|
||||||
|
if la2:
|
||||||
|
G = clean_text(la2.get_text())
|
||||||
|
rows.append({'D': D, 'E': E, 'F': F, 'G': G,
|
||||||
|
'href': extract_href(la2),
|
||||||
|
'H': '', 'I': '', 'J': ''})
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def parse_dl_dt_dd(soup, base, use_onclick=False):
|
||||||
|
"""홍성군·예산군·천안시·금산군 dl 형: div.sitemap (옵션 .type2) > dl > dt + dd > b > a + ul > li > a (+ ul > li > a).
|
||||||
|
|
||||||
|
use_onclick=True 인 경우 onclick의 encodeURI/location.href에서 URL을 추출.
|
||||||
|
"""
|
||||||
|
rows = []
|
||||||
|
sm = soup.select_one('div.sitemap[class*=type2]') or soup.select_one('div.sitemap')
|
||||||
|
if not sm:
|
||||||
|
return rows
|
||||||
|
href_fn = extract_href_with_onclick if use_onclick else extract_href
|
||||||
|
|
||||||
|
def walk(ul, depth, base_path, out):
|
||||||
|
for li in ul.find_all('li', recursive=False):
|
||||||
|
a = li.find('a', recursive=False)
|
||||||
|
if not a:
|
||||||
|
continue
|
||||||
|
text = clean_text(a.get_text())
|
||||||
|
href = href_fn(a)
|
||||||
|
path = base_path[:depth] + [(text, href)]
|
||||||
|
nested = li.find('ul', recursive=False)
|
||||||
|
if nested:
|
||||||
|
out.append({'path': list(path), 'href': href})
|
||||||
|
walk(nested, depth + 1, path, out)
|
||||||
|
else:
|
||||||
|
out.append({'path': list(path), 'href': href})
|
||||||
|
|
||||||
|
for dl in sm.find_all('dl', recursive=False):
|
||||||
|
dt = dl.find('dt')
|
||||||
|
dt_a = dt.find('a') if dt else None
|
||||||
|
D = clean_text(dt_a.get_text() if dt_a else (dt.get_text() if dt else ''))
|
||||||
|
for dd in dl.find_all('dd', recursive=False):
|
||||||
|
b = dd.find('b')
|
||||||
|
b_a = b.find('a') if b else None
|
||||||
|
E = clean_text(b_a.get_text()) if b_a else ''
|
||||||
|
E_href = href_fn(b_a) if b_a else ''
|
||||||
|
nested = dd.find('ul', recursive=False)
|
||||||
|
if nested:
|
||||||
|
tmp = []
|
||||||
|
walk(nested, 0, [], tmp)
|
||||||
|
for item in tmp:
|
||||||
|
p = item['path']
|
||||||
|
row = {'D': D, 'E': E, 'href': item['href'],
|
||||||
|
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''}
|
||||||
|
for di, (t, _) in enumerate(p):
|
||||||
|
col = 'FGHIJ'[di] if di < 5 else 'J'
|
||||||
|
row[col] = t
|
||||||
|
rows.append(row)
|
||||||
|
else:
|
||||||
|
rows.append({'D': D, 'E': E, 'href': E_href,
|
||||||
|
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def parse_seosan_top_menu(soup, base):
|
||||||
|
"""서산시: ul.top_menu > li.depth1 > a.depth1_ti (D) + div...div.depth2_wrap > ul.depth2 > li > a (E) + ul.depth3 > li > a (F)."""
|
||||||
|
rows = []
|
||||||
|
tm = soup.select_one('ul.top_menu')
|
||||||
|
if not tm:
|
||||||
|
return rows
|
||||||
|
for top_li in tm.find_all('li', class_='depth1', recursive=False):
|
||||||
|
a1 = top_li.find('a', class_='depth1_ti')
|
||||||
|
D = clean_text(a1.get_text()) if a1 else ''
|
||||||
|
depth2 = top_li.select_one('ul.depth2')
|
||||||
|
if not depth2:
|
||||||
|
continue
|
||||||
|
for d2_li in depth2.find_all('li', recursive=False):
|
||||||
|
d2_a = d2_li.find('a', recursive=False)
|
||||||
|
if not d2_a:
|
||||||
|
continue
|
||||||
|
E = clean_text(d2_a.get_text())
|
||||||
|
E_href = extract_href(d2_a)
|
||||||
|
d3 = d2_li.find('ul', class_='depth3', recursive=False)
|
||||||
|
if not d3:
|
||||||
|
rows.append({'D': D, 'E': E, 'href': E_href,
|
||||||
|
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
continue
|
||||||
|
rows.append({'D': D, 'E': E, 'href': E_href,
|
||||||
|
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
for d3_li in d3.find_all('li', recursive=False):
|
||||||
|
d3_a = d3_li.find('a', recursive=False)
|
||||||
|
if not d3_a:
|
||||||
|
continue
|
||||||
|
F = clean_text(d3_a.get_text())
|
||||||
|
rows.append({'D': D, 'E': E, 'F': F, 'href': extract_href(d3_a),
|
||||||
|
'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def parse_yesan_depth(soup, base):
|
||||||
|
"""예산군: #gnb > ul.depth1_ul > li > a.th_1st + div.item > ul.depth2_ul > li > a + ul.depth3_ul > li > a."""
|
||||||
|
rows = []
|
||||||
|
container = soup.select_one('#gnb ul.depth1_ul') or soup.select_one('ul.depth1_ul')
|
||||||
|
if not container:
|
||||||
|
return rows
|
||||||
|
for top_li in container.find_all('li', recursive=False):
|
||||||
|
a1 = top_li.find('a', class_='th_1st', recursive=False)
|
||||||
|
D = clean_text(a1.get_text()) if a1 else ''
|
||||||
|
item = top_li.find('div', class_='item')
|
||||||
|
if not item:
|
||||||
|
continue
|
||||||
|
d2_ul = item.find('ul', class_='depth2_ul')
|
||||||
|
if not d2_ul:
|
||||||
|
continue
|
||||||
|
for d2_li in d2_ul.find_all('li', recursive=False):
|
||||||
|
d2_a = d2_li.find('a', recursive=False)
|
||||||
|
if not d2_a:
|
||||||
|
continue
|
||||||
|
E = clean_text(d2_a.get_text())
|
||||||
|
E_href = extract_href(d2_a)
|
||||||
|
d3_ul = d2_li.find('ul', class_='depth3_ul', recursive=False)
|
||||||
|
if not d3_ul:
|
||||||
|
rows.append({'D': D, 'E': E, 'href': E_href,
|
||||||
|
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
continue
|
||||||
|
rows.append({'D': D, 'E': E, 'href': E_href,
|
||||||
|
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
for d3_li in d3_ul.find_all('li', recursive=False):
|
||||||
|
d3_a = d3_li.find('a', recursive=False)
|
||||||
|
if not d3_a:
|
||||||
|
continue
|
||||||
|
F = clean_text(d3_a.get_text())
|
||||||
|
rows.append({'D': D, 'E': E, 'F': F, 'href': extract_href(d3_a),
|
||||||
|
'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def parse_cheonan_depth(soup, base):
|
||||||
|
"""천안시: ul.depth1-ul > li > a.depth1-btn + div.depth1-content > div.layout > ul.depth2-ul > li > a.depth2-btn + div.depth2-content > ul.depth3-ul > li > a.depth3-btn."""
|
||||||
|
rows = []
|
||||||
|
container = soup.select_one('ul.depth1-ul')
|
||||||
|
if not container:
|
||||||
|
return rows
|
||||||
|
for top_li in container.find_all('li', recursive=False):
|
||||||
|
a1 = top_li.find('a', class_='depth1-btn', recursive=False)
|
||||||
|
D = clean_text(a1.get_text()) if a1 else ''
|
||||||
|
D_href = extract_href(a1) if a1 else ''
|
||||||
|
d2_ul = top_li.select_one('ul.depth2-ul')
|
||||||
|
if not d2_ul:
|
||||||
|
rows.append({'D': D, 'E': '', 'href': D_href,
|
||||||
|
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
continue
|
||||||
|
for d2_li in d2_ul.find_all('li', recursive=False):
|
||||||
|
d2_a = d2_li.find('a', class_='depth2-btn', recursive=False)
|
||||||
|
if not d2_a:
|
||||||
|
continue
|
||||||
|
E = clean_text(d2_a.get_text())
|
||||||
|
E_href = extract_href(d2_a)
|
||||||
|
d3_ul = d2_li.select_one('ul.depth3-ul')
|
||||||
|
if not d3_ul:
|
||||||
|
rows.append({'D': D, 'E': E, 'href': E_href,
|
||||||
|
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
continue
|
||||||
|
rows.append({'D': D, 'E': E, 'href': E_href,
|
||||||
|
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
for d3_li in d3_ul.find_all('li', recursive=False):
|
||||||
|
d3_a = d3_li.find('a', class_='depth3-btn', recursive=False)
|
||||||
|
if not d3_a:
|
||||||
|
continue
|
||||||
|
F = clean_text(d3_a.get_text())
|
||||||
|
rows.append({'D': D, 'E': E, 'F': F, 'href': extract_href(d3_a),
|
||||||
|
'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def parse_cheongyang_depth1(soup, base):
|
||||||
|
"""청양군: ul.depth1_ul > li > a.th_1st + div.item > ul.depth2_ul > li > a + ul.depth3_ul > li > a (예산군 형)."""
|
||||||
|
# 청양군은 예산군과 동일한 e-Gov GNB 구조 — yesan 파서 재사용
|
||||||
|
return parse_yesan_depth(soup, base)
|
||||||
|
|
||||||
|
|
||||||
|
def parse_asan_gnb(soup, base):
|
||||||
|
"""아산시: div#mGnb-anchor{n}.gnb-sub-list > ul > li > a.gnb-sub-trigger + ul.sub-ul > li > a.subm."""
|
||||||
|
rows = []
|
||||||
|
# 대분류 이름 — mobile-nav의 gnb-main-trigger 텍스트 + href(#mGnb-anchorN) 매핑
|
||||||
|
nav_main = soup.select_one('nav#mobile-nav')
|
||||||
|
main_categories = [] # [(D, anchor_id)]
|
||||||
|
if nav_main:
|
||||||
|
for trig in nav_main.find_all(['a', 'button'], class_='gnb-main-trigger'):
|
||||||
|
text = clean_text(trig.get_text())
|
||||||
|
target = trig.get('href') or trig.get('data-target') or ''
|
||||||
|
if target.startswith('#mGnb-anchor'):
|
||||||
|
main_categories.append((text, target.lstrip('#')))
|
||||||
|
# Fallback: sections without name mapping
|
||||||
|
if not main_categories:
|
||||||
|
for sec in soup.select('div[id^=mGnb-anchor]'):
|
||||||
|
main_categories.append((sec.get('id'), sec.get('id')))
|
||||||
|
|
||||||
|
for D, anchor_id in main_categories:
|
||||||
|
section = soup.find('div', id=anchor_id)
|
||||||
|
if not section:
|
||||||
|
continue
|
||||||
|
# section > ul > li > a.gnb-sub-trigger + ul.sub-ul > li > a.subm
|
||||||
|
for ul in section.find_all('ul', recursive=False):
|
||||||
|
for li in ul.find_all('li', recursive=False):
|
||||||
|
a_E = li.find('a', class_='gnb-sub-trigger', recursive=False) or li.find('a', recursive=False)
|
||||||
|
if not a_E:
|
||||||
|
continue
|
||||||
|
E = clean_text(a_E.get_text())
|
||||||
|
E_href = extract_href(a_E)
|
||||||
|
sub_ul = li.find('ul', class_='sub-ul', recursive=False) or li.find('ul', recursive=False)
|
||||||
|
if not sub_ul:
|
||||||
|
rows.append({'D': D, 'E': E, 'href': E_href,
|
||||||
|
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
continue
|
||||||
|
rows.append({'D': D, 'E': E, 'href': E_href,
|
||||||
|
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
for sub_li in sub_ul.find_all('li', recursive=False):
|
||||||
|
sub_a = sub_li.find('a', recursive=False)
|
||||||
|
if not sub_a:
|
||||||
|
continue
|
||||||
|
F = clean_text(sub_a.get_text())
|
||||||
|
rows.append({'D': D, 'E': E, 'F': F, 'href': extract_href(sub_a),
|
||||||
|
'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def parse_buyeo_topmenu(soup, base):
|
||||||
|
"""부여군: ul#tm > li.th1 > a.th1_lnk (D) + div.summry > ul.th2 > li > a.th2_lnk (E) + ul.th3 > li > a (F)."""
|
||||||
|
rows = []
|
||||||
|
tm = soup.select_one('ul#tm')
|
||||||
|
if not tm:
|
||||||
|
return rows
|
||||||
|
for top_li in tm.find_all('li', class_=re.compile(r'th1?'), recursive=False):
|
||||||
|
a1 = top_li.find('a', class_='th1_lnk', recursive=False)
|
||||||
|
if not a1:
|
||||||
|
a1 = top_li.find('a', recursive=False)
|
||||||
|
if not a1:
|
||||||
|
continue
|
||||||
|
D = clean_text(a1.get_text())
|
||||||
|
D_href = extract_href(a1)
|
||||||
|
# Find ul.th2 inside div.summry
|
||||||
|
summry = top_li.find('div', class_=re.compile(r'summry'), recursive=False)
|
||||||
|
th2 = (summry.find('ul', class_='th2') if summry else None) or top_li.find('ul', class_='th2')
|
||||||
|
if not th2:
|
||||||
|
rows.append({'D': D, 'E': '', 'href': D_href,
|
||||||
|
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
continue
|
||||||
|
for li2 in th2.find_all('li', recursive=False):
|
||||||
|
a2 = li2.find('a', class_='th2_lnk', recursive=False) or li2.find('a', recursive=False)
|
||||||
|
if not a2:
|
||||||
|
continue
|
||||||
|
E = clean_text(a2.get_text())
|
||||||
|
E_href = extract_href(a2)
|
||||||
|
th3 = li2.find('ul', class_='th3', recursive=False)
|
||||||
|
if not th3:
|
||||||
|
rows.append({'D': D, 'E': E, 'href': E_href,
|
||||||
|
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
continue
|
||||||
|
rows.append({'D': D, 'E': E, 'href': E_href,
|
||||||
|
'F': '', 'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
for li3 in th3.find_all('li', recursive=False):
|
||||||
|
a3 = li3.find('a', recursive=False)
|
||||||
|
if not a3:
|
||||||
|
continue
|
||||||
|
F = clean_text(a3.get_text())
|
||||||
|
rows.append({'D': D, 'E': E, 'F': F, 'href': extract_href(a3),
|
||||||
|
'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
# ================================================================
|
||||||
|
# 사이트 설정
|
||||||
|
# ================================================================
|
||||||
|
|
||||||
|
SITES = [
|
||||||
|
{
|
||||||
|
'idx': 5, 'name': '당진시', 'base': 'https://www.dangjin.go.kr',
|
||||||
|
'sitemap': 'https://www.dangjin.go.kr/kor/sitemap_11.do',
|
||||||
|
'sheet': '05_당진시', 'parser': parse_amThum,
|
||||||
|
'domain': 'dangjin.go.kr',
|
||||||
|
'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\5.당진시',
|
||||||
|
},
|
||||||
|
{
|
||||||
|
'idx': 6, 'name': '보령시', 'base': 'https://www.brcn.go.kr',
|
||||||
|
'sitemap': 'https://www.brcn.go.kr/kor/sitemap_11.do',
|
||||||
|
'sheet': '06_보령시', 'parser': parse_ul_sitemap_h4,
|
||||||
|
'domain': 'brcn.go.kr',
|
||||||
|
'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\6.보령시',
|
||||||
|
},
|
||||||
|
{
|
||||||
|
'idx': 7, 'name': '부여군', 'base': 'https://www.buyeo.go.kr',
|
||||||
|
'sitemap': 'https://www.buyeo.go.kr/html/kr/',
|
||||||
|
'sheet': '07_부여군', 'parser': parse_buyeo_topmenu,
|
||||||
|
'domain': 'buyeo.go.kr',
|
||||||
|
'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\7.부여군',
|
||||||
|
},
|
||||||
|
{
|
||||||
|
'idx': 8, 'name': '서산시', 'base': 'https://www.seosan.go.kr',
|
||||||
|
'sitemap': 'https://www.seosan.go.kr/www/index.do',
|
||||||
|
'sheet': '08_서산시', 'parser': parse_seosan_top_menu,
|
||||||
|
'domain': 'seosan.go.kr',
|
||||||
|
'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\8.서산시',
|
||||||
|
},
|
||||||
|
{
|
||||||
|
'idx': 9, 'name': '서천군', 'base': 'https://www.seocheon.go.kr',
|
||||||
|
'sitemap': 'https://www.seocheon.go.kr/kor/sitemap_11.do',
|
||||||
|
'sheet': '09_서천군', 'parser': parse_ul_sitemap_h4,
|
||||||
|
'domain': 'seocheon.go.kr',
|
||||||
|
'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\9.서천군',
|
||||||
|
},
|
||||||
|
{
|
||||||
|
'idx': 10, 'name': '아산시', 'base': 'https://www.asan.go.kr',
|
||||||
|
'sitemap': 'https://www.asan.go.kr/main/',
|
||||||
|
'sheet': '10_아산시', 'parser': parse_asan_gnb,
|
||||||
|
'domain': 'asan.go.kr',
|
||||||
|
'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\10.아산시',
|
||||||
|
},
|
||||||
|
{
|
||||||
|
'idx': 11, 'name': '예산군', 'base': 'https://www.yesan.go.kr',
|
||||||
|
'sitemap': 'https://www.yesan.go.kr/kor/sitemap.do',
|
||||||
|
'sheet': '11_예산군', 'parser': parse_yesan_depth,
|
||||||
|
'domain': 'yesan.go.kr',
|
||||||
|
'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\11.예산군',
|
||||||
|
},
|
||||||
|
{
|
||||||
|
'idx': 12, 'name': '천안시', 'base': 'https://www.cheonan.go.kr',
|
||||||
|
'sitemap': 'https://www.cheonan.go.kr/kor/sitemap.do',
|
||||||
|
'sheet': '12_천안시', 'parser': parse_cheonan_depth,
|
||||||
|
'domain': 'cheonan.go.kr',
|
||||||
|
'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\12.천안시',
|
||||||
|
},
|
||||||
|
{
|
||||||
|
'idx': 13, 'name': '청양군', 'base': 'https://www.cheongyang.go.kr',
|
||||||
|
'sitemap': 'https://www.cheongyang.go.kr/kor/sitemap_11.do',
|
||||||
|
'sheet': '13_청양군', 'parser': parse_cheongyang_depth1,
|
||||||
|
'domain': 'cheongyang.go.kr',
|
||||||
|
'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\13.청양군',
|
||||||
|
},
|
||||||
|
{
|
||||||
|
'idx': 14, 'name': '태안군', 'base': 'https://www.taean.go.kr',
|
||||||
|
'sitemap': 'https://www.taean.go.kr/kor/sitemap_11.do',
|
||||||
|
'sheet': '14_태안군', 'parser': parse_amThum,
|
||||||
|
'domain': 'taean.go.kr',
|
||||||
|
'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\14.태안군',
|
||||||
|
},
|
||||||
|
{
|
||||||
|
'idx': 15, 'name': '홍성군', 'base': 'https://www.hongseong.go.kr',
|
||||||
|
'sitemap': 'https://www.hongseong.go.kr/kor/sitemap.do',
|
||||||
|
'sheet': '15_홍성군',
|
||||||
|
'parser': lambda soup, base: parse_dl_dt_dd(soup, base, use_onclick=True),
|
||||||
|
'domain': 'hongseong.go.kr',
|
||||||
|
'folder': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\15.홍성군',
|
||||||
|
},
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
# ================================================================
|
||||||
|
# 엑셀 생성 (공통)
|
||||||
|
# ================================================================
|
||||||
|
|
||||||
|
|
||||||
|
def write_excel(site, raw_rows):
|
||||||
|
name = site['name']
|
||||||
|
base = site['base']
|
||||||
|
domain = site['domain']
|
||||||
|
output = f"{site['folder']}\\{os.path.basename(os.path.dirname(site['folder']))}_{name}.xlsx"
|
||||||
|
|
||||||
|
def abs_url(href):
|
||||||
|
if not href:
|
||||||
|
return ''
|
||||||
|
href = href.strip()
|
||||||
|
if href.startswith(('javascript:', '#')):
|
||||||
|
return ''
|
||||||
|
if href.startswith(('http://', 'https://')):
|
||||||
|
return href
|
||||||
|
return urljoin(base + '/', href)
|
||||||
|
|
||||||
|
def is_external(url):
|
||||||
|
return url.startswith(('http://', 'https://')) and domain not in url
|
||||||
|
|
||||||
|
# 부모-자식 URL 중복 제거
|
||||||
|
final_rows = []
|
||||||
|
i = 0
|
||||||
|
removed = 0
|
||||||
|
while i < len(raw_rows):
|
||||||
|
row = raw_rows[i]
|
||||||
|
if (i + 1 < len(raw_rows)
|
||||||
|
and row.get('G', '') == ''
|
||||||
|
and raw_rows[i + 1].get('D') == row.get('D')
|
||||||
|
and raw_rows[i + 1].get('E') == row.get('E')
|
||||||
|
and raw_rows[i + 1].get('F') == row.get('F')
|
||||||
|
and raw_rows[i + 1].get('G', '') != ''
|
||||||
|
and raw_rows[i + 1].get('href') == row.get('href')):
|
||||||
|
removed += 1
|
||||||
|
i += 1
|
||||||
|
continue
|
||||||
|
final_rows.append(row)
|
||||||
|
i += 1
|
||||||
|
|
||||||
|
print(f' [{name}] 원본 {len(raw_rows)} → 중복 {removed} → 최종 {len(final_rows)}')
|
||||||
|
if not final_rows:
|
||||||
|
print(f' [{name}] !! 행 0개 — 파서 점검 필요. 엑셀 생성 스킵.')
|
||||||
|
return False
|
||||||
|
|
||||||
|
shutil.copy(TEMPLATE, output)
|
||||||
|
wb = openpyxl.load_workbook(output)
|
||||||
|
ws = wb.active
|
||||||
|
ws.title = site['sheet']
|
||||||
|
|
||||||
|
HEADER = {'B1:R1', 'S1:W1', 'Y1:AA1'}
|
||||||
|
for rng in [str(m) for m in ws.merged_cells.ranges if str(m) not in HEADER]:
|
||||||
|
ws.unmerge_cells(rng)
|
||||||
|
for row in ws.iter_rows(min_row=3, max_row=ws.max_row, min_col=1, max_col=ws.max_column):
|
||||||
|
for cell in row:
|
||||||
|
cell.value = None
|
||||||
|
|
||||||
|
START = 3
|
||||||
|
template_r = 3
|
||||||
|
cur_max = ws.max_row
|
||||||
|
for idx, item in enumerate(final_rows, start=START):
|
||||||
|
if idx > cur_max:
|
||||||
|
for c in range(1, ws.max_column + 1):
|
||||||
|
srcc = ws.cell(template_r, c)
|
||||||
|
tgt = ws.cell(idx, c)
|
||||||
|
if srcc.has_style:
|
||||||
|
tgt.font = copy(srcc.font)
|
||||||
|
tgt.fill = copy(srcc.fill)
|
||||||
|
tgt.border = copy(srcc.border)
|
||||||
|
tgt.alignment = copy(srcc.alignment)
|
||||||
|
tgt.number_format = srcc.number_format
|
||||||
|
tgt.protection = copy(srcc.protection)
|
||||||
|
url = abs_url(item.get('href', ''))
|
||||||
|
ws.cell(idx, 2).value = idx - 2
|
||||||
|
ws.cell(idx, 3).value = name
|
||||||
|
ws.cell(idx, 4).value = item.get('D', '')
|
||||||
|
ws.cell(idx, 5).value = item.get('E', '')
|
||||||
|
ws.cell(idx, 6).value = item.get('F', '')
|
||||||
|
ws.cell(idx, 7).value = item.get('G', '')
|
||||||
|
ws.cell(idx, 8).value = item.get('H', '')
|
||||||
|
ws.cell(idx, 9).value = item.get('I', '')
|
||||||
|
ws.cell(idx, 10).value = item.get('J', '')
|
||||||
|
ws.cell(idx, 11).value = url
|
||||||
|
if is_external(url):
|
||||||
|
ws.cell(idx, 19).value = '외부링크'
|
||||||
|
|
||||||
|
END = START + len(final_rows) - 1
|
||||||
|
center = Alignment(horizontal='center', vertical='center', wrap_text=True)
|
||||||
|
left = Alignment(horizontal='left', vertical='center', wrap_text=False)
|
||||||
|
|
||||||
|
def merge_runs(col_letter, col_idx, group_cols=()):
|
||||||
|
runs = []
|
||||||
|
cur_val = ws.cell(START, col_idx).value
|
||||||
|
cur_grp = tuple(ws.cell(START, g).value for g in group_cols)
|
||||||
|
run_start = START
|
||||||
|
for r in range(START + 1, END + 1):
|
||||||
|
v = ws.cell(r, col_idx).value
|
||||||
|
g = tuple(ws.cell(r, gg).value for gg in group_cols)
|
||||||
|
if v == cur_val and g == cur_grp:
|
||||||
|
continue
|
||||||
|
if cur_val not in (None, '') and r - 1 > run_start:
|
||||||
|
runs.append((run_start, r - 1))
|
||||||
|
cur_val, cur_grp, run_start = v, g, r
|
||||||
|
if cur_val not in (None, '') and END > run_start:
|
||||||
|
runs.append((run_start, END))
|
||||||
|
for s, e in runs:
|
||||||
|
ws.merge_cells(f'{col_letter}{s}:{col_letter}{e}')
|
||||||
|
ws.cell(s, col_idx).alignment = center
|
||||||
|
return len(runs)
|
||||||
|
|
||||||
|
# 병합 순서 F→E→D (D를 먼저 병합하면 2행부터 D=None이 되어 E 그룹키가 깨짐)
|
||||||
|
n_f = merge_runs('F', 6, group_cols=(4, 5))
|
||||||
|
n_e = merge_runs('E', 5, group_cols=(4,))
|
||||||
|
n_d = merge_runs('D', 4)
|
||||||
|
for r in range(START, END + 1):
|
||||||
|
for c in (4, 5, 6, 7):
|
||||||
|
if ws.cell(r, c).value is not None:
|
||||||
|
ws.cell(r, c).alignment = center
|
||||||
|
|
||||||
|
for r in range(1, END + 1):
|
||||||
|
ws.row_dimensions[r].height = 15
|
||||||
|
|
||||||
|
link_n = 0
|
||||||
|
for r in range(START, END + 1):
|
||||||
|
cell = ws.cell(r, 11)
|
||||||
|
u = cell.value
|
||||||
|
if u and isinstance(u, str) and u.startswith(('http://', 'https://')):
|
||||||
|
cell.hyperlink = u
|
||||||
|
old = cell.font
|
||||||
|
cell.font = Font(name=old.name or '맑은 고딕', size=old.size or 11,
|
||||||
|
bold=old.bold, italic=old.italic,
|
||||||
|
color='0000FF', underline='single')
|
||||||
|
cell.alignment = left
|
||||||
|
link_n += 1
|
||||||
|
|
||||||
|
wb.save(output)
|
||||||
|
ext_n = sum(1 for r in range(START, END + 1) if ws.cell(r, 19).value == '외부링크')
|
||||||
|
print(f' [{name}] 병합 D:{n_d} E:{n_e} F:{n_f} | 외부링크 {ext_n} | K링크 {link_n} → {output}')
|
||||||
|
return True
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
targets = sys.argv[1:] if len(sys.argv) > 1 else [s['name'] for s in SITES]
|
||||||
|
for site in SITES:
|
||||||
|
if site['name'] not in targets:
|
||||||
|
continue
|
||||||
|
print(f"\n{'='*60}\n[{site['idx']}.{site['name']}] {site['sitemap']}\n{'='*60}")
|
||||||
|
try:
|
||||||
|
html = fetch_html(site['sitemap'])
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
raw_rows = site['parser'](soup, site['base'])
|
||||||
|
write_excel(site, raw_rows)
|
||||||
|
except Exception as e:
|
||||||
|
print(f' [{site["name"]}] !! 실패: {e}')
|
||||||
|
import traceback
|
||||||
|
traceback.print_exc()
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
378
_스크립트/_chungnam_phase234_all.py
Normal file
378
_스크립트/_chungnam_phase234_all.py
Normal file
@ -0,0 +1,378 @@
|
|||||||
|
"""충청남도 11개 시·군 Phase 2~4 일괄 처리 (L/M/N/O/P/Q).
|
||||||
|
|
||||||
|
매뉴얼: D:\\01.프로젝트\\DB수집\\사이트맵_수집_매뉴얼.md
|
||||||
|
입력: 각 폴더의 {기관명}.xlsx
|
||||||
|
처리: K열 URL 접근 → L(게시판형태), M(수량), N(저작물 유형), O/P/Q(공공누리)
|
||||||
|
"""
|
||||||
|
import re
|
||||||
|
import sys
|
||||||
|
import time
|
||||||
|
import warnings
|
||||||
|
from urllib.parse import urljoin
|
||||||
|
from concurrent.futures import ThreadPoolExecutor, as_completed
|
||||||
|
|
||||||
|
import openpyxl
|
||||||
|
import requests
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
|
||||||
|
# ================================================================
|
||||||
|
# 정규식
|
||||||
|
# ================================================================
|
||||||
|
TOTAL_PAT = re.compile(r'총\s*(?:게시물\s*)?(\d[\d,]*)\s*(?:건|개|page|페이지)', re.I)
|
||||||
|
TOTAL_PAT_LOOSE = re.compile(r'(?:전체|총)\s*[:\-]?\s*(\d[\d,]*)\s*건', re.I)
|
||||||
|
KOGL_IMG_PAT = re.compile(r'(?:new_)?img_opentype(\d{2})\.png', re.I)
|
||||||
|
KOGL_LINK_PAT = re.compile(r'kogl\.or\.kr/info/licenseType(\d)', re.I)
|
||||||
|
YOUTUBE_PAT = re.compile(r'(?:youtube\.com|youtu\.be)', re.I)
|
||||||
|
VIDEO_EXT = re.compile(r'\.(mp4|webm|mov|avi)(?:\?|$)', re.I)
|
||||||
|
|
||||||
|
|
||||||
|
# ================================================================
|
||||||
|
# 사이트 설정
|
||||||
|
# ================================================================
|
||||||
|
|
||||||
|
SITES = {
|
||||||
|
'논산시': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\4.논산시\충청남도_논산시.xlsx',
|
||||||
|
'body_sel': ['#txt', '#contents', 'main']},
|
||||||
|
'당진시': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\5.당진시\충청남도_당진시.xlsx',
|
||||||
|
'body_sel': ['#txt', '#contents', 'main']},
|
||||||
|
'보령시': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\6.보령시\충청남도_보령시.xlsx',
|
||||||
|
'body_sel': ['#txt', '#contents', 'main']},
|
||||||
|
'부여군': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\7.부여군\충청남도_부여군.xlsx',
|
||||||
|
'body_sel': ['#txt', '#contents', 'main']},
|
||||||
|
'서산시': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\8.서산시\충청남도_서산시.xlsx',
|
||||||
|
'body_sel': ['#contents', '#txt', 'main']},
|
||||||
|
'서천군': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\9.서천군\충청남도_서천군.xlsx',
|
||||||
|
'body_sel': ['#txt', '#contents', 'main']},
|
||||||
|
'아산시': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\10.아산시\충청남도_아산시.xlsx',
|
||||||
|
'body_sel': ['.contents', '#contents', '#txt', 'main']},
|
||||||
|
'예산군': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\11.예산군\충청남도_예산군.xlsx',
|
||||||
|
'body_sel': ['#txt', '#contents', 'main']},
|
||||||
|
'천안시': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\12.천안시\충청남도_천안시.xlsx',
|
||||||
|
'body_sel': ['#txt', '#contents', 'main']},
|
||||||
|
'청양군': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\13.청양군\충청남도_청양군.xlsx',
|
||||||
|
'body_sel': ['#txt', '#contents', 'main']},
|
||||||
|
'태안군': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\14.태안군\충청남도_태안군.xlsx',
|
||||||
|
'body_sel': ['#txt', '#contents', 'main']},
|
||||||
|
'홍성군': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\15.홍성군\충청남도_홍성군.xlsx',
|
||||||
|
'body_sel': ['#txt', '#contents', 'main']},
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
# ================================================================
|
||||||
|
# 크롤링 공통 함수
|
||||||
|
# ================================================================
|
||||||
|
|
||||||
|
|
||||||
|
def fetch(url, timeout=12):
|
||||||
|
try:
|
||||||
|
r = requests.get(url, headers=H, timeout=timeout, verify=False, allow_redirects=True)
|
||||||
|
r.encoding = r.apparent_encoding
|
||||||
|
if r.status_code == 200:
|
||||||
|
return BeautifulSoup(r.text, 'html.parser')
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def get_body(soup, selectors):
|
||||||
|
for sel in selectors:
|
||||||
|
el = soup.select_one(sel)
|
||||||
|
if el:
|
||||||
|
return el
|
||||||
|
return soup
|
||||||
|
|
||||||
|
|
||||||
|
def detect_form(body):
|
||||||
|
"""L 판별. 페이징·검색·총건수 3요소 중 하나라도 있으면 '게시판'."""
|
||||||
|
has_paging = bool(body.select('.pagination, .paging, nav.paging, .page_nav'))
|
||||||
|
text_inputs = [i for i in body.find_all('input')
|
||||||
|
if (i.get('type') or 'text').lower() in ('text', 'search')]
|
||||||
|
has_search = len(text_inputs) >= 1
|
||||||
|
txt = body.get_text(' ', strip=True)
|
||||||
|
m = TOTAL_PAT.search(txt)
|
||||||
|
if not m:
|
||||||
|
m = TOTAL_PAT_LOOSE.search(txt)
|
||||||
|
total = None
|
||||||
|
if m:
|
||||||
|
digits = m.group(1).replace(',', '')
|
||||||
|
if digits.isdigit():
|
||||||
|
total = int(digits)
|
||||||
|
is_board = has_paging or has_search or (total is not None)
|
||||||
|
if is_board:
|
||||||
|
return '게시판', total if total is not None else 0
|
||||||
|
return '페이지', 1
|
||||||
|
|
||||||
|
|
||||||
|
DETAIL_PAT = re.compile(r'(mode=V|view\.do|bbtSn=|seqRepeat=|nttId=|articleNo=|boardSeq=)', re.I)
|
||||||
|
|
||||||
|
|
||||||
|
def extract_detail_urls(body, base_url, limit=5):
|
||||||
|
urls = []
|
||||||
|
seen = set()
|
||||||
|
for a in body.find_all('a', href=True):
|
||||||
|
h = a['href']
|
||||||
|
if not h or h.startswith('#'):
|
||||||
|
continue
|
||||||
|
if DETAIL_PAT.search(h):
|
||||||
|
full = urljoin(base_url, h)
|
||||||
|
if full not in seen:
|
||||||
|
seen.add(full)
|
||||||
|
urls.append(full)
|
||||||
|
if len(urls) >= limit:
|
||||||
|
break
|
||||||
|
# Also try fn_search_detail JS pattern (공주시 형)
|
||||||
|
if len(urls) < limit:
|
||||||
|
for a in body.find_all('a'):
|
||||||
|
onclick = a.get('onclick', '')
|
||||||
|
m = re.search(r"fn_(?:search_)?detail\(['\"]([^'\"]+)['\"]", onclick)
|
||||||
|
if m:
|
||||||
|
ntt_id = m.group(1)
|
||||||
|
# Build URL by replacing list.do with view.do?nttId=
|
||||||
|
view_url = re.sub(r'list\.do[^\'"]*', f'view.do?nttId={ntt_id}', base_url)
|
||||||
|
if view_url not in seen:
|
||||||
|
seen.add(view_url)
|
||||||
|
urls.append(view_url)
|
||||||
|
if len(urls) >= limit:
|
||||||
|
break
|
||||||
|
return urls
|
||||||
|
|
||||||
|
|
||||||
|
def detect_media(body):
|
||||||
|
has_text = len(body.get_text(strip=True)) > 30
|
||||||
|
has_image = False
|
||||||
|
for img in body.find_all('img'):
|
||||||
|
src = img.get('src', '')
|
||||||
|
if KOGL_IMG_PAT.search(src):
|
||||||
|
continue
|
||||||
|
if not src:
|
||||||
|
continue
|
||||||
|
has_image = True
|
||||||
|
break
|
||||||
|
has_video = False
|
||||||
|
for iframe in body.find_all('iframe'):
|
||||||
|
if YOUTUBE_PAT.search(iframe.get('src', '')):
|
||||||
|
has_video = True
|
||||||
|
break
|
||||||
|
if not has_video:
|
||||||
|
for a in body.find_all('a', href=True):
|
||||||
|
if YOUTUBE_PAT.search(a['href']):
|
||||||
|
has_video = True
|
||||||
|
break
|
||||||
|
if not has_video and body.find_all('video'):
|
||||||
|
has_video = True
|
||||||
|
if not has_video and VIDEO_EXT.search(str(body)):
|
||||||
|
has_video = True
|
||||||
|
return has_image, has_video, has_text
|
||||||
|
|
||||||
|
|
||||||
|
def n_string(has_text, has_image, has_video):
|
||||||
|
parts = []
|
||||||
|
if has_text:
|
||||||
|
parts.append('어문')
|
||||||
|
if has_image:
|
||||||
|
parts.append('이미지')
|
||||||
|
if has_video:
|
||||||
|
parts.append('영상')
|
||||||
|
return ','.join(parts) if parts else '없음'
|
||||||
|
|
||||||
|
|
||||||
|
def img_has_valid_anchor(img):
|
||||||
|
p = img.parent
|
||||||
|
while p is not None:
|
||||||
|
if p.name == 'a':
|
||||||
|
href = p.get('href', '')
|
||||||
|
if href and not href.startswith('#') and not href.lower().startswith('javascript:'):
|
||||||
|
return True
|
||||||
|
return False
|
||||||
|
p = p.parent
|
||||||
|
return False
|
||||||
|
|
||||||
|
|
||||||
|
def detect_kogl(body):
|
||||||
|
types = set()
|
||||||
|
q_any_y = False
|
||||||
|
q_any_n = False
|
||||||
|
for a in body.find_all('a', href=True):
|
||||||
|
m = KOGL_LINK_PAT.search(a['href'])
|
||||||
|
if m:
|
||||||
|
types.add(int(m.group(1)))
|
||||||
|
q_any_y = True
|
||||||
|
for img in body.find_all('img'):
|
||||||
|
src = img.get('src', '')
|
||||||
|
m = KOGL_IMG_PAT.search(src)
|
||||||
|
if m:
|
||||||
|
types.add(int(m.group(1)))
|
||||||
|
if img_has_valid_anchor(img):
|
||||||
|
q_any_y = True
|
||||||
|
else:
|
||||||
|
q_any_n = True
|
||||||
|
for el in body.find_all(style=True):
|
||||||
|
m = KOGL_IMG_PAT.search(el.get('style', ''))
|
||||||
|
if m:
|
||||||
|
types.add(int(m.group(1)))
|
||||||
|
q_any_n = True
|
||||||
|
if not types:
|
||||||
|
return set(), None
|
||||||
|
return types, ('Y' if q_any_y else 'N')
|
||||||
|
|
||||||
|
|
||||||
|
def process_row(url, body_selectors):
|
||||||
|
out = {'L': '', 'M': '', 'N': '', 'O': '', 'P': '', 'Q': '', 'note': ''}
|
||||||
|
soup = fetch(url)
|
||||||
|
if soup is None:
|
||||||
|
out['note'] = '접근 실패'
|
||||||
|
return out
|
||||||
|
body = get_body(soup, body_selectors)
|
||||||
|
|
||||||
|
form, count = detect_form(body)
|
||||||
|
out['L'] = form
|
||||||
|
out['M'] = count if form == '게시판' else 1
|
||||||
|
|
||||||
|
has_img, has_vid, has_txt = detect_media(body)
|
||||||
|
types_main, q_main = detect_kogl(body)
|
||||||
|
P = '게시판' if types_main else ''
|
||||||
|
types_all = set(types_main)
|
||||||
|
q_flags = []
|
||||||
|
if q_main:
|
||||||
|
q_flags.append(q_main)
|
||||||
|
|
||||||
|
if form == '게시판':
|
||||||
|
detail_urls = extract_detail_urls(body, url, limit=5)
|
||||||
|
for du in detail_urls:
|
||||||
|
d_soup = fetch(du, timeout=10)
|
||||||
|
if not d_soup:
|
||||||
|
continue
|
||||||
|
d_body = get_body(d_soup, body_selectors)
|
||||||
|
di, dv, dt = detect_media(d_body)
|
||||||
|
has_img = has_img or di
|
||||||
|
has_vid = has_vid or dv
|
||||||
|
has_txt = has_txt or dt
|
||||||
|
dt_types, dt_q = detect_kogl(d_body)
|
||||||
|
if dt_types and not types_main and not P:
|
||||||
|
P = '게시물'
|
||||||
|
types_all |= dt_types
|
||||||
|
if dt_q:
|
||||||
|
q_flags.append(dt_q)
|
||||||
|
|
||||||
|
out['N'] = n_string(has_txt, has_img, has_vid)
|
||||||
|
|
||||||
|
if not types_all:
|
||||||
|
out['O'] = '미부착'
|
||||||
|
else:
|
||||||
|
sorted_types = sorted(types_all)
|
||||||
|
out['O'] = ','.join(f'{n}유형' for n in sorted_types)
|
||||||
|
out['P'] = P if P else '게시판'
|
||||||
|
out['Q'] = 'Y' if 'Y' in q_flags else 'N'
|
||||||
|
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
# ================================================================
|
||||||
|
# 사이트 단위 실행
|
||||||
|
# ================================================================
|
||||||
|
|
||||||
|
|
||||||
|
def run_site(name, xlsx, body_selectors, workers=10):
|
||||||
|
print(f'\n{"="*60}\n[{name}] {xlsx}\n{"="*60}')
|
||||||
|
wb = openpyxl.load_workbook(xlsx)
|
||||||
|
ws = wb.active
|
||||||
|
START = 3
|
||||||
|
# 진짜 데이터 행만 가져옴 (B열에 순번 있어야)
|
||||||
|
END = START - 1
|
||||||
|
for r in range(START, ws.max_row + 1):
|
||||||
|
if ws.cell(r, 2).value is None:
|
||||||
|
break
|
||||||
|
END = r
|
||||||
|
|
||||||
|
tasks = []
|
||||||
|
for r in range(START, END + 1):
|
||||||
|
url = ws.cell(r, 11).value
|
||||||
|
is_ext = (ws.cell(r, 19).value == '외부링크')
|
||||||
|
tasks.append((r, url, is_ext))
|
||||||
|
n_ext = sum(1 for t in tasks if t[2])
|
||||||
|
print(f' 처리 대상: 총 {len(tasks)}행 (외부링크 {n_ext})')
|
||||||
|
|
||||||
|
t0 = time.time()
|
||||||
|
results = {}
|
||||||
|
|
||||||
|
def worker(task):
|
||||||
|
row, url, is_ext = task
|
||||||
|
if is_ext:
|
||||||
|
return row, {'L': '사이트', 'M': '', 'N': '', 'O': '', 'P': '', 'Q': '', 'note': ''}
|
||||||
|
if not url or not isinstance(url, str):
|
||||||
|
return row, {'L': '', 'M': '', 'N': '', 'O': '', 'P': '', 'Q': '', 'note': 'URL 없음'}
|
||||||
|
return row, process_row(url, body_selectors)
|
||||||
|
|
||||||
|
done = 0
|
||||||
|
with ThreadPoolExecutor(max_workers=workers) as ex:
|
||||||
|
futs = [ex.submit(worker, t) for t in tasks]
|
||||||
|
for fut in as_completed(futs):
|
||||||
|
row, res = fut.result()
|
||||||
|
results[row] = res
|
||||||
|
done += 1
|
||||||
|
if done % 50 == 0 or done == len(tasks):
|
||||||
|
print(f' 진행 {done}/{len(tasks)} ({time.time()-t0:.0f}s)')
|
||||||
|
|
||||||
|
for r in range(START, END + 1):
|
||||||
|
res = results.get(r, {})
|
||||||
|
if not res:
|
||||||
|
continue
|
||||||
|
if res.get('L'):
|
||||||
|
ws.cell(r, 12).value = res['L']
|
||||||
|
if res.get('M') != '':
|
||||||
|
ws.cell(r, 13).value = res['M']
|
||||||
|
if res.get('N'):
|
||||||
|
ws.cell(r, 14).value = res['N']
|
||||||
|
if res.get('O'):
|
||||||
|
ws.cell(r, 15).value = res['O']
|
||||||
|
if res.get('P'):
|
||||||
|
ws.cell(r, 16).value = res['P']
|
||||||
|
if res.get('Q'):
|
||||||
|
ws.cell(r, 17).value = res['Q']
|
||||||
|
if res.get('note'):
|
||||||
|
existing = ws.cell(r, 19).value
|
||||||
|
if not existing:
|
||||||
|
ws.cell(r, 19).value = res['note']
|
||||||
|
|
||||||
|
wb.save(xlsx)
|
||||||
|
forms = {}
|
||||||
|
attach = {'미부착': 0, '부착': 0, '기타': 0}
|
||||||
|
q_dist = {'Y': 0, 'N': 0, '': 0}
|
||||||
|
for r, res in results.items():
|
||||||
|
forms[res.get('L', '')] = forms.get(res.get('L', ''), 0) + 1
|
||||||
|
o = res.get('O', '')
|
||||||
|
if o == '미부착':
|
||||||
|
attach['미부착'] += 1
|
||||||
|
elif o and '유형' in o:
|
||||||
|
attach['부착'] += 1
|
||||||
|
else:
|
||||||
|
attach['기타'] += 1
|
||||||
|
q = res.get('Q', '')
|
||||||
|
q_dist[q] = q_dist.get(q, 0) + 1
|
||||||
|
print(f' [{name}] L분포 {forms} | 부착 {attach} | Q {q_dist} | 시간 {time.time()-t0:.0f}s')
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
targets = sys.argv[1:] if len(sys.argv) > 1 else list(SITES.keys())
|
||||||
|
total_t0 = time.time()
|
||||||
|
for name in targets:
|
||||||
|
if name not in SITES:
|
||||||
|
print(f' 알 수 없음: {name}')
|
||||||
|
continue
|
||||||
|
cfg = SITES[name]
|
||||||
|
try:
|
||||||
|
run_site(name, cfg['xlsx'], cfg['body_sel'])
|
||||||
|
except Exception as e:
|
||||||
|
print(f' [{name}] 실패: {e}')
|
||||||
|
import traceback
|
||||||
|
traceback.print_exc()
|
||||||
|
print(f'\n총 소요: {time.time()-total_t0:.0f}s')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
127
_스크립트/_collapse_all.py
Normal file
127
_스크립트/_collapse_all.py
Normal file
@ -0,0 +1,127 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""규칙 B 전 사이트 적용: F 하나에 G가 여러 개이고 그 G들이 전부 L=페이지면
|
||||||
|
→ 1행으로 합치고(G/H/I/J 비움), K=첫 G URL, 수량(M)=G들의 M 합. (공주시와 동일)
|
||||||
|
|
||||||
|
안전장치:
|
||||||
|
· G 중 하나라도 페이지가 아니면(게시판/사이트) 합치지 않음.
|
||||||
|
· 합칠 때 공공누리: O유형=그룹 합집합(숫자 오름차순,콤마,공백없음), Q=하나라도 Y면 Y,
|
||||||
|
P=게시판 우선(없으면 게시물). 부착 정보 손실 방지하며 합침.
|
||||||
|
· G가 잎일 때만(H/I/J 비어있음). 더 깊은 중첩은 건드리지 않음.
|
||||||
|
|
||||||
|
공주시(완료)·계룡시(검수완료)는 제외.
|
||||||
|
|
||||||
|
사용: python -X utf8 _collapse_all.py dry [기관...]
|
||||||
|
python -X utf8 _collapse_all.py run [기관...]
|
||||||
|
"""
|
||||||
|
import os, re, sys, shutil, importlib.util, warnings
|
||||||
|
import openpyxl
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
TYPE_PAT = re.compile(r'(\d)\s*유형')
|
||||||
|
|
||||||
|
|
||||||
|
def merge_kogl(grp):
|
||||||
|
"""그룹 G행들의 O/P/Q 통합. 반환 (O, P, Q)."""
|
||||||
|
types = set()
|
||||||
|
for g in grp:
|
||||||
|
o = g['vals'].get(15)
|
||||||
|
if isinstance(o, str):
|
||||||
|
for m in TYPE_PAT.finditer(o):
|
||||||
|
types.add(int(m.group(1)))
|
||||||
|
if types:
|
||||||
|
O = ','.join(f'{n}유형' for n in sorted(types))
|
||||||
|
else:
|
||||||
|
O = '미부착'
|
||||||
|
Ps = [g['vals'].get(16) for g in grp]
|
||||||
|
P = '게시판' if '게시판' in Ps else ('게시물' if '게시물' in Ps else None)
|
||||||
|
Qs = [g['vals'].get(17) for g in grp]
|
||||||
|
Q = 'Y' if 'Y' in Qs else ('N' if 'N' in Qs else None)
|
||||||
|
return O, (P if O != '미부착' else None), (Q if O != '미부착' else None)
|
||||||
|
|
||||||
|
spec = importlib.util.spec_from_file_location('dd', __file__.replace('_collapse_all.py', '_dedup_all.py'))
|
||||||
|
dd = importlib.util.module_from_spec(spec); spec.loader.exec_module(dd)
|
||||||
|
|
||||||
|
EXCLUDE = {'공주시', '계룡시'}
|
||||||
|
|
||||||
|
|
||||||
|
def mval(x):
|
||||||
|
try:
|
||||||
|
return int(x)
|
||||||
|
except Exception:
|
||||||
|
return 1
|
||||||
|
|
||||||
|
|
||||||
|
def collapse(rows):
|
||||||
|
out, groups, held = [], [], []
|
||||||
|
i, n = 0, len(rows)
|
||||||
|
while i < n:
|
||||||
|
v = rows[i]['vals']
|
||||||
|
D, E, F, G = v.get(4), v.get(5), v.get(6), v.get(7)
|
||||||
|
if F not in (None, '') and G not in (None, ''):
|
||||||
|
j = i
|
||||||
|
while (j < n and rows[j]['vals'].get(4) == D and rows[j]['vals'].get(5) == E
|
||||||
|
and rows[j]['vals'].get(6) == F and rows[j]['vals'].get(7) not in (None, '')):
|
||||||
|
j += 1
|
||||||
|
grp = rows[i:j]
|
||||||
|
all_page = all(g['vals'].get(12) == '페이지' for g in grp)
|
||||||
|
leaf = all(all(g['vals'].get(c) in (None, '') for c in (8, 9, 10)) for g in grp)
|
||||||
|
has_attach = any(g['vals'].get(15) not in (None, '', '미부착') for g in grp)
|
||||||
|
if len(grp) >= 2 and all_page and leaf:
|
||||||
|
msum = sum(mval(g['vals'].get(13)) for g in grp)
|
||||||
|
O, P, Q = merge_kogl(grp)
|
||||||
|
first = dict(grp[0]); nv = dict(first['vals'])
|
||||||
|
for c in (7, 8, 9, 10):
|
||||||
|
nv[c] = None
|
||||||
|
nv[13] = msum
|
||||||
|
nv[15] = O; nv[16] = P; nv[17] = Q
|
||||||
|
first['vals'] = nv
|
||||||
|
out.append(first)
|
||||||
|
groups.append({'E': E, 'F': F, 'n': len(grp), 'sum': msum, 'O': O})
|
||||||
|
if has_attach:
|
||||||
|
held.append({'E': E, 'F': F, 'O': O})
|
||||||
|
i = j
|
||||||
|
continue
|
||||||
|
out.extend(grp); i = j; continue
|
||||||
|
out.append(rows[i]); i += 1
|
||||||
|
return out, groups, held
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
mode = sys.argv[1] if len(sys.argv) > 1 else 'dry'
|
||||||
|
only = set(a for a in sys.argv[2:] if not a.startswith('--'))
|
||||||
|
sites = [s for s in dd.SITES if s[2] not in EXCLUDE and (not only or s[2] in only)]
|
||||||
|
print(f'대상 기관: {len(sites)}개 (공주·계룡 제외) 모드={mode}\n')
|
||||||
|
print(f'{"기관":<8}{"현재":>6}{"합친그룹":>7}{"제거행":>6}{"→남음":>7}{"부착합침":>7}')
|
||||||
|
print('-' * 55)
|
||||||
|
grand_g = grand_r = grand_h = 0
|
||||||
|
attach_detail = []
|
||||||
|
for prov, idx, name in sites:
|
||||||
|
xp = dd.xpath(prov, idx, name)
|
||||||
|
if not os.path.exists(xp):
|
||||||
|
print(f'{name:<8} 엑셀 없음'); continue
|
||||||
|
wb = openpyxl.load_workbook(xp); ws = wb.active
|
||||||
|
rows = dd.load_flat(ws)
|
||||||
|
out, groups, held = collapse(rows)
|
||||||
|
removed = len(rows) - len(out)
|
||||||
|
grand_g += len(groups); grand_r += removed; grand_h += len(held)
|
||||||
|
print(f'{name:<8}{len(rows):>6}{len(groups):>7}{removed:>6}{len(out):>7}{len(held):>7}')
|
||||||
|
for h in held:
|
||||||
|
attach_detail.append((name, h['E'], h['F'], h['O']))
|
||||||
|
if mode == 'run' and groups:
|
||||||
|
bak = xp.replace('.xlsx', '_backup_collapse전.xlsx')
|
||||||
|
shutil.copy(xp, bak)
|
||||||
|
dd.write_back(ws, out)
|
||||||
|
try:
|
||||||
|
wb.save(xp)
|
||||||
|
except PermissionError:
|
||||||
|
wb.save(xp.replace('.xlsx', '_LP.xlsx')); print(f' !! {name} 잠김→_LP')
|
||||||
|
print('-' * 55)
|
||||||
|
print(f'합계: 합친그룹 {grand_g} / 제거행 {grand_r} / 부착합침 {grand_h} ({mode})')
|
||||||
|
if attach_detail:
|
||||||
|
print(f'\n=== 부착 유형 합쳐진 그룹 {len(attach_detail)}개 (O 결과 확인용) ===')
|
||||||
|
for nm, E, F, O in attach_detail:
|
||||||
|
print(f' [{nm}] {E} > {F} → O={O}')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
126
_스크립트/_collapse_landing.py
Normal file
126
_스크립트/_collapse_landing.py
Normal file
@ -0,0 +1,126 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""중분류/소분류 '랜딩행' 제거 — 매뉴얼 1-5 확장(③ K동일 조건 제거).
|
||||||
|
|
||||||
|
_dedup_all.py 는 부모행 URL == 첫 자식 URL 일 때만 부모행을 삭제한다(③).
|
||||||
|
eGov '누리집지도'(고창·임실·정읍·진안 등)는 중분류 랜딩(예 …002000000 '민원안내')과
|
||||||
|
첫 소분류(…002007000)의 URL이 달라 ③에 안 걸리고 랜딩행이 남아 있다.
|
||||||
|
사용자 요청(2026-05-31): ③조건을 빼고, 자식을 거느린 부모 랜딩행을 전부 삭제.
|
||||||
|
· 부모행의 가장 깊은 카테고리 컬럼 lc, lc+1은 빈칸(부모는 그 깊이의 잎이 아님)
|
||||||
|
· 직후 자식행이 D~lc 값 동일 & lc+1 채워짐 → 부모행 삭제(자식이 병합으로 라벨 승계)
|
||||||
|
URL 무관. 카테고리 텍스트(D/E/F)는 재병합으로 보존.
|
||||||
|
|
||||||
|
평탄화/재병합/순번/하이퍼링크는 _dedup_all 재사용. 전 컬럼(B~AA, L~Q) 보존.
|
||||||
|
|
||||||
|
사용:
|
||||||
|
python -X utf8 _collapse_landing.py dry <기관명...|경로>
|
||||||
|
python -X utf8 _collapse_landing.py run <기관명...|경로>
|
||||||
|
--maxlc N : lc<=N 깊이까지만 삭제(기본 9=I, 즉 D~I 랜딩 모두). E만 원하면 --maxlc 5.
|
||||||
|
"""
|
||||||
|
import os, sys, shutil, importlib.util, warnings
|
||||||
|
import openpyxl
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
spec = importlib.util.spec_from_file_location('dd', __file__.replace('_collapse_landing.py', '_dedup_all.py'))
|
||||||
|
dd = importlib.util.module_from_spec(spec); spec.loader.exec_module(dd)
|
||||||
|
|
||||||
|
CAT_COLS = dd.CAT_COLS # D E F G H I J = 4..10
|
||||||
|
EXCLUDE = {'계룡시'} # 검수완료
|
||||||
|
|
||||||
|
|
||||||
|
def is_safe_shell(v):
|
||||||
|
"""순수 메뉴 셸인가: L=페이지(또는 빈칸) & M<=1 & 공공누리 미부착."""
|
||||||
|
L = v.get(12); M = v.get(13); O = v.get(15)
|
||||||
|
if isinstance(O, str) and '유형' in O:
|
||||||
|
return False
|
||||||
|
if L not in (None, '', '페이지'):
|
||||||
|
return False
|
||||||
|
if isinstance(M, (int, float)) and M and M > 1:
|
||||||
|
return False
|
||||||
|
return True
|
||||||
|
|
||||||
|
|
||||||
|
def collapse(rows, maxlc=9, safe=False):
|
||||||
|
"""랜딩행 제거. safe=True면 순수 셸(데이터 무보유)만. 반환 (남은행, 제거목록)."""
|
||||||
|
out, removed = [], []
|
||||||
|
i = 0
|
||||||
|
while i < len(rows):
|
||||||
|
v = rows[i]['vals']
|
||||||
|
if i + 1 < len(rows):
|
||||||
|
nv = rows[i + 1]['vals']
|
||||||
|
lc = dd.leaf_depth(v) # 부모의 가장 깊은 카테고리 컬럼
|
||||||
|
if 4 <= lc <= maxlc:
|
||||||
|
same_upper = all((v.get(c) or '') == (nv.get(c) or '')
|
||||||
|
for c in CAT_COLS if c <= lc)
|
||||||
|
child_next = nv.get(lc + 1) not in (None, '')
|
||||||
|
parent_is_landing = v.get(lc + 1) in (None, '')
|
||||||
|
if safe and not is_safe_shell(v):
|
||||||
|
same_upper = False # 데이터 보유 랜딩은 보존
|
||||||
|
if same_upper and child_next and parent_is_landing:
|
||||||
|
removed.append({'src': rows[i]['src'], 'lc': lc,
|
||||||
|
'label': v.get(lc), 'child': nv.get(lc + 1),
|
||||||
|
'url': v.get(11), 'curl': nv.get(11)})
|
||||||
|
i += 1
|
||||||
|
continue
|
||||||
|
out.append(rows[i]); i += 1
|
||||||
|
return out, removed
|
||||||
|
|
||||||
|
|
||||||
|
LV = {4: 'D', 5: 'E', 6: 'F', 7: 'G', 8: 'H', 9: 'I', 10: 'J'}
|
||||||
|
|
||||||
|
|
||||||
|
def resolve(args):
|
||||||
|
paths = []
|
||||||
|
for a in args:
|
||||||
|
if a.lower().endswith('.xlsx') or os.path.sep in a:
|
||||||
|
paths.append((os.path.splitext(os.path.basename(a))[0], a))
|
||||||
|
names = [a for a in args if not (a.lower().endswith('.xlsx') or os.path.sep in a)]
|
||||||
|
for prov, idx, name in dd.SITES:
|
||||||
|
if names and name in names:
|
||||||
|
paths.append((name, dd.xpath(prov, idx, name)))
|
||||||
|
return paths
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
mode = sys.argv[1] if len(sys.argv) > 1 else 'dry'
|
||||||
|
rest = [a for a in sys.argv[2:] if not a.startswith('--')]
|
||||||
|
maxlc = 9
|
||||||
|
if '--maxlc' in sys.argv:
|
||||||
|
maxlc = int(sys.argv[sys.argv.index('--maxlc') + 1])
|
||||||
|
safe = '--safe' in sys.argv
|
||||||
|
targets = resolve(rest)
|
||||||
|
if not targets:
|
||||||
|
print('대상 없음.'); return
|
||||||
|
grand = 0
|
||||||
|
for name, xp in targets:
|
||||||
|
if name in EXCLUDE:
|
||||||
|
print(f'{name}: 검수완료 제외'); continue
|
||||||
|
if not os.path.exists(xp):
|
||||||
|
print(f'{name}: 엑셀 없음'); continue
|
||||||
|
wb = openpyxl.load_workbook(xp); ws = wb.active
|
||||||
|
rows = dd.load_flat(ws)
|
||||||
|
out, removed = collapse(rows, maxlc, safe)
|
||||||
|
grand += len(removed)
|
||||||
|
bylv = {}
|
||||||
|
for r in removed:
|
||||||
|
bylv[r['lc']] = bylv.get(r['lc'], 0) + 1
|
||||||
|
lvstr = ' '.join(f'{LV[k]}:{v}' for k, v in sorted(bylv.items()))
|
||||||
|
print(f'\n=== {name} === {len(rows)}→{len(out)} (제거 {len(removed)}) [{lvstr}]')
|
||||||
|
for r in removed[:6]:
|
||||||
|
print(f" [{LV[r['lc']]}] {r['label']} ({r['url']}) → 자식 첫행 {r['child']} ({r['curl']})")
|
||||||
|
if len(removed) > 6:
|
||||||
|
print(f' ... 외 {len(removed)-6}건')
|
||||||
|
if mode == 'run' and removed:
|
||||||
|
bak = xp.replace('.xlsx', '_backup_랜딩제거전.xlsx')
|
||||||
|
shutil.copy(xp, bak)
|
||||||
|
dd.write_back(ws, out)
|
||||||
|
try:
|
||||||
|
wb.save(xp)
|
||||||
|
except PermissionError:
|
||||||
|
wb.save(xp.replace('.xlsx', '_LP.xlsx')); print(' !! 잠김→_LP')
|
||||||
|
else:
|
||||||
|
print(f' 저장 완료. 백업: {os.path.basename(bak)}')
|
||||||
|
print(f'\n총 제거 대상: {grand}행 ({mode}, maxlc={maxlc})')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
166
_스크립트/_collapse_recursive.py
Normal file
166
_스크립트/_collapse_recursive.py
Normal file
@ -0,0 +1,166 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""규칙 B 재귀판: 탭 레벨(G→H→I…)에서 '자식이 전부 페이지'면 1행으로 합치고 M=합.
|
||||||
|
메뉴 카테고리(D·E·F)는 보존 — 합치기는 lc>=7(G 이하 탭 레벨)에서만 수행.
|
||||||
|
|
||||||
|
각 레벨 lc 에서:
|
||||||
|
부모(D..lc-1) 동일 + lc 비어있지 않음 으로 형제 묶음 → 그 묶음이
|
||||||
|
· 전부 L=페이지, · lc보다 깊은 카테고리열 모두 비어있음(잎)
|
||||||
|
이면 1행으로 합침: lc 이하 카테고리열 비움, M=합, K=첫 행 URL,
|
||||||
|
공공누리 O=유형 합집합(오름차순,콤마,공백없음)/P=게시판우선/Q=하나라도 Y.
|
||||||
|
깊은→얕은 순으로 반복(안정될 때까지) → G→H 합친 뒤 F→G까지 자연 연쇄.
|
||||||
|
|
||||||
|
평탄화/재병합(D/E/F)/순번/하이퍼링크는 _dedup_all.write_back 재사용.
|
||||||
|
|
||||||
|
사용:
|
||||||
|
python -X utf8 _collapse_recursive.py dry <엑셀경로 | 기관명...>
|
||||||
|
python -X utf8 _collapse_recursive.py run <엑셀경로 | 기관명...>
|
||||||
|
(기관명 없이 경로 1개만 줘도 됨)
|
||||||
|
"""
|
||||||
|
import os, re, sys, shutil, importlib.util, warnings
|
||||||
|
import openpyxl
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
spec = importlib.util.spec_from_file_location('dd', __file__.replace('_collapse_recursive.py', '_dedup_all.py'))
|
||||||
|
dd = importlib.util.module_from_spec(spec); spec.loader.exec_module(dd)
|
||||||
|
|
||||||
|
CAT = [4, 5, 6, 7, 8, 9, 10] # D E F G H I J
|
||||||
|
COLLAPSE_LEVELS = [10, 9, 8, 7] # 탭 레벨만(깊은→얕은). E(6)·D(5)는 제외=카테고리 보존
|
||||||
|
TYPE_PAT = re.compile(r'(\d)\s*유형')
|
||||||
|
EXCLUDE = {'계룡시'} # 검수완료
|
||||||
|
|
||||||
|
|
||||||
|
def mval(x):
|
||||||
|
try:
|
||||||
|
return int(x)
|
||||||
|
except Exception:
|
||||||
|
return 1
|
||||||
|
|
||||||
|
|
||||||
|
def merge_kogl(grp):
|
||||||
|
types = set()
|
||||||
|
for g in grp:
|
||||||
|
o = g['vals'].get(15)
|
||||||
|
if isinstance(o, str):
|
||||||
|
for m in TYPE_PAT.finditer(o):
|
||||||
|
types.add(int(m.group(1)))
|
||||||
|
if types:
|
||||||
|
O = ','.join(f'{n}유형' for n in sorted(types))
|
||||||
|
else:
|
||||||
|
O = '미부착'
|
||||||
|
Ps = [g['vals'].get(16) for g in grp]
|
||||||
|
P = '게시판' if '게시판' in Ps else ('게시물' if '게시물' in Ps else None)
|
||||||
|
Qs = [g['vals'].get(17) for g in grp]
|
||||||
|
Q = 'Y' if 'Y' in Qs else ('N' if 'N' in Qs else None)
|
||||||
|
return O, (P if O != '미부착' else None), (Q if O != '미부착' else None)
|
||||||
|
|
||||||
|
|
||||||
|
def collapse_level(rows, lc):
|
||||||
|
parent = [c for c in CAT if c < lc]
|
||||||
|
deeper = [c for c in CAT if c > lc]
|
||||||
|
out, groups = [], []
|
||||||
|
i, n = 0, len(rows)
|
||||||
|
while i < n:
|
||||||
|
v = rows[i]['vals']
|
||||||
|
if v.get(lc) not in (None, ''):
|
||||||
|
j = i
|
||||||
|
while (j < n and all(rows[j]['vals'].get(c) == v.get(c) for c in parent)
|
||||||
|
and rows[j]['vals'].get(lc) not in (None, '')):
|
||||||
|
j += 1
|
||||||
|
grp = rows[i:j]
|
||||||
|
all_page = all(g['vals'].get(12) == '페이지' for g in grp)
|
||||||
|
leaf = all(all(g['vals'].get(c) in (None, '') for c in deeper) for g in grp)
|
||||||
|
has_attach = any(g['vals'].get(15) not in (None, '', '미부착') for g in grp)
|
||||||
|
if len(grp) >= 2 and all_page and leaf:
|
||||||
|
msum = sum(mval(g['vals'].get(13)) for g in grp)
|
||||||
|
O, P, Q = merge_kogl(grp)
|
||||||
|
first = dict(grp[0]); nv = dict(first['vals'])
|
||||||
|
for c in [lc] + deeper:
|
||||||
|
nv[c] = None
|
||||||
|
nv[13] = msum; nv[15] = O; nv[16] = P; nv[17] = Q
|
||||||
|
first['vals'] = nv
|
||||||
|
out.append(first)
|
||||||
|
groups.append({'lc': lc, 'path': [v.get(c) for c in parent],
|
||||||
|
'label': v.get(lc), 'n': len(grp),
|
||||||
|
'ms': [mval(g['vals'].get(13)) for g in grp],
|
||||||
|
'sum': msum, 'O': O, 'attach': has_attach,
|
||||||
|
'names': [g['vals'].get(lc) for g in grp]})
|
||||||
|
i = j; continue
|
||||||
|
out.extend(grp); i = j; continue
|
||||||
|
out.append(rows[i]); i += 1
|
||||||
|
return out, groups
|
||||||
|
|
||||||
|
|
||||||
|
def collapse_all_levels(rows):
|
||||||
|
all_groups = []
|
||||||
|
changed = True
|
||||||
|
while changed:
|
||||||
|
changed = False
|
||||||
|
for lc in COLLAPSE_LEVELS:
|
||||||
|
rows, groups = collapse_level(rows, lc)
|
||||||
|
if groups:
|
||||||
|
changed = True
|
||||||
|
all_groups.extend(groups)
|
||||||
|
return rows, all_groups
|
||||||
|
|
||||||
|
|
||||||
|
def resolve_paths(args):
|
||||||
|
"""경로 또는 기관명 목록 → [(name, xlsx경로)]"""
|
||||||
|
paths = []
|
||||||
|
for a in args:
|
||||||
|
if a.lower().endswith('.xlsx') or os.path.sep in a:
|
||||||
|
paths.append((os.path.splitext(os.path.basename(a))[0], a))
|
||||||
|
names = [a for a in args if not (a.lower().endswith('.xlsx') or os.path.sep in a)]
|
||||||
|
for prov, idx, name in dd.SITES:
|
||||||
|
if names and name not in names:
|
||||||
|
continue
|
||||||
|
if not names:
|
||||||
|
continue
|
||||||
|
paths.append((name, dd.xpath(prov, idx, name)))
|
||||||
|
return paths
|
||||||
|
|
||||||
|
|
||||||
|
LV = {7: 'F→G', 8: 'G→H', 9: 'H→I', 10: 'I→J'}
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
mode = sys.argv[1] if len(sys.argv) > 1 else 'dry'
|
||||||
|
rest = [a for a in sys.argv[2:] if not a.startswith('--')]
|
||||||
|
targets = resolve_paths(rest)
|
||||||
|
if not targets:
|
||||||
|
print('대상 없음. 경로 또는 기관명을 지정하세요.'); return
|
||||||
|
for name, xp in targets:
|
||||||
|
if name in EXCLUDE:
|
||||||
|
print(f'{name}: 검수완료 제외'); continue
|
||||||
|
if not os.path.exists(xp):
|
||||||
|
print(f'{name}: 엑셀 없음 ({xp})'); continue
|
||||||
|
wb = openpyxl.load_workbook(xp); ws = wb.active
|
||||||
|
rows = dd.load_flat(ws)
|
||||||
|
out, groups = collapse_all_levels(rows)
|
||||||
|
removed = len(rows) - len(out)
|
||||||
|
bylv = {}
|
||||||
|
for g in groups:
|
||||||
|
bylv.setdefault(g['lc'], 0)
|
||||||
|
bylv[g['lc']] += 1
|
||||||
|
lvstr = ' '.join(f"{LV.get(k,k)}:{v}" for k, v in sorted(bylv.items()))
|
||||||
|
print(f'\n=== {name} === {len(rows)}→{len(out)} (제거 {removed}) [{lvstr}]')
|
||||||
|
for g in groups:
|
||||||
|
ms = '+'.join(str(m) for m in g['ms'])
|
||||||
|
star = ' ★부착' if g['attach'] else ''
|
||||||
|
print(f" [{LV.get(g['lc'],g['lc'])}] {' > '.join(str(x) for x in g['path'] if x)} > {g['label']}"
|
||||||
|
f" ({g['n']}개) → M={ms}={g['sum']} O={g['O']}{star}")
|
||||||
|
if mode == 'run' and groups:
|
||||||
|
bak = xp.replace('.xlsx', '_backup_재귀합치기전.xlsx')
|
||||||
|
shutil.copy(xp, bak)
|
||||||
|
dd.write_back(ws, out)
|
||||||
|
try:
|
||||||
|
wb.save(xp)
|
||||||
|
except PermissionError:
|
||||||
|
wb.save(xp.replace('.xlsx', '_LP.xlsx')); print(f' !! {name} 잠김→_LP')
|
||||||
|
else:
|
||||||
|
print(f' 저장 완료. 백업: {os.path.basename(bak)}')
|
||||||
|
elif mode != 'run':
|
||||||
|
print(' (DRY)')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
160
_스크립트/_collapse_run.log
Normal file
160
_스크립트/_collapse_run.log
Normal file
@ -0,0 +1,160 @@
|
|||||||
|
대상 기관: 40개 (공주·계룡 제외) 모드=run
|
||||||
|
|
||||||
|
기관 현재 합친그룹 제거행 →남음 부착합침
|
||||||
|
-------------------------------------------------------
|
||||||
|
금산군 533 54 167 366 0
|
||||||
|
논산시 758 75 155 603 0
|
||||||
|
당진시 309 0 0 309 0
|
||||||
|
보령시 547 0 0 547 0
|
||||||
|
부여군 222 0 0 222 0
|
||||||
|
서산시 306 0 0 306 0
|
||||||
|
서천군 275 0 0 275 0
|
||||||
|
아산시 261 0 0 261 0
|
||||||
|
예산군 566 51 95 471 0
|
||||||
|
천안시 640 55 152 488 0
|
||||||
|
청양군 257 0 0 257 0
|
||||||
|
태안군 154 0 0 154 0
|
||||||
|
홍성군 524 47 138 386 0
|
||||||
|
괴산군 240 2 6 234 0
|
||||||
|
단양군 428 20 109 319 10
|
||||||
|
보은군 512 26 102 410 0
|
||||||
|
영동군 764 25 66 698 0
|
||||||
|
옥천군 521 37 165 356 0
|
||||||
|
음성군 632 45 244 388 0
|
||||||
|
제천시 594 41 224 370 1
|
||||||
|
증평군 404 0 0 404 0
|
||||||
|
진천군 297 0 0 297 0
|
||||||
|
청주시 406 0 0 406 0
|
||||||
|
충주시 376 0 0 376 0
|
||||||
|
고창군 379 19 53 326 0
|
||||||
|
군산시 637 9 35 602 9
|
||||||
|
김제시 452 0 0 452 0
|
||||||
|
남원시 331 38 90 241 38
|
||||||
|
무주군 83 0 0 83 0
|
||||||
|
부안군 201 0 0 201 0
|
||||||
|
순창군 452 28 152 300 28
|
||||||
|
완주군 328 0 0 328 0
|
||||||
|
익산시 535 26 131 404 26
|
||||||
|
임실군 313 16 127 186 0
|
||||||
|
장수군 78 0 0 78 0
|
||||||
|
전주시 263 0 0 263 0
|
||||||
|
정읍시 245 0 0 245 0
|
||||||
|
진안군 307 0 0 307 0
|
||||||
|
서귀포시 388 1 1 387 0
|
||||||
|
제주시 310 0 0 310 0
|
||||||
|
-------------------------------------------------------
|
||||||
|
합계: 합친그룹 615 / 제거행 2212 / 부착합침 112 (run)
|
||||||
|
|
||||||
|
=== 부착 유형 합쳐진 그룹 112개 (O 결과 확인용) ===
|
||||||
|
[단양군] 상징 > 브랜드 → O=2유형
|
||||||
|
[단양군] 군정안내 > 청사배치도 → O=2유형
|
||||||
|
[단양군] 아동복지 > 아동복지정책 → O=2유형
|
||||||
|
[단양군] 청소년복지 > 청소년복지시설 → O=2유형
|
||||||
|
[단양군] 여성·가족 복지 > 여성복지시설 → O=2유형
|
||||||
|
[단양군] 여성·가족 복지 > 여성·가족복지정책 → O=2유형
|
||||||
|
[단양군] 노인복지 > 노인복지시설 → O=2유형
|
||||||
|
[단양군] 노인복지 > 노인복지정책 → O=2유형
|
||||||
|
[단양군] 장애인복지 > 장애인복지시설 → O=2유형
|
||||||
|
[단양군] 장애인복지 > 장애인복지정책 → O=2유형
|
||||||
|
[제천시] 제천시소개 > 공공저작물 → O=1유형,2유형,3유형
|
||||||
|
[군산시] 산업인프라 > 항만/여객/공항/철도/컨벤션 → O=4유형
|
||||||
|
[군산시] 농업/축산업 > 농산물 유통 → O=4유형
|
||||||
|
[군산시] 건설 > 자전거 → O=4유형
|
||||||
|
[군산시] 에너지 > 태양광 → O=4유형
|
||||||
|
[군산시] 에너지 > 가스/석유 → O=4유형
|
||||||
|
[군산시] 군산시 소개 > 행정구역/행정지도 → O=4유형
|
||||||
|
[군산시] 군산시 소개 > 자매결연/국제협력 도시 → O=4유형
|
||||||
|
[군산시] 군산시 소개 > 군산의 상징 → O=4유형
|
||||||
|
[군산시] 시청안내 > 전화번호안내 → O=4유형
|
||||||
|
[남원시] 행복민원실 > 민원발급안내 → O=4유형
|
||||||
|
[남원시] 부동산정보 > 지적재조사 → O=4유형
|
||||||
|
[남원시] 행정정보공개 > 정보공개제도안내 → O=4유형
|
||||||
|
[남원시] 예산공개 > 2026년도 → O=4유형
|
||||||
|
[남원시] 예산공개 > 2025년도 → O=4유형
|
||||||
|
[남원시] 예산공개 > 2024년도 → O=4유형
|
||||||
|
[남원시] 예산공개 > 2023년도 → O=4유형
|
||||||
|
[남원시] 예산공개 > 2022년도 → O=4유형
|
||||||
|
[남원시] 예산공개 > 2021년도 → O=4유형
|
||||||
|
[남원시] 예산공개 > 2020년도 → O=4유형
|
||||||
|
[남원시] 예산공개 > 2019년도 → O=4유형
|
||||||
|
[남원시] 예산공개 > 2018년도 → O=4유형
|
||||||
|
[남원시] 예산공개 > 2017년도 → O=4유형
|
||||||
|
[남원시] 예산공개 > 2016년도 → O=4유형
|
||||||
|
[남원시] 예산공개 > 2015년도 → O=4유형
|
||||||
|
[남원시] 예산공개 > 2014년도 → O=4유형
|
||||||
|
[남원시] 예산공개 > 2013년도 → O=4유형
|
||||||
|
[남원시] 예산공개 > 2012년도 → O=4유형
|
||||||
|
[남원시] 예산공개 > 2011년도 → O=4유형
|
||||||
|
[남원시] 예산공개 > 2009년도 → O=4유형
|
||||||
|
[남원시] 재정공시 > 2025년 → O=4유형
|
||||||
|
[남원시] 재정공시 > 2024년 → O=4유형
|
||||||
|
[남원시] 재정공시 > 2023년 → O=4유형
|
||||||
|
[남원시] 재정공시 > 2022년 → O=4유형
|
||||||
|
[남원시] 재정공시 > 2021년 → O=4유형
|
||||||
|
[남원시] 재정공시 > 2020년 → O=4유형
|
||||||
|
[남원시] 재정공시 > 2019년 → O=4유형
|
||||||
|
[남원시] 재정공시 > 2018년 → O=4유형
|
||||||
|
[남원시] 재정공시 > 2017년 → O=4유형
|
||||||
|
[남원시] 재정공시 > 2016년 → O=4유형
|
||||||
|
[남원시] 재정공시 > 2015년 → O=4유형
|
||||||
|
[남원시] 열린행정 > 행정서비스헌장 → O=4유형
|
||||||
|
[남원시] 남원의역사 > 시대별 → O=4유형
|
||||||
|
[남원시] 남원의상징 > 기관상징 → O=4유형
|
||||||
|
[남원시] 자매 우호 결연 > 자매결연 → O=4유형
|
||||||
|
[남원시] 자매 우호 결연 > 우호결연 → O=4유형
|
||||||
|
[남원시] 시청안내 > 청사(시설물)안내 → O=4유형
|
||||||
|
[남원시] 시청안내 > 찾아오시는길 → O=4유형
|
||||||
|
[순창군] 주요민원안내 > 자동차등록안내 → O=4유형
|
||||||
|
[순창군] 인허가제도 > 축사시설 허가 → O=4유형
|
||||||
|
[순창군] 인허가제도 > 정보통신설비 허가 → O=4유형
|
||||||
|
[순창군] 여성/가족 > 여성친화도시 → O=4유형
|
||||||
|
[순창군] 여성/가족 > 아동/청소년 → O=4유형
|
||||||
|
[순창군] 행복순창! 인구정책 길라잡이 > 임신/출산 → O=4유형
|
||||||
|
[순창군] 행복순창! 인구정책 길라잡이 > 영유아 → O=4유형
|
||||||
|
[순창군] 행복순창! 인구정책 길라잡이 > 아동/청소년 → O=4유형
|
||||||
|
[순창군] 행복순창! 인구정책 길라잡이 > 결혼 → O=4유형
|
||||||
|
[순창군] 행복순창! 인구정책 길라잡이 > 청년/중장년 → O=4유형
|
||||||
|
[순창군] 행복순창! 인구정책 길라잡이 > 노년 → O=4유형
|
||||||
|
[순창군] 행복순창! 인구정책 길라잡이 > 다문화 → O=4유형
|
||||||
|
[순창군] 행복순창! 인구정책 길라잡이 > 공통사항 → O=4유형
|
||||||
|
[순창군] 군민생활 > 복지·주거 → O=4유형
|
||||||
|
[순창군] 군민생활 > 식품·공중위생 → O=4유형
|
||||||
|
[순창군] 군민생활 > 군민안전보험 → O=4유형
|
||||||
|
[순창군] 군민생활 > 교육·인재양성 → O=4유형
|
||||||
|
[순창군] 군민생활 > 일자리·고용 → O=4유형
|
||||||
|
[순창군] 경제·산업 > 농공단지 → O=4유형
|
||||||
|
[순창군] 상하수도·환경 > 상수도 → O=4유형
|
||||||
|
[순창군] 상하수도·환경 > 생활환경정보 → O=4유형
|
||||||
|
[순창군] 농촌개발·축산·산림 > 일반농산어촌개발사업 → O=4유형
|
||||||
|
[순창군] 농촌개발·축산·산림 > 농촌체험마을 → O=4유형
|
||||||
|
[순창군] 농촌개발·축산·산림 > 축산·산림 → O=4유형
|
||||||
|
[순창군] 재난·안전·민방위 > 민방위 → O=4유형
|
||||||
|
[순창군] 순창 장류산업 지역특구 > 주요사업 → O=4유형
|
||||||
|
[순창군] 순창 장류산업 지역특구 > 주요시설 → O=4유형
|
||||||
|
[순창군] 재정정보 > 공유재산공개 → O=4유형
|
||||||
|
[익산시] 기부美 > 명예의전당 → O=4유형
|
||||||
|
[익산시] 교통 > 시내버스 → O=4유형
|
||||||
|
[익산시] 교통 > 주정차 → O=4유형
|
||||||
|
[익산시] 복지 > 아동/청소년 → O=4유형
|
||||||
|
[익산시] 복지 > 국민생활복지 → O=4유형
|
||||||
|
[익산시] 복지 > 의료급여제도 → O=4유형
|
||||||
|
[익산시] 위생 > 식품위생 → O=4유형
|
||||||
|
[익산시] 환경/보건 > 정신재활시설 → O=4유형
|
||||||
|
[익산시] 상하수도사업단 > 사업단소개 → O=4유형
|
||||||
|
[익산시] 상하수도사업단 > 하수도 → O=4유형
|
||||||
|
[익산시] 상하수도사업단 > 주요시책사업 → O=4유형
|
||||||
|
[익산시] 산업/경제 > 소상공인 정책 → O=4유형
|
||||||
|
[익산시] 전입혜택 > 학생지원 → O=4유형
|
||||||
|
[익산시] 전입혜택 > 일반시민 → O=4유형
|
||||||
|
[익산시] 전입혜택 > 여성보육 → O=4유형
|
||||||
|
[익산시] 전입혜택 > 전입청년 → O=4유형
|
||||||
|
[익산시] 재난안전 > 대피장소 현황 → O=4유형
|
||||||
|
[익산시] 익산의 상징 > 익산브랜드 → O=4유형
|
||||||
|
[익산시] 익산의 역사 > 역사와 유래 → O=4유형
|
||||||
|
[익산시] 익산의 역사 > 시대별 익산의 역사 → O=4유형
|
||||||
|
[익산시] 익산의통계 > 도표로 보는 통계 → O=4유형
|
||||||
|
[익산시] 시청안내 > 조직도 → O=4유형
|
||||||
|
[익산시] 시청안내 > 부서소개 → O=4유형
|
||||||
|
[익산시] 시청안내 > 시청사소개 → O=4유형
|
||||||
|
[익산시] 시청안내 > 전화번호 → O=4유형
|
||||||
|
[익산시] 상호결연·우호도시 > 국제도시 → O=4유형
|
||||||
215
_스크립트/_dedup_all.py
Normal file
215
_스크립트/_dedup_all.py
Normal file
@ -0,0 +1,215 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""부모-자식 URL 중복 제거 — 전 깊이 일반화 (매뉴얼 1-5 확장).
|
||||||
|
|
||||||
|
기존 phase1 dedup은 F→G 한 깊이만 처리. 서산시처럼 중분류(E)가 클릭 랜딩이고
|
||||||
|
그 URL이 첫 소분류(F)와 같으면(E→F 중복) 안 지워졌다. 본 도구는 깊이 무관:
|
||||||
|
부모행의 가장 깊은 카테고리 컬럼이 lc 이고, 직후 자식행이
|
||||||
|
· D~lc 까지 값이 모두 동일, · lc+1 컬럼이 채워짐, · K(URL)가 부모와 '정확히' 동일
|
||||||
|
이면 부모행을 삭제(자식이 카테고리 텍스트를 병합으로 승계). URL이 다르면 보존.
|
||||||
|
|
||||||
|
평탄화→재병합(D/E/F)→순번/행높이/하이퍼링크 재설정은 _tab_expand 와 동일 로직.
|
||||||
|
모든 컬럼(B~AA, L~Q 등 Phase2~4 데이터 포함) 보존.
|
||||||
|
|
||||||
|
사용:
|
||||||
|
python -X utf8 _dedup_all.py dry [기관...] # 제거대상 집계만(읽기전용)
|
||||||
|
python -X utf8 _dedup_all.py run [기관...] # 백업(*_backup_dedup전.xlsx) 후 제거
|
||||||
|
"""
|
||||||
|
import os
|
||||||
|
import sys
|
||||||
|
import shutil
|
||||||
|
import warnings
|
||||||
|
from copy import copy
|
||||||
|
|
||||||
|
import openpyxl
|
||||||
|
from openpyxl.styles import Alignment, Font
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
MAXCOL = 27
|
||||||
|
CAT_COLS = [4, 5, 6, 7, 8, 9, 10] # D E F G H I J
|
||||||
|
HEADER_MERGES = {'B1:R1', 'S1:W1', 'Y1:AA1'}
|
||||||
|
|
||||||
|
HERE = os.path.dirname(os.path.abspath(__file__))
|
||||||
|
MAP = os.path.join(os.path.dirname(HERE), '작업파일', '광역_사이트맵')
|
||||||
|
|
||||||
|
# 전 42개 시·군 (광역, idx, 기관)
|
||||||
|
SITES = (
|
||||||
|
[('충청남도', i, n) for i, n in [
|
||||||
|
(1, '계룡시'), (2, '공주시'), (3, '금산군'), (4, '논산시'), (5, '당진시'),
|
||||||
|
(6, '보령시'), (7, '부여군'), (8, '서산시'), (9, '서천군'), (10, '아산시'),
|
||||||
|
(11, '예산군'), (12, '천안시'), (13, '청양군'), (14, '태안군'), (15, '홍성군')]]
|
||||||
|
+ [('충청북도', i, n) for i, n in [
|
||||||
|
(1, '괴산군'), (2, '단양군'), (3, '보은군'), (4, '영동군'), (5, '옥천군'),
|
||||||
|
(6, '음성군'), (7, '제천시'), (8, '증평군'), (9, '진천군'), (10, '청주시'), (11, '충주시')]]
|
||||||
|
+ [('전북특별자치도', i, n) for i, n in [
|
||||||
|
(1, '고창군'), (2, '군산시'), (3, '김제시'), (4, '남원시'), (5, '무주군'),
|
||||||
|
(6, '부안군'), (7, '순창군'), (8, '완주군'), (9, '익산시'), (10, '임실군'),
|
||||||
|
(11, '장수군'), (12, '전주시'), (13, '정읍시'), (14, '진안군')]]
|
||||||
|
+ [('제주특별자치도', i, n) for i, n in [(1, '서귀포시'), (2, '제주시')]]
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def xpath(prov, idx, name):
|
||||||
|
return os.path.join(MAP, prov, f'{idx}.{name}', f'{prov}_{name}.xlsx')
|
||||||
|
|
||||||
|
|
||||||
|
def load_flat(ws):
|
||||||
|
for mr in list(ws.merged_cells.ranges):
|
||||||
|
s = str(mr)
|
||||||
|
if s in HEADER_MERGES:
|
||||||
|
continue
|
||||||
|
top = ws.cell(mr.min_row, mr.min_col).value
|
||||||
|
ws.unmerge_cells(s)
|
||||||
|
for rr in range(mr.min_row, mr.max_row + 1):
|
||||||
|
for cc in range(mr.min_col, mr.max_col + 1):
|
||||||
|
if ws.cell(rr, cc).value in (None, ''):
|
||||||
|
ws.cell(rr, cc).value = top
|
||||||
|
rows = []
|
||||||
|
for r in range(3, ws.max_row + 1):
|
||||||
|
if all(ws.cell(r, c).value in (None, '') for c in range(2, MAXCOL + 1)):
|
||||||
|
continue
|
||||||
|
vals = {c: ws.cell(r, c).value for c in range(1, MAXCOL + 1)}
|
||||||
|
styles = {}
|
||||||
|
for c in range(1, MAXCOL + 1):
|
||||||
|
sc = ws.cell(r, c)
|
||||||
|
styles[c] = (copy(sc.font), copy(sc.fill), copy(sc.border),
|
||||||
|
copy(sc.alignment), sc.number_format, copy(sc.protection))
|
||||||
|
hl = ws.cell(r, 11).hyperlink
|
||||||
|
rows.append({'src': r, 'vals': vals, 'styles': styles,
|
||||||
|
'hyperlink': hl.target if hl else None})
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def leaf_depth(vals):
|
||||||
|
deep = 4
|
||||||
|
for c in CAT_COLS:
|
||||||
|
if vals.get(c) not in (None, ''):
|
||||||
|
deep = c
|
||||||
|
return deep
|
||||||
|
|
||||||
|
|
||||||
|
def dedup(rows):
|
||||||
|
"""전 깊이 부모-자식 URL 중복 제거. 반환: (남은행, 제거목록)."""
|
||||||
|
out, removed = [], []
|
||||||
|
i = 0
|
||||||
|
while i < len(rows):
|
||||||
|
v = rows[i]['vals']
|
||||||
|
if i + 1 < len(rows):
|
||||||
|
nv = rows[i + 1]['vals']
|
||||||
|
lc = leaf_depth(v)
|
||||||
|
ku, kc = v.get(11), nv.get(11)
|
||||||
|
if lc < 10 and isinstance(ku, str) and ku.startswith('http') and ku == kc:
|
||||||
|
same_upper = all((v.get(c) or '') == (nv.get(c) or '')
|
||||||
|
for c in CAT_COLS if c <= lc)
|
||||||
|
child_next = nv.get(lc + 1) not in (None, '')
|
||||||
|
if same_upper and child_next:
|
||||||
|
removed.append({'src': rows[i]['src'], 'lc': lc,
|
||||||
|
'E': v.get(5), 'F_child': nv.get(lc + 1),
|
||||||
|
'url': ku})
|
||||||
|
i += 1
|
||||||
|
continue
|
||||||
|
out.append(rows[i])
|
||||||
|
i += 1
|
||||||
|
return out, removed
|
||||||
|
|
||||||
|
|
||||||
|
def write_back(ws, out_rows):
|
||||||
|
"""남은 행으로 데이터영역 재작성 + D/E/F 재병합 + 순번/행높이/하이퍼링크."""
|
||||||
|
for r in range(3, ws.max_row + 1):
|
||||||
|
for c in range(1, MAXCOL + 1):
|
||||||
|
ws.cell(r, c).value = None
|
||||||
|
ws.cell(r, c).hyperlink = None
|
||||||
|
START = 3
|
||||||
|
for idx, orow in enumerate(out_rows):
|
||||||
|
r = START + idx
|
||||||
|
sty = orow['styles']
|
||||||
|
for c in range(1, MAXCOL + 1):
|
||||||
|
cell = ws.cell(r, c)
|
||||||
|
f, fl, bd, al, nf, pr = sty[c]
|
||||||
|
cell.font = copy(f); cell.fill = copy(fl); cell.border = copy(bd)
|
||||||
|
cell.alignment = copy(al); cell.number_format = nf; cell.protection = copy(pr)
|
||||||
|
v = orow['vals']
|
||||||
|
for c in range(1, MAXCOL + 1):
|
||||||
|
ws.cell(r, c).value = v.get(c)
|
||||||
|
ws.cell(r, 2).value = idx + 1
|
||||||
|
END = START + len(out_rows) - 1
|
||||||
|
center = Alignment(horizontal='center', vertical='center', wrap_text=True)
|
||||||
|
left = Alignment(horizontal='left', vertical='center', wrap_text=False)
|
||||||
|
|
||||||
|
def merge_runs(letter, col, group_cols=()):
|
||||||
|
cur = ws.cell(START, col).value
|
||||||
|
grp = tuple(ws.cell(START, g).value for g in group_cols)
|
||||||
|
run = START
|
||||||
|
runs = []
|
||||||
|
for r in range(START + 1, END + 1):
|
||||||
|
val = ws.cell(r, col).value
|
||||||
|
g = tuple(ws.cell(r, gg).value for gg in group_cols)
|
||||||
|
if val == cur and g == grp:
|
||||||
|
continue
|
||||||
|
if cur not in (None, '') and r - 1 > run:
|
||||||
|
runs.append((run, r - 1))
|
||||||
|
cur, grp, run = val, g, r
|
||||||
|
if cur not in (None, '') and END > run:
|
||||||
|
runs.append((run, END))
|
||||||
|
for s, e in runs:
|
||||||
|
ws.merge_cells(f'{letter}{s}:{letter}{e}')
|
||||||
|
ws.cell(s, col).alignment = center
|
||||||
|
|
||||||
|
merge_runs('F', 6, (4, 5))
|
||||||
|
merge_runs('E', 5, (4,))
|
||||||
|
merge_runs('D', 4)
|
||||||
|
for r in range(START, END + 1):
|
||||||
|
for c in (4, 5, 6, 7):
|
||||||
|
if ws.cell(r, c).value is not None:
|
||||||
|
ws.cell(r, c).alignment = center
|
||||||
|
for r in range(1, END + 1):
|
||||||
|
ws.row_dimensions[r].height = 15
|
||||||
|
for r in range(START, END + 1):
|
||||||
|
cell = ws.cell(r, 11)
|
||||||
|
u = cell.value
|
||||||
|
if isinstance(u, str) and u.startswith(('http://', 'https://')):
|
||||||
|
cell.hyperlink = u
|
||||||
|
old = cell.font
|
||||||
|
cell.font = Font(name=old.name or '맑은 고딕', size=old.size or 11,
|
||||||
|
bold=old.bold, italic=old.italic,
|
||||||
|
color='0000FF', underline='single')
|
||||||
|
cell.alignment = left
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
mode = sys.argv[1] if len(sys.argv) > 1 else 'dry'
|
||||||
|
only = set(a for a in sys.argv[2:] if not a.startswith('--'))
|
||||||
|
sites = [s for s in SITES if not only or s[2] in only]
|
||||||
|
grand = 0
|
||||||
|
print(f'{"기관":<8}{"현재행":>6}{"제거":>5}{"→남음":>7} 유형(중분류 예시)')
|
||||||
|
print('-' * 80)
|
||||||
|
for prov, idx, name in sites:
|
||||||
|
xp = xpath(prov, idx, name)
|
||||||
|
if not os.path.exists(xp):
|
||||||
|
print(f'{name:<8} 엑셀 없음')
|
||||||
|
continue
|
||||||
|
wb = openpyxl.load_workbook(xp)
|
||||||
|
ws = wb.active
|
||||||
|
rows = load_flat(ws)
|
||||||
|
out, removed = dedup(rows)
|
||||||
|
grand += len(removed)
|
||||||
|
ex = ''
|
||||||
|
if removed:
|
||||||
|
sample = removed[0]
|
||||||
|
ex = f"E={sample['E']} = {sample['F_child']}"
|
||||||
|
print(f'{name:<8}{len(rows):>6}{len(removed):>5}{len(out):>7} {ex}')
|
||||||
|
if mode == 'run' and removed:
|
||||||
|
bak = xp.replace('.xlsx', '_backup_dedup전.xlsx')
|
||||||
|
shutil.copy(xp, bak)
|
||||||
|
write_back(ws, out)
|
||||||
|
try:
|
||||||
|
wb.save(xp)
|
||||||
|
except PermissionError:
|
||||||
|
wb.save(xp.replace('.xlsx', '_LP.xlsx'))
|
||||||
|
print(f' !! 원본 잠김 → _LP.xlsx 저장')
|
||||||
|
print('-' * 80)
|
||||||
|
print(f'총 제거 대상: {grand}행 ({mode})')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
23
_스크립트/_extract_chungbuk.py
Normal file
23
_스크립트/_extract_chungbuk.py
Normal file
@ -0,0 +1,23 @@
|
|||||||
|
"""충청북도 시·군 목록을 마스터 엑셀에서 추출."""
|
||||||
|
import openpyxl
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
ROOT = Path(r'D:\01.프로젝트\DB수집')
|
||||||
|
XLSX = ROOT / '붙임1_공공저작물 개방 대상기관(1160개) 실태조사 목록(신유형개방지원사업) 양식_260518.xlsx'
|
||||||
|
wb = openpyxl.load_workbook(XLSX, data_only=True)
|
||||||
|
ws = wb['1단계_홈페이지']
|
||||||
|
|
||||||
|
targets = ['괴산군', '단양군', '보은군', '영동군', '옥천군', '음성군',
|
||||||
|
'제천시', '증평군', '진천군', '청주시', '충주시']
|
||||||
|
|
||||||
|
for name in targets:
|
||||||
|
for r in range(5, ws.max_row + 1):
|
||||||
|
cell_name = ws.cell(row=r, column=4).value
|
||||||
|
if cell_name and ('충청북도' in str(cell_name) or '충북' in str(cell_name)) and name in str(cell_name):
|
||||||
|
daebun = ws.cell(row=r, column=2).value
|
||||||
|
jungbun = ws.cell(row=r, column=3).value
|
||||||
|
url = ws.cell(row=r, column=5).value
|
||||||
|
print(f'Row {r}: 대={daebun!r} 중={jungbun!r} 기관명={cell_name!r} URL={url!r}')
|
||||||
|
break
|
||||||
|
else:
|
||||||
|
print(f'NOT FOUND: {name}')
|
||||||
27
_스크립트/_extract_chungnam.py
Normal file
27
_스크립트/_extract_chungnam.py
Normal file
@ -0,0 +1,27 @@
|
|||||||
|
"""Find how Chungnam cities are categorized in the master file."""
|
||||||
|
import openpyxl
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
ROOT = Path(r'D:\01.프로젝트\DB수집')
|
||||||
|
XLSX = ROOT / '붙임1_공공저작물 개방 대상기관(1160개) 실태조사 목록(신유형개방지원사업) 양식_260518.xlsx'
|
||||||
|
wb = openpyxl.load_workbook(XLSX, data_only=True)
|
||||||
|
ws = wb['1단계_홈페이지']
|
||||||
|
|
||||||
|
# Known 충청남도 cities to search
|
||||||
|
targets = ['계룡시', '공주시', '금산군', '논산시', '당진시', '보령시',
|
||||||
|
'부여군', '서산시', '서천군', '아산시', '예산군', '천안시',
|
||||||
|
'청양군', '태안군', '홍성군']
|
||||||
|
|
||||||
|
for name in targets:
|
||||||
|
found = False
|
||||||
|
for r in range(5, ws.max_row + 1):
|
||||||
|
cell_name = ws.cell(row=r, column=4).value
|
||||||
|
if cell_name and name in str(cell_name):
|
||||||
|
daebun = ws.cell(row=r, column=2).value
|
||||||
|
jungbun = ws.cell(row=r, column=3).value
|
||||||
|
url = ws.cell(row=r, column=5).value
|
||||||
|
print(f'Row {r}: 대={daebun!r} 중={jungbun!r} 기관명={cell_name!r} URL={url!r}')
|
||||||
|
found = True
|
||||||
|
break
|
||||||
|
if not found:
|
||||||
|
print(f'NOT FOUND: {name}')
|
||||||
52
_스크립트/_fill_C_province.py
Normal file
52
_스크립트/_fill_C_province.py
Normal file
@ -0,0 +1,52 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""미검수 27곳 C열(사이트명)을 "광역 기관"(공백구분)으로 채움. 검수완료 제외.
|
||||||
|
사용: python -X utf8 _fill_C_province.py [--write]
|
||||||
|
"""
|
||||||
|
import sys, os, re, shutil, importlib.util
|
||||||
|
import openpyxl
|
||||||
|
|
||||||
|
HERE=os.path.dirname(os.path.abspath(__file__))
|
||||||
|
MODULES=['_chungnam_phase234_all.py','_chungbuk_phase234_all.py','_jeonbuk_phase234_all.py']
|
||||||
|
TARGETS={'서산시','아산시','천안시','청양군','홍성군',
|
||||||
|
'괴산군','단양군','보은군','영동군','옥천군','음성군','제천시','증평군','진천군','청주시','충주시',
|
||||||
|
'부안군','순창군','완주군','익산시','임실군','장수군','전주시','정읍시','진안군','서귀포시','제주시'}
|
||||||
|
|
||||||
|
def load():
|
||||||
|
sites={}
|
||||||
|
for f in MODULES:
|
||||||
|
sp=importlib.util.spec_from_file_location(f[:-3],os.path.join(HERE,f));m=importlib.util.module_from_spec(sp);sp.loader.exec_module(m)
|
||||||
|
sites.update(m.SITES)
|
||||||
|
return sites
|
||||||
|
|
||||||
|
def province_of(path):
|
||||||
|
m=re.search(r'광역_사이트맵[\\/]([^\\/]+)[\\/]',path)
|
||||||
|
return m.group(1) if m else '?'
|
||||||
|
|
||||||
|
def main():
|
||||||
|
write='--write' in sys.argv
|
||||||
|
sites=load(); locked=[]; done=[]
|
||||||
|
for org in sorted(TARGETS):
|
||||||
|
if org not in sites: print(org,'SITES없음'); continue
|
||||||
|
xlsx=sites[org]['xlsx']
|
||||||
|
prov=province_of(xlsx); newC=f'{prov} {org}'
|
||||||
|
wb=openpyxl.load_workbook(xlsx); ws=wb.active
|
||||||
|
ch=0
|
||||||
|
for r in range(3,ws.max_row+1):
|
||||||
|
content = any(ws.cell(r,c).value not in (None,'') for c in (2,4,5,6,7,8,9,10,11,12))
|
||||||
|
cur=ws.cell(r,3).value
|
||||||
|
if (cur not in (None,'')) or content:
|
||||||
|
if cur!=newC:
|
||||||
|
if write: ws.cell(r,3).value=newC
|
||||||
|
ch+=1
|
||||||
|
if write and ch:
|
||||||
|
try:
|
||||||
|
bak=xlsx.replace('.xlsx','_backup_C광역전.xlsx')
|
||||||
|
if not os.path.exists(bak): shutil.copy(xlsx,bak)
|
||||||
|
wb.save(xlsx)
|
||||||
|
except PermissionError:
|
||||||
|
locked.append(org); print(f'{org:6} ❌파일열림 (미적용)'); continue
|
||||||
|
done.append((org,newC,ch))
|
||||||
|
print(f'{org:6} → "{newC}" 변경 {ch}행 {"[적용]" if write else "[DRY]"}')
|
||||||
|
if locked: print('\n⚠️ 파일열림으로 미적용:',', '.join(locked))
|
||||||
|
|
||||||
|
if __name__=='__main__': main()
|
||||||
116
_스크립트/_fix_emerge_all.py
Normal file
116
_스크립트/_fix_emerge_all.py
Normal file
@ -0,0 +1,116 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""E열(대분류) 병합 누락 보정 — 재크롤링 없이 기존 엑셀의 D/E/F 병합만 바로잡는다.
|
||||||
|
|
||||||
|
원인: phase1이 D를 먼저 병합 → 블록 2행부터 D=None → E 병합(같은 D 안에서만)이
|
||||||
|
D블록 첫 행에서 끊겨 첫 행 E가 단독으로 남음.
|
||||||
|
해결: 현재 병합에서 값 복원 → D/E/F 병합 해제 → 전 행에 값 채움 →
|
||||||
|
F → E → D 순서로 재병합(컬럼을 병합하면 그 컬럼 2행부터 None이 되므로 순서가 중요).
|
||||||
|
|
||||||
|
데이터(D/E/F 텍스트)는 그대로. 헤더 병합(B1:R1 등)은 건드리지 않음.
|
||||||
|
파일별 *_backup_emerge전.xlsx 백업.
|
||||||
|
"""
|
||||||
|
import sys, os, glob, shutil, warnings, openpyxl
|
||||||
|
from openpyxl.styles import Alignment
|
||||||
|
from openpyxl.utils import get_column_letter
|
||||||
|
|
||||||
|
sys.stdout.reconfigure(encoding='utf-8')
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
START = 3
|
||||||
|
COLS = (4, 5, 6) # D, E, F
|
||||||
|
CENTER = Alignment(horizontal='center', vertical='center', wrap_text=False)
|
||||||
|
|
||||||
|
|
||||||
|
def filled_values(ws, col):
|
||||||
|
"""현재 병합 상태에서 각 행의 실제 값을 복원(병합 top값으로 채움)."""
|
||||||
|
end = ws.max_row
|
||||||
|
vals = {r: ws.cell(r, col).value for r in range(START, end + 1)}
|
||||||
|
for mr in ws.merged_cells.ranges:
|
||||||
|
if mr.min_col == col == mr.max_col and mr.min_row >= START:
|
||||||
|
top = ws.cell(mr.min_row, col).value
|
||||||
|
for r in range(mr.min_row, mr.max_row + 1):
|
||||||
|
vals[r] = top
|
||||||
|
return vals, end
|
||||||
|
|
||||||
|
|
||||||
|
def merge_runs(ws, col, end, group_keys, fills):
|
||||||
|
"""fills[col][r] 기준으로 연속 동일 구간 병합. group_keys: 상위 컬럼 튜플."""
|
||||||
|
letter = get_column_letter(col)
|
||||||
|
runs = []
|
||||||
|
def key(r):
|
||||||
|
return tuple(fills[g][r] for g in group_keys)
|
||||||
|
cur_val = fills[col][START]
|
||||||
|
cur_grp = key(START)
|
||||||
|
run_start = START
|
||||||
|
for r in range(START + 1, end + 1):
|
||||||
|
v, g = fills[col][r], key(r)
|
||||||
|
if v == cur_val and g == cur_grp:
|
||||||
|
continue
|
||||||
|
if cur_val not in (None, '') and r - 1 > run_start:
|
||||||
|
runs.append((run_start, r - 1))
|
||||||
|
cur_val, cur_grp, run_start = v, g, r
|
||||||
|
if cur_val not in (None, '') and end > run_start:
|
||||||
|
runs.append((run_start, end))
|
||||||
|
for s, e in runs:
|
||||||
|
ws.merge_cells(f'{letter}{s}:{letter}{e}')
|
||||||
|
ws.cell(s, col).alignment = CENTER
|
||||||
|
return len(runs)
|
||||||
|
|
||||||
|
|
||||||
|
def fix(xlsx):
|
||||||
|
wb = openpyxl.load_workbook(xlsx)
|
||||||
|
ws = wb.active
|
||||||
|
# 1) 값 복원
|
||||||
|
fills, end = {}, None
|
||||||
|
for c in COLS:
|
||||||
|
fills[c], end = filled_values(ws, c)
|
||||||
|
# 2) D/E/F 데이터 병합만 해제 (단일컬럼·row>=START)
|
||||||
|
to_unmerge = [str(mr) for mr in ws.merged_cells.ranges
|
||||||
|
if mr.min_col == mr.max_col and mr.min_col in COLS and mr.min_row >= START]
|
||||||
|
for rng in to_unmerge:
|
||||||
|
ws.unmerge_cells(rng)
|
||||||
|
# 3) 전 행에 값 다시 기입
|
||||||
|
for c in COLS:
|
||||||
|
for r in range(START, end + 1):
|
||||||
|
ws.cell(r, c).value = fills[c][r]
|
||||||
|
# 4) F → E → D 순서 재병합
|
||||||
|
n_f = merge_runs(ws, 6, end, (4, 5), fills)
|
||||||
|
n_e = merge_runs(ws, 5, end, (4,), fills)
|
||||||
|
n_d = merge_runs(ws, 4, end, (), fills)
|
||||||
|
return wb, (n_d, n_e, n_f, len(to_unmerge))
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
sel = [a for a in sys.argv[1:] if not a.startswith('-')]
|
||||||
|
files = []
|
||||||
|
for region in ['충청남도', '충청북도']:
|
||||||
|
for d in sorted(glob.glob(region + r'\*')):
|
||||||
|
if not os.path.isdir(d):
|
||||||
|
continue
|
||||||
|
base = os.path.basename(d)
|
||||||
|
nm = base.split('.', 1)[1] if '.' in base else base
|
||||||
|
prov = os.path.basename(os.path.dirname(d))
|
||||||
|
f = os.path.join(d, f'{prov}_{nm}.xlsx')
|
||||||
|
if os.path.exists(f) and (not sel or nm in sel):
|
||||||
|
files.append((nm, f))
|
||||||
|
print(f'대상 {len(files)}개: {[n for n, _ in files]}')
|
||||||
|
ok = []
|
||||||
|
for nm, f in files:
|
||||||
|
try:
|
||||||
|
bak = f.replace('.xlsx', '') + '_backup_emerge전.xlsx'
|
||||||
|
if not os.path.exists(bak):
|
||||||
|
shutil.copy2(f, bak)
|
||||||
|
wb, (nd, ne, nf, un) = fix(f)
|
||||||
|
wb.save(f)
|
||||||
|
print(f' ✅ {nm:<6} 병합해제 {un:>3} → 재병합 D:{nd} E:{ne} F:{nf}')
|
||||||
|
ok.append(nm)
|
||||||
|
except PermissionError:
|
||||||
|
print(f' ⚠️ {nm:<6} 파일 잠김(열려있음) → 건너뜀')
|
||||||
|
except Exception as e:
|
||||||
|
import traceback; traceback.print_exc()
|
||||||
|
print(f' !! {nm:<6} 오류: {e}')
|
||||||
|
print(f'\n완료 {len(ok)}/{len(files)}')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
91
_스크립트/_geumsan_subtree_M.py
Normal file
91
_스크립트/_geumsan_subtree_M.py
Normal file
@ -0,0 +1,91 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""금산군 메뉴 랜딩행 M = 서브트리 전체 leaf 합 (2026-05-31, 사용자 규칙)
|
||||||
|
여성가족=10 패턴: 메뉴(6자리 prefix)의 모든 자식페이지 X01..X0N에 대해
|
||||||
|
leaf = 라이브 ui-nav_tabs 개수 (탭 없으면 1)
|
||||||
|
M(랜딩행) = Σ leaf. 단 메뉴의 자식 중 게시판이 있으면 SKIP(분리 필요, 수동).
|
||||||
|
대상: 랜딩행 >= 143(여성가족) 이고 그 prefix의 시트행이 1개(sheet=1)인 메뉴.
|
||||||
|
sheet>1(이미 G확장됨: 군민복지050307·하수처리050601·주민참여060603)은 SKIP.
|
||||||
|
"""
|
||||||
|
import sys, re, json, openpyxl
|
||||||
|
import urllib.request, ssl
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
from concurrent.futures import ThreadPoolExecutor
|
||||||
|
|
||||||
|
PATH = r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\3.금산군\충청남도_금산군.xlsx'
|
||||||
|
WRITE = '--write' in sys.argv
|
||||||
|
ctx = ssl.create_default_context(); ctx.check_hostname = False; ctx.verify_mode = ssl.CERT_NONE
|
||||||
|
|
||||||
|
def fetch(u):
|
||||||
|
req = urllib.request.Request(u, headers={'User-Agent': 'Mozilla/5.0'})
|
||||||
|
try:
|
||||||
|
return urllib.request.urlopen(req, context=ctx, timeout=15).read().decode('utf-8', 'replace')
|
||||||
|
except Exception:
|
||||||
|
return None
|
||||||
|
|
||||||
|
def analyze(code):
|
||||||
|
sub = 'sub06' if code.startswith('06') else 'sub05'
|
||||||
|
h = fetch(f'https://www.geumsan.go.kr/kr/html/{sub}/{code}.html')
|
||||||
|
if h is None:
|
||||||
|
return None
|
||||||
|
soup = BeautifulSoup(h, 'html.parser')
|
||||||
|
t = soup.select_one('.location_wrap li:last-child, h2.h2')
|
||||||
|
title = t.get_text(strip=True) if t else '?'
|
||||||
|
ntab = len(soup.select('ul.ui-nav_tabs a.ui-tabs_link'))
|
||||||
|
board = bool(soup.select('table.bbs,.board_list,.bbs_list,.pagination')) or bool(re.search(r'총\s*[\d,]+\s*건', str(soup)))
|
||||||
|
return dict(code=code, title=title, ntab=ntab, board=board, leaf=max(1, ntab))
|
||||||
|
|
||||||
|
wb = openpyxl.load_workbook(PATH); ws = wb.active
|
||||||
|
|
||||||
|
# prefix -> sheet rows (sub05/sub06 8-digit)
|
||||||
|
pref_rows = {}
|
||||||
|
for r in range(3, ws.max_row + 1):
|
||||||
|
k = ws.cell(r, 11).value
|
||||||
|
if not k:
|
||||||
|
continue
|
||||||
|
m = re.search(r'/sub0[56]/(\d{6})(\d{2})\.html', str(k))
|
||||||
|
if m:
|
||||||
|
pref_rows.setdefault(m.group(1), []).append((r, m.group(0)))
|
||||||
|
|
||||||
|
targets = {}
|
||||||
|
for pfx, rows in pref_rows.items():
|
||||||
|
landing = min(r for r, _ in rows)
|
||||||
|
if landing < 143:
|
||||||
|
continue # 여성가족(143) 위는 제외
|
||||||
|
if len(rows) != 1:
|
||||||
|
print(f'SKIP {pfx} (sheet={len(rows)} rows={[r for r,_ in rows]}) — 이미 확장됨')
|
||||||
|
continue
|
||||||
|
targets[pfx] = landing
|
||||||
|
|
||||||
|
print(f'대상 sheet=1 메뉴: {len(targets)}')
|
||||||
|
plan = []
|
||||||
|
for pfx, landing in sorted(targets.items(), key=lambda x: x[1]):
|
||||||
|
codes = [f'{pfx}{xx:02d}' for xx in range(1, 16)]
|
||||||
|
kids = [k for k in ThreadPoolExecutor(8).map(analyze, codes) if k]
|
||||||
|
leaf_sum = sum(k['leaf'] for k in kids)
|
||||||
|
nboard = sum(1 for k in kids if k['board'])
|
||||||
|
cur_m = ws.cell(landing, 13).value
|
||||||
|
cur_f = ws.cell(landing, 6).value or ws.cell(landing, 7).value
|
||||||
|
flag = ' ⚠GESIPAN' if nboard else ''
|
||||||
|
plan.append(dict(pfx=pfx, row=landing, m_old=cur_m, m_new=leaf_sum, nkids=len(kids), nboard=nboard, kids=kids))
|
||||||
|
print(f'row{landing} {pfx} [{cur_f}] M {cur_m}->{leaf_sum} (자식{len(kids)},board{nboard}){flag}')
|
||||||
|
if nboard:
|
||||||
|
for k in kids:
|
||||||
|
print(f' {k["code"]} leaf{k["leaf"]} board={k["board"]} {k["title"][:20]}')
|
||||||
|
|
||||||
|
# 적용: 게시판 없는 메뉴만 M 갱신
|
||||||
|
applied = 0
|
||||||
|
for p in plan:
|
||||||
|
if p['nboard'] == 0:
|
||||||
|
ws.cell(p['row'], 13).value = p['m_new']
|
||||||
|
applied += 1
|
||||||
|
else:
|
||||||
|
print(f' ⚠ row{p["row"]} {p["pfx"]} 게시판 포함 → M 미적용(수동 검토)')
|
||||||
|
print(f'적용 대상(게시판없음): {applied}/{len(plan)}')
|
||||||
|
|
||||||
|
json.dump([{k: v for k, v in p.items() if k != 'kids'} for p in plan],
|
||||||
|
open(r'D:\01.프로젝트\DB수집\_temp\_geumsan_subtreeM_plan.json', 'w', encoding='utf8'), ensure_ascii=False)
|
||||||
|
|
||||||
|
if WRITE:
|
||||||
|
wb.save(PATH); print('SAVED')
|
||||||
|
else:
|
||||||
|
print('DRY-RUN (--write 로 저장)')
|
||||||
102
_스크립트/_geumsan_tab_recheck.py
Normal file
102
_스크립트/_geumsan_tab_recheck.py
Normal file
@ -0,0 +1,102 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""금산군 ui-nav_tabs 재검토 + 군민복지 cross-product 정리 (2026-05-31)
|
||||||
|
- PHASE1: ui-nav_tabs 인페이지 탭 페이지 → 1행 유지, M=라이브 탭 개수
|
||||||
|
여성가족 05030301=3, 어르신 05030401=3, 사회복지 05030601=5(무변경),
|
||||||
|
읍·면 행정복지센터 06X0101 ×10 =2
|
||||||
|
- PHASE2: 군민복지(05030701) tab-ul type1 4탭 cross-product(147~159, 13행)
|
||||||
|
→ 4행으로 축소(공지사항/금산군자원봉사센터/복지시설/관련사이트). 151~159 삭제.
|
||||||
|
"""
|
||||||
|
import sys, openpyxl
|
||||||
|
from copy import copy
|
||||||
|
from openpyxl.utils import range_boundaries
|
||||||
|
|
||||||
|
PATH = r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\3.금산군\충청남도_금산군.xlsx'
|
||||||
|
WRITE = '--write' in sys.argv
|
||||||
|
|
||||||
|
wb = openpyxl.load_workbook(PATH)
|
||||||
|
ws = wb.active
|
||||||
|
MAXC = ws.max_column
|
||||||
|
|
||||||
|
# ---------- PHASE 1: M 정규화 (URL 매칭) ----------
|
||||||
|
MTARGETS = {
|
||||||
|
'05030301': 3, '05030401': 3, '05030601': 5,
|
||||||
|
'06010101': 2, '06020101': 2, '06030101': 2, '06040101': 2, '06050101': 2,
|
||||||
|
'06060101': 2, '06070101': 2, '06080101': 2, '06090101': 2, '06100101': 2,
|
||||||
|
}
|
||||||
|
print('=== PHASE1: M 정규화 ===')
|
||||||
|
for r in range(3, ws.max_row + 1):
|
||||||
|
k = ws.cell(r, 11).value
|
||||||
|
if not k:
|
||||||
|
continue
|
||||||
|
ks = str(k)
|
||||||
|
for code, mval in MTARGETS.items():
|
||||||
|
if code + '.html' in ks:
|
||||||
|
old_m = ws.cell(r, 13).value
|
||||||
|
old_l = ws.cell(r, 12).value
|
||||||
|
ws.cell(r, 13).value = mval
|
||||||
|
print(f' row{r} {code} L={old_l} M:{old_m}->{mval}')
|
||||||
|
break
|
||||||
|
|
||||||
|
# ---------- PHASE 2: 군민복지 restructure ----------
|
||||||
|
print('=== PHASE2: 군민복지 정리 ===')
|
||||||
|
# Step A: 147~150 G<-H, clear H
|
||||||
|
for r in range(147, 151):
|
||||||
|
g_old = ws.cell(r, 7).value
|
||||||
|
h = ws.cell(r, 8).value
|
||||||
|
ws.cell(r, 7).value = h
|
||||||
|
ws.cell(r, 8).value = None
|
||||||
|
print(f' row{r} G:{g_old}->{h} (H clear)')
|
||||||
|
|
||||||
|
DEL_START, DEL_COUNT = 151, 9
|
||||||
|
|
||||||
|
# Step B: capture & clear merges
|
||||||
|
merges = [str(mc) for mc in ws.merged_cells.ranges]
|
||||||
|
for mc in list(ws.merged_cells.ranges):
|
||||||
|
ws.unmerge_cells(str(mc))
|
||||||
|
|
||||||
|
# Step C: manual shift up (delete 151~159)
|
||||||
|
max_row = ws.max_row
|
||||||
|
for r in range(DEL_START, max_row - DEL_COUNT + 1):
|
||||||
|
src = r + DEL_COUNT
|
||||||
|
for c in range(1, MAXC + 1):
|
||||||
|
s = ws.cell(src, c); d = ws.cell(r, c)
|
||||||
|
d.value = s.value
|
||||||
|
if s.has_style:
|
||||||
|
d._style = copy(s._style)
|
||||||
|
d.number_format = s.number_format
|
||||||
|
d.hyperlink = None
|
||||||
|
if s.hyperlink:
|
||||||
|
d.hyperlink = copy(s.hyperlink)
|
||||||
|
d.hyperlink.ref = d.coordinate
|
||||||
|
for r in range(max_row - DEL_COUNT + 1, max_row + 1):
|
||||||
|
for c in range(1, MAXC + 1):
|
||||||
|
d = ws.cell(r, c); d.value = None; d.hyperlink = None
|
||||||
|
|
||||||
|
# Step D: rebuild merges with row mapping
|
||||||
|
def surv(x):
|
||||||
|
return x < DEL_START or x >= DEL_START + DEL_COUNT
|
||||||
|
def mp(x):
|
||||||
|
return x if x < DEL_START else x - DEL_COUNT
|
||||||
|
for mc in merges:
|
||||||
|
c1, r1, c2, r2 = range_boundaries(mc)
|
||||||
|
s = [x for x in range(r1, r2 + 1) if surv(x)]
|
||||||
|
if not s:
|
||||||
|
continue
|
||||||
|
nr1, nr2 = mp(min(s)), mp(max(s))
|
||||||
|
if nr1 == nr2 and c1 == c2:
|
||||||
|
continue # collapsed to single cell, no merge
|
||||||
|
ws.merge_cells(start_row=nr1, start_column=c1, end_row=nr2, end_column=c2)
|
||||||
|
|
||||||
|
# Step E: 순번 재번호 (151행 이후 -9 → 연속성 유지)
|
||||||
|
for r in range(DEL_START, ws.max_row + 1):
|
||||||
|
v = ws.cell(r, 2).value
|
||||||
|
if isinstance(v, int):
|
||||||
|
ws.cell(r, 2).value = v - DEL_COUNT
|
||||||
|
|
||||||
|
print(f' deleted rows {DEL_START}~{DEL_START+DEL_COUNT-1} (9 rows)')
|
||||||
|
|
||||||
|
if WRITE:
|
||||||
|
wb.save(PATH)
|
||||||
|
print('SAVED', PATH)
|
||||||
|
else:
|
||||||
|
print('DRY-RUN (use --write to save)')
|
||||||
87
_스크립트/_gongju_apply.py
Normal file
87
_스크립트/_gongju_apply.py
Normal file
@ -0,0 +1,87 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""공주시 227~414행 재검수 적용.
|
||||||
|
규칙(226행 검수 캘리브레이션 결과):
|
||||||
|
R1 관광지 dongList(tursmCn): L 페이지->게시판, M=썸네일수
|
||||||
|
R2 게시판 글수 재계산: M=현재 크롤 글수 (board_count 있는 행, 값이 다를 때만)
|
||||||
|
R3 찾아오시는길/오시는길: N=어문,이미지
|
||||||
|
DRY 기본. --write 시 백업 후 저장.
|
||||||
|
"""
|
||||||
|
import sys, json, shutil
|
||||||
|
import openpyxl
|
||||||
|
|
||||||
|
XLSX = r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\2.공주시\충청남도_공주시.xlsx'
|
||||||
|
EV = r'D:\01.프로젝트\DB수집\_스크립트\_evidence.json'
|
||||||
|
LO, HI = 227, 414
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
write = '--write' in sys.argv
|
||||||
|
ev = json.load(open(EV, encoding='utf-8'))
|
||||||
|
wb = openpyxl.load_workbook(XLSX)
|
||||||
|
ws = wb.active
|
||||||
|
changes = [] # (row, col, colname, old, new, rule)
|
||||||
|
drift = [] # 미세 드리프트(원본 유지) 보고용
|
||||||
|
|
||||||
|
for r in range(LO, HI + 1):
|
||||||
|
e = ev.get(str(r))
|
||||||
|
if not e:
|
||||||
|
continue
|
||||||
|
url = (ws.cell(r, 11).value or '')
|
||||||
|
L = ws.cell(r, 12).value
|
||||||
|
M = ws.cell(r, 13).value
|
||||||
|
N = ws.cell(r, 14).value
|
||||||
|
Gname = ws.cell(r, 7).value or ''
|
||||||
|
Fname = ws.cell(r, 6).value or ''
|
||||||
|
name = '%s %s' % (Fname, Gname)
|
||||||
|
|
||||||
|
# R1 관광지 dongList
|
||||||
|
if 'tursmCn' in url and e.get('tursm_count'):
|
||||||
|
tc = e['tursm_count']
|
||||||
|
if L != '게시판':
|
||||||
|
changes.append((r, 12, 'L', L, '게시판', 'R1관광지'))
|
||||||
|
if M != tc:
|
||||||
|
changes.append((r, 13, 'M', M, tc, 'R1관광지'))
|
||||||
|
continue # 관광지는 M 규칙2 적용 안함
|
||||||
|
|
||||||
|
# R2 게시판 글수 재계산 (0/빈 값만 복구; 유효 비0값의 미세 드리프트는 원본 유지)
|
||||||
|
if L == '게시판' and e.get('board_count') is not None:
|
||||||
|
bc = e['board_count']
|
||||||
|
if (M in (None, 0, '0')) and bc:
|
||||||
|
changes.append((r, 13, 'M', M, bc, 'R2게시판글수복구'))
|
||||||
|
elif M != bc:
|
||||||
|
drift.append((r, M, bc, ws.cell(r, 7).value or ws.cell(r, 6).value or ''))
|
||||||
|
|
||||||
|
# R3 찾아오시는길 N
|
||||||
|
if ('찾아오시는' in name) or ('오시는길' in name) or ('오시는 길' in name):
|
||||||
|
if N == '어문':
|
||||||
|
changes.append((r, 14, 'N', N, '어문,이미지', 'R3오시는길'))
|
||||||
|
|
||||||
|
# report
|
||||||
|
print('=== 제안 변경 %d건 (DRY) ===' % len(changes))
|
||||||
|
by = {}
|
||||||
|
for r, c, cn, old, new, rule in changes:
|
||||||
|
by.setdefault(rule, []).append((r, cn, old, new))
|
||||||
|
for rule in sorted(by):
|
||||||
|
print('\n[%s] %d건' % (rule, len(by[rule])))
|
||||||
|
for r, cn, old, new in by[rule]:
|
||||||
|
g = ws.cell(r, 7).value or ws.cell(r, 6).value or ws.cell(r, 5).value or ''
|
||||||
|
print(' r%d %s: %r -> %r (%s)' % (r, cn, old, new, g))
|
||||||
|
|
||||||
|
if drift:
|
||||||
|
print('\n[미세 드리프트: 원본 유지, 참고용] %d건' % len(drift))
|
||||||
|
for r, old, new, g in drift:
|
||||||
|
print(' r%d 게시판글수 원본 %s (현재 크롤 %s, 차이 %+d) %s' % (r, old, new, new - (old or 0), g))
|
||||||
|
|
||||||
|
if write:
|
||||||
|
bak = XLSX.replace('.xlsx', '_backup_재검수227전.xlsx')
|
||||||
|
shutil.copy(XLSX, bak)
|
||||||
|
for r, c, cn, old, new, rule in changes:
|
||||||
|
ws.cell(r, c).value = new
|
||||||
|
wb.save(XLSX)
|
||||||
|
print('\n저장 완료. 백업:', bak)
|
||||||
|
else:
|
||||||
|
print('\n(DRY — 적용하려면 --write)')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
82
_스크립트/_gongju_collapse.py
Normal file
82
_스크립트/_gongju_collapse.py
Normal file
@ -0,0 +1,82 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""공주시 전용 규칙 B: F 하나에 G가 여러 개이고 그 G들이 전부 L=페이지면
|
||||||
|
→ 1행으로 합치고(G/H/I/J 비움), K=첫 G의 URL, 수량(M)=G들의 M 합.
|
||||||
|
G 중 하나라도 페이지가 아니면(게시판/사이트 등) 그 그룹은 합치지 않음.
|
||||||
|
|
||||||
|
평탄화/재병합/스타일은 _dedup_all 재사용.
|
||||||
|
사용: python -X utf8 _gongju_collapse.py [--write]
|
||||||
|
"""
|
||||||
|
import sys, shutil, importlib.util, warnings
|
||||||
|
import openpyxl
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
spec = importlib.util.spec_from_file_location('dd', __file__.replace('_gongju_collapse.py', '_dedup_all.py'))
|
||||||
|
dd = importlib.util.module_from_spec(spec); spec.loader.exec_module(dd)
|
||||||
|
|
||||||
|
XLSX = r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\2.공주시\충청남도_공주시.xlsx'
|
||||||
|
|
||||||
|
|
||||||
|
def mval(x):
|
||||||
|
try:
|
||||||
|
return int(x)
|
||||||
|
except Exception:
|
||||||
|
return 1
|
||||||
|
|
||||||
|
|
||||||
|
def collapse(rows):
|
||||||
|
out, groups = [], []
|
||||||
|
i, n = 0, len(rows)
|
||||||
|
while i < n:
|
||||||
|
v = rows[i]['vals']
|
||||||
|
D, E, F, G = v.get(4), v.get(5), v.get(6), v.get(7)
|
||||||
|
if F not in (None, '') and G not in (None, ''):
|
||||||
|
j = i
|
||||||
|
while (j < n and rows[j]['vals'].get(4) == D and rows[j]['vals'].get(5) == E
|
||||||
|
and rows[j]['vals'].get(6) == F and rows[j]['vals'].get(7) not in (None, '')):
|
||||||
|
j += 1
|
||||||
|
grp = rows[i:j]
|
||||||
|
all_page = all(g['vals'].get(12) == '페이지' for g in grp)
|
||||||
|
leaf = all(all(g['vals'].get(c) in (None, '') for c in (8, 9, 10)) for g in grp)
|
||||||
|
if len(grp) >= 2 and all_page and leaf:
|
||||||
|
msum = sum(mval(g['vals'].get(13)) for g in grp)
|
||||||
|
first = dict(grp[0]); nv = dict(first['vals'])
|
||||||
|
for c in (7, 8, 9, 10):
|
||||||
|
nv[c] = None
|
||||||
|
nv[13] = msum
|
||||||
|
first['vals'] = nv
|
||||||
|
out.append(first)
|
||||||
|
groups.append({'D': D, 'E': E, 'F': F, 'n': len(grp),
|
||||||
|
'ms': [mval(g['vals'].get(13)) for g in grp],
|
||||||
|
'gs': [g['vals'].get(7) for g in grp],
|
||||||
|
'sum': msum, 'k': nv.get(11)})
|
||||||
|
i = j
|
||||||
|
continue
|
||||||
|
out.extend(grp); i = j; continue
|
||||||
|
out.append(rows[i]); i += 1
|
||||||
|
return out, groups
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
write = '--write' in sys.argv
|
||||||
|
wb = openpyxl.load_workbook(XLSX)
|
||||||
|
ws = wb.active
|
||||||
|
rows = dd.load_flat(ws)
|
||||||
|
out, groups = collapse(rows)
|
||||||
|
removed = len(rows) - len(out)
|
||||||
|
print(f'합칠 F그룹: {len(groups)}개 / 제거행 {removed} ({len(rows)}→{len(out)})\n')
|
||||||
|
for g in groups:
|
||||||
|
ms = '+'.join(str(m) for m in g['ms'])
|
||||||
|
print(f" [{g['E']} > {g['F']}] G {g['n']}개({'/'.join(str(x) for x in g['gs'])[:50]})")
|
||||||
|
print(f" → 1행, M={ms}={g['sum']}, K={g['k']}")
|
||||||
|
if write and groups:
|
||||||
|
bak = XLSX.replace('.xlsx', '_backup_collapse전.xlsx')
|
||||||
|
shutil.copy(XLSX, bak)
|
||||||
|
dd.write_back(ws, out)
|
||||||
|
wb.save(XLSX)
|
||||||
|
print(f'\n저장 완료. 백업: {bak}')
|
||||||
|
elif not write:
|
||||||
|
print('\n(DRY — 실제 적용하려면 --write)')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
120
_스크립트/_gongju_evidence.py
Normal file
120
_스크립트/_gongju_evidence.py
Normal file
@ -0,0 +1,120 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""공주시 227~414행 재검수용 증거 수집(쓰기 없음).
|
||||||
|
각 URL을 크롤링해 게시판 글수/관광지 썸네일수/외부링크/공공누리마크/#nav탭수/본문이미지·영상 신호를 수집.
|
||||||
|
검증용으로 226행까지 검수에서 바뀐 행 일부도 같이 수집.
|
||||||
|
출력: _스크립트/_evidence.json + 콘솔 요약
|
||||||
|
"""
|
||||||
|
import sys, json, re, warnings
|
||||||
|
from urllib.parse import urlparse
|
||||||
|
from concurrent.futures import ThreadPoolExecutor, as_completed
|
||||||
|
import requests
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
import openpyxl
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
XLSX = r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\2.공주시\충청남도_공주시.xlsx'
|
||||||
|
H = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) Chrome/120 Safari/537.36'}
|
||||||
|
COMMON_IMG = re.compile(r'(flag\.jpg|slogan|/common/|/template/|btn_|/btn|icon|blank|_mark\.png|sns|share|loading)', re.I)
|
||||||
|
|
||||||
|
|
||||||
|
def fetch(row, url):
|
||||||
|
try:
|
||||||
|
r = requests.get(url, headers=H, timeout=25, verify=False)
|
||||||
|
r.encoding = r.apparent_encoding or 'utf-8'
|
||||||
|
return row, url, r.status_code, r.text
|
||||||
|
except Exception as e:
|
||||||
|
return row, url, None, ('ERR:%s' % e)
|
||||||
|
|
||||||
|
|
||||||
|
def analyze(url, html):
|
||||||
|
ev = {}
|
||||||
|
host = urlparse(url).netloc
|
||||||
|
ev['host'] = host
|
||||||
|
ev['external'] = ('gongju.go.kr' not in host)
|
||||||
|
if not html or html.startswith('ERR:'):
|
||||||
|
ev['error'] = html
|
||||||
|
return ev
|
||||||
|
s = BeautifulSoup(html, 'html.parser')
|
||||||
|
# board total count
|
||||||
|
cnt_el = s.select_one('.program--count strong')
|
||||||
|
if cnt_el:
|
||||||
|
m = re.sub(r'[^0-9]', '', cnt_el.get_text())
|
||||||
|
ev['board_count'] = int(m) if m else None
|
||||||
|
ev['is_board'] = True
|
||||||
|
else:
|
||||||
|
ev['is_board'] = False
|
||||||
|
# tursmCn thumbnails (관광지)
|
||||||
|
tt = len(re.findall(r'/thumbnail/tursmCn/', html))
|
||||||
|
if tt:
|
||||||
|
ev['tursm_count'] = tt
|
||||||
|
# content images (static html, excludes template/common) — JS maps missed
|
||||||
|
cimgs = []
|
||||||
|
for im in s.find_all('img'):
|
||||||
|
src = im.get('src') or im.get('data-src') or ''
|
||||||
|
if src and not COMMON_IMG.search(src):
|
||||||
|
cimgs.append(src)
|
||||||
|
ev['content_img'] = len(cimgs)
|
||||||
|
ev['content_img_sample'] = cimgs[:4]
|
||||||
|
# video signals
|
||||||
|
low = html.lower()
|
||||||
|
ev['video'] = bool(re.search(r'youtube\.com/embed|youtu\.be/|player\.vimeo|<video|\.mp4|data-video', low))
|
||||||
|
# KOGL / 공공누리
|
||||||
|
kogl_hits = re.findall(r'(opentype0?[1-4]|kogl[_\-]?[1-4]?|공공누리)', html, re.I)
|
||||||
|
ev['kogl'] = bool(kogl_hits)
|
||||||
|
ev['kogl_sample'] = list(dict.fromkeys(kogl_hits))[:6]
|
||||||
|
# explicit opentype number
|
||||||
|
ot = re.findall(r'opentype0?([1-4])', html, re.I)
|
||||||
|
if ot:
|
||||||
|
ev['kogl_types'] = sorted(set(int(x) for x in ot))
|
||||||
|
# #nav tab count (규칙 A)
|
||||||
|
best = 0
|
||||||
|
for ul in s.find_all('ul'):
|
||||||
|
cls = ' '.join(ul.get('class') or []).lower()
|
||||||
|
if 'tab-ul' in cls:
|
||||||
|
anchors = [a for a in ul.find_all('a')
|
||||||
|
if (a.get('href') or '').strip().startswith('#') and a.get_text(strip=True)]
|
||||||
|
if len(anchors) >= 2:
|
||||||
|
best = max(best, len(anchors))
|
||||||
|
if best:
|
||||||
|
ev['nav_tabs'] = best
|
||||||
|
return ev
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
lo, hi = 227, 414
|
||||||
|
if len(sys.argv) > 2:
|
||||||
|
lo, hi = int(sys.argv[1]), int(sys.argv[2])
|
||||||
|
wb = openpyxl.load_workbook(XLSX)
|
||||||
|
ws = wb.active
|
||||||
|
targets = []
|
||||||
|
rowinfo = {}
|
||||||
|
for r in range(lo, hi + 1):
|
||||||
|
u = ws.cell(r, 11).value
|
||||||
|
if isinstance(u, str) and u.strip().startswith('http'):
|
||||||
|
targets.append((r, u.strip()))
|
||||||
|
rowinfo[r] = {
|
||||||
|
'D': ws.cell(r, 4).value, 'E': ws.cell(r, 5).value, 'F': ws.cell(r, 6).value,
|
||||||
|
'G': ws.cell(r, 7).value, 'K': u.strip(), 'L': ws.cell(r, 12).value,
|
||||||
|
'M': ws.cell(r, 13).value, 'N': ws.cell(r, 14).value, 'O': ws.cell(r, 15).value,
|
||||||
|
}
|
||||||
|
print('크롤 대상:', len(targets), '행', lo, '~', hi)
|
||||||
|
out = {}
|
||||||
|
done = 0
|
||||||
|
with ThreadPoolExecutor(max_workers=10) as ex:
|
||||||
|
futs = [ex.submit(fetch, r, u) for r, u in targets]
|
||||||
|
for f in as_completed(futs):
|
||||||
|
row, url, st, html = f.result()
|
||||||
|
ev = analyze(url, html)
|
||||||
|
ev['status'] = st
|
||||||
|
ev['row'] = row
|
||||||
|
out[row] = {**rowinfo[row], **ev}
|
||||||
|
done += 1
|
||||||
|
if done % 25 == 0:
|
||||||
|
print(' ...', done, '/', len(targets))
|
||||||
|
json.dump(out, open(r'D:\01.프로젝트\DB수집\_스크립트\_evidence.json', 'w', encoding='utf-8'),
|
||||||
|
ensure_ascii=False, indent=1)
|
||||||
|
print('저장: _evidence.json (', len(out), '행 )')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
91
_스크립트/_gongju_navcount.py
Normal file
91
_스크립트/_gongju_navcount.py
Normal file
@ -0,0 +1,91 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""공주시 전용: 본문에 인페이지 앵커 탭(ul.tab-ul 안 a[href^="#"], 예: #nav1~4)이 있는
|
||||||
|
페이지는 1행 유지하되 수량(M, 13열)을 그 탭 개수로 기입.
|
||||||
|
|
||||||
|
판별: class 에 'tab-ul' 포함한 ul 안에서 href 가 '#' 로 시작하는 탭 <a> 개수(>=2).
|
||||||
|
한 페이지에 그런 탭그룹이 여럿이면 가장 큰 그룹의 탭 수.
|
||||||
|
|
||||||
|
사용: python -X utf8 _gongju_navcount.py [--write]
|
||||||
|
"""
|
||||||
|
import sys, warnings
|
||||||
|
from concurrent.futures import ThreadPoolExecutor, as_completed
|
||||||
|
import openpyxl, requests, shutil
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
XLSX = r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\2.공주시\충청남도_공주시.xlsx'
|
||||||
|
DOMAIN = 'gongju.go.kr'
|
||||||
|
H = {'User-Agent': 'Mozilla/5.0 Chrome/120 Safari/537.36'}
|
||||||
|
|
||||||
|
|
||||||
|
def nav_count(html):
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
best = 0
|
||||||
|
for ul in soup.find_all('ul'):
|
||||||
|
cls = ' '.join(ul.get('class') or []).lower()
|
||||||
|
if 'tab-ul' not in cls:
|
||||||
|
continue
|
||||||
|
anchors = [a for a in ul.find_all('a')
|
||||||
|
if (a.get('href') or '').strip().startswith('#') and a.get_text(strip=True)]
|
||||||
|
if len(anchors) >= 2:
|
||||||
|
best = max(best, len(anchors))
|
||||||
|
return best
|
||||||
|
|
||||||
|
|
||||||
|
def fetch(r, u):
|
||||||
|
try:
|
||||||
|
x = requests.get(u, headers=H, timeout=15, verify=False)
|
||||||
|
return r, u, x.content
|
||||||
|
except Exception:
|
||||||
|
return r, u, None
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
write = '--write' in sys.argv
|
||||||
|
wb = openpyxl.load_workbook(XLSX)
|
||||||
|
ws = wb.active
|
||||||
|
targets = []
|
||||||
|
for r in range(3, ws.max_row + 1):
|
||||||
|
u = ws.cell(r, 11).value
|
||||||
|
if isinstance(u, str) and DOMAIN in u:
|
||||||
|
targets.append((r, u))
|
||||||
|
print(f'스캔 대상(동일도메인): {len(targets)}행')
|
||||||
|
|
||||||
|
res = {}
|
||||||
|
with ThreadPoolExecutor(max_workers=8) as ex:
|
||||||
|
for f in as_completed([ex.submit(fetch, r, u) for r, u in targets]):
|
||||||
|
r, u, html = f.result()
|
||||||
|
if html:
|
||||||
|
c = nav_count(html)
|
||||||
|
if c >= 2:
|
||||||
|
res[r] = (u, c)
|
||||||
|
|
||||||
|
print(f'\n인페이지 탭(#) 보유 페이지: {len(res)}건')
|
||||||
|
print(f"{'행':>4} {'현L':<5}{'현M':>4} → {'새M':>4} F/카테고리 | URL")
|
||||||
|
print('-' * 90)
|
||||||
|
chg = 0
|
||||||
|
for r in sorted(res):
|
||||||
|
u, c = res[r]
|
||||||
|
L = ws.cell(r, 12).value or ''
|
||||||
|
M = ws.cell(r, 13).value
|
||||||
|
cat = ws.cell(r, 6).value or ws.cell(r, 5).value or ws.cell(r, 4).value or ''
|
||||||
|
mark = '' if str(M) == str(c) else '★'
|
||||||
|
if str(M) != str(c):
|
||||||
|
chg += 1
|
||||||
|
print(f"{r:>4} {str(L):<5}{str(M):>4} → {c:>4}{mark} {cat} | {u}")
|
||||||
|
if write:
|
||||||
|
ws.cell(r, 13).value = c
|
||||||
|
|
||||||
|
print('-' * 90)
|
||||||
|
print(f'변경 대상: {chg}행')
|
||||||
|
if write and chg:
|
||||||
|
bak = XLSX.replace('.xlsx', '_backup_navcount전.xlsx')
|
||||||
|
shutil.copy(XLSX, bak)
|
||||||
|
wb.save(XLSX)
|
||||||
|
print(f'저장 완료. 백업: {bak}')
|
||||||
|
elif not write:
|
||||||
|
print('(DRY — 실제 기입하려면 --write)')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
109
_스크립트/_gongju_navtab_board.py
Normal file
109
_스크립트/_gongju_navtab_board.py
Normal file
@ -0,0 +1,109 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""공주시: 인페이지(#nav) 탭을 가진 페이지의 각 탭 패널 안에 '게시판'이 있는지 점검.
|
||||||
|
|
||||||
|
대상: _gongju_navcount 에서 잡힌 14개 페이지(동일도메인 자동 재탐지).
|
||||||
|
각 탭 <a href="#navN"> → 패널 element(id=navN) 내부에서 게시판 신호 탐지:
|
||||||
|
· 목록 table(td 안 링크 다수) · 페이징 · '총 N건' · 상세링크(view.do/mode=V/nttId/BBSMSTR) · iframe
|
||||||
|
판정: 신호 1+ → 그 탭은 '게시판 포함' 가능.
|
||||||
|
|
||||||
|
사용: python -X utf8 _gongju_navtab_board.py
|
||||||
|
"""
|
||||||
|
import re, warnings
|
||||||
|
from concurrent.futures import ThreadPoolExecutor, as_completed
|
||||||
|
import openpyxl, requests
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
XLSX = r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\2.공주시\충청남도_공주시.xlsx'
|
||||||
|
DOMAIN = 'gongju.go.kr'
|
||||||
|
H = {'User-Agent': 'Mozilla/5.0 Chrome/120 Safari/537.36'}
|
||||||
|
DETAIL = re.compile(r'(view\.do|mode=V|nttId|BBSMSTR|selectBoard|selectBbs)', re.I)
|
||||||
|
TOTAL = re.compile(r'총\s*[\d,]+\s*(건|개)')
|
||||||
|
|
||||||
|
|
||||||
|
def get_navtabs(soup):
|
||||||
|
"""type3(#nav) 탭그룹의 [(label, panel_id)] 반환 (가장 큰 그룹)."""
|
||||||
|
best = []
|
||||||
|
for ul in soup.find_all('ul'):
|
||||||
|
cls = ' '.join(ul.get('class') or []).lower()
|
||||||
|
if 'tab-ul' not in cls:
|
||||||
|
continue
|
||||||
|
tabs = []
|
||||||
|
for a in ul.find_all('a'):
|
||||||
|
h = (a.get('href') or '').strip()
|
||||||
|
t = a.get_text(strip=True)
|
||||||
|
if h.startswith('#') and len(h) > 1 and t:
|
||||||
|
tabs.append((t, h[1:]))
|
||||||
|
if len(tabs) >= 2 and len(tabs) > len(best):
|
||||||
|
best = tabs
|
||||||
|
return best
|
||||||
|
|
||||||
|
|
||||||
|
def board_signals(panel):
|
||||||
|
if panel is None:
|
||||||
|
return []
|
||||||
|
sig = []
|
||||||
|
# 목록 테이블: 행 3+ 이고 링크 포함
|
||||||
|
for tb in panel.find_all('table'):
|
||||||
|
rows = tb.find_all('tr')
|
||||||
|
links = tb.find_all('a', href=True)
|
||||||
|
if len(rows) >= 3 and len(links) >= 3:
|
||||||
|
sig.append('목록table')
|
||||||
|
break
|
||||||
|
if panel.select('.paging,.pagination,.board_paging,.bbs_paging'):
|
||||||
|
sig.append('페이징')
|
||||||
|
if TOTAL.search(panel.get_text(' ', strip=True)):
|
||||||
|
sig.append('총건수')
|
||||||
|
if any(DETAIL.search(a['href']) for a in panel.find_all('a', href=True)):
|
||||||
|
sig.append('상세링크')
|
||||||
|
if panel.find('iframe'):
|
||||||
|
sig.append('iframe')
|
||||||
|
return sig
|
||||||
|
|
||||||
|
|
||||||
|
def fetch(r, u):
|
||||||
|
try:
|
||||||
|
return r, u, requests.get(u, headers=H, timeout=15, verify=False).content
|
||||||
|
except Exception:
|
||||||
|
return r, u, None
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
ws = openpyxl.load_workbook(XLSX).active
|
||||||
|
targets = [(r, ws.cell(r, 11).value) for r in range(3, ws.max_row + 1)
|
||||||
|
if isinstance(ws.cell(r, 11).value, str) and DOMAIN in ws.cell(r, 11).value]
|
||||||
|
htmls = {}
|
||||||
|
with ThreadPoolExecutor(max_workers=8) as ex:
|
||||||
|
for f in as_completed([ex.submit(fetch, r, u) for r, u in targets]):
|
||||||
|
r, u, h = f.result()
|
||||||
|
if h:
|
||||||
|
htmls[r] = (u, h)
|
||||||
|
|
||||||
|
found_pages = 0
|
||||||
|
board_hits = 0
|
||||||
|
for r in sorted(htmls):
|
||||||
|
u, h = htmls[r]
|
||||||
|
soup = BeautifulSoup(h, 'html.parser')
|
||||||
|
tabs = get_navtabs(soup)
|
||||||
|
if len(tabs) < 2:
|
||||||
|
continue
|
||||||
|
found_pages += 1
|
||||||
|
cat = ws.cell(r, 6).value or ws.cell(r, 5).value or ws.cell(r, 4).value or ''
|
||||||
|
per = []
|
||||||
|
for label, pid in tabs:
|
||||||
|
sig = board_signals(soup.find(id=pid))
|
||||||
|
per.append((label, sig))
|
||||||
|
any_board = any(s for _, s in per)
|
||||||
|
flag = '🟢게시판 발견' if any_board else '— 게시판 없음(정적 콘텐츠)'
|
||||||
|
print(f'\n[행{r}] {cat} 탭{len(tabs)}개 {flag}')
|
||||||
|
print(f' {u}')
|
||||||
|
for label, sig in per:
|
||||||
|
mk = ('게시판? ' + ','.join(sig)) if sig else '정적'
|
||||||
|
print(f' · {label[:30]:<30} {mk}')
|
||||||
|
if sig:
|
||||||
|
board_hits += 1
|
||||||
|
print(f'\n===== 요약: {found_pages}개 페이지 / 게시판 신호 탭 {board_hits}개 =====')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
171
_스크립트/_gongju_recurse.py
Normal file
171
_스크립트/_gongju_recurse.py
Normal file
@ -0,0 +1,171 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""공주시 전용: type1 탭 메뉴를 재귀 크롤링해 각 시트 행의 '진짜 수량'(잎 페이지 수)을
|
||||||
|
끝까지 세서 재계산. 잎 weight = 인페이지 #nav 개수(>=2) 또는 1.
|
||||||
|
하위 subtree에 게시판이 하나라도 있으면 그 행은 '합치지 않음'으로 플래그(M 미변경).
|
||||||
|
|
||||||
|
판별:
|
||||||
|
· type1 메뉴 = ul.tab-ul(단 type3/#nav 제외) 안 같은도메인 .do 링크 집합 중 self 포함하는 것
|
||||||
|
· #nav 개수 = ul.tab-ul 안 href^='#' 탭 수
|
||||||
|
· 게시판 = 목록table/페이징/총N건/상세링크(view.do·mode=V·nttId·BBSMSTR)
|
||||||
|
|
||||||
|
사용: python -X utf8 _gongju_recurse.py (DRY 리포트)
|
||||||
|
python -X utf8 _gongju_recurse.py --write (M 갱신)
|
||||||
|
"""
|
||||||
|
import re, sys, shutil, warnings
|
||||||
|
from urllib.parse import urljoin, urlsplit
|
||||||
|
from concurrent.futures import ThreadPoolExecutor, as_completed
|
||||||
|
import openpyxl, requests
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
XLSX = r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\2.공주시\충청남도_공주시.xlsx'
|
||||||
|
DOMAIN = 'gongju.go.kr'
|
||||||
|
H = {'User-Agent': 'Mozilla/5.0 Chrome/120 Safari/537.36'}
|
||||||
|
DETAIL = re.compile(r'(view\.do|mode=V|nttId|BBSMSTR|selectBoard|selectBbs)', re.I)
|
||||||
|
TOTAL = re.compile(r'총\s*[\d,]+\s*건')
|
||||||
|
|
||||||
|
S = requests.Session(); S.headers.update(H)
|
||||||
|
CACHE = {} # url -> (menu_set, menu_list, nav, is_board)
|
||||||
|
|
||||||
|
|
||||||
|
def norm(u):
|
||||||
|
s = urlsplit(u); return urljoin('http://x/', s.path).split('//', 1)[-1].rstrip('/').lower()
|
||||||
|
|
||||||
|
|
||||||
|
def analyze(url):
|
||||||
|
try:
|
||||||
|
html = S.get(url, timeout=15, verify=False).content
|
||||||
|
except Exception:
|
||||||
|
return (frozenset(), [], 0, False)
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
self_n = norm(url)
|
||||||
|
nav = 0
|
||||||
|
type1_groups = []
|
||||||
|
for ul in soup.find_all('ul'):
|
||||||
|
cls = ' '.join(ul.get('class') or []).lower()
|
||||||
|
if 'tab-ul' not in cls:
|
||||||
|
continue
|
||||||
|
do_links = []
|
||||||
|
navc = 0
|
||||||
|
for a in ul.find_all('a'):
|
||||||
|
h = (a.get('href') or '').strip()
|
||||||
|
t = a.get_text(strip=True)
|
||||||
|
if not t:
|
||||||
|
continue
|
||||||
|
if h.startswith('#'):
|
||||||
|
navc += 1
|
||||||
|
elif h and not h.startswith('javascript:'):
|
||||||
|
au = urljoin(url, h)
|
||||||
|
if DOMAIN in au and au.lower().endswith('.do'):
|
||||||
|
do_links.append((t, au))
|
||||||
|
if navc >= 2:
|
||||||
|
nav = max(nav, navc)
|
||||||
|
if len(do_links) >= 2:
|
||||||
|
type1_groups.append(do_links)
|
||||||
|
# self 포함하는 type1 그룹 우선, 없으면 최대
|
||||||
|
menu = []
|
||||||
|
for g in type1_groups:
|
||||||
|
if any(norm(u) == self_n for _, u in g):
|
||||||
|
menu = g; break
|
||||||
|
if not menu and type1_groups:
|
||||||
|
menu = max(type1_groups, key=len)
|
||||||
|
menu_set = frozenset(norm(u) for _, u in menu)
|
||||||
|
# 게시판 판정
|
||||||
|
body = soup.select_one('#txt') or soup
|
||||||
|
txt = body.get_text(' ', strip=True)
|
||||||
|
is_board = bool(TOTAL.search(txt)) or bool(body.select('.paging,.pagination,.board_paging')) \
|
||||||
|
or any(DETAIL.search(a['href']) for a in body.find_all('a', href=True))
|
||||||
|
return (menu_set, menu, nav, is_board)
|
||||||
|
|
||||||
|
|
||||||
|
def get(url):
|
||||||
|
n = norm(url)
|
||||||
|
if n not in CACHE:
|
||||||
|
CACHE[n] = analyze(url)
|
||||||
|
return CACHE[n]
|
||||||
|
|
||||||
|
|
||||||
|
def crawl(seed_urls):
|
||||||
|
seen = set()
|
||||||
|
queue = list(seed_urls)
|
||||||
|
while queue:
|
||||||
|
batch = [u for u in queue if norm(u) not in seen]
|
||||||
|
for u in batch:
|
||||||
|
seen.add(norm(u))
|
||||||
|
queue = []
|
||||||
|
with ThreadPoolExecutor(max_workers=8) as ex:
|
||||||
|
futs = {ex.submit(get, u): u for u in batch}
|
||||||
|
for f in as_completed(futs):
|
||||||
|
_ms, menu, _nav, _b = f.result()
|
||||||
|
for _, cu in menu:
|
||||||
|
if norm(cu) not in seen:
|
||||||
|
queue.append(cu)
|
||||||
|
|
||||||
|
|
||||||
|
def weight(url):
|
||||||
|
_ms, _menu, nav, _b = get(url)
|
||||||
|
return nav if nav >= 2 else 1
|
||||||
|
|
||||||
|
|
||||||
|
def expand(url, parent_set, path):
|
||||||
|
n = norm(url)
|
||||||
|
if n in path:
|
||||||
|
return 1, False
|
||||||
|
ms, menu, nav, is_board = get(url)
|
||||||
|
if not menu or ms == parent_set:
|
||||||
|
return (nav if nav >= 2 else 1), is_board
|
||||||
|
tot = 0; board = is_board
|
||||||
|
for _, cu in menu:
|
||||||
|
w, b = expand(cu, ms, path | {n})
|
||||||
|
tot += w; board = board or b
|
||||||
|
return tot, board
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
write = '--write' in sys.argv
|
||||||
|
wb = openpyxl.load_workbook(XLSX); ws = wb.active
|
||||||
|
rows = []
|
||||||
|
seeds = []
|
||||||
|
for r in range(3, ws.max_row + 1):
|
||||||
|
u = ws.cell(r, 11).value
|
||||||
|
if isinstance(u, str) and DOMAIN in u and u.lower().endswith('.do'):
|
||||||
|
rows.append((r, u, ws.cell(r, 12).value, ws.cell(r, 13).value,
|
||||||
|
ws.cell(r, 6).value or ws.cell(r, 5).value or ws.cell(r, 4).value or ''))
|
||||||
|
seeds.append(u)
|
||||||
|
print(f'시트 .do 페이지행: {len(rows)} — 재귀 크롤 시작...')
|
||||||
|
crawl(seeds)
|
||||||
|
print(f'크롤한 고유 페이지: {len(CACHE)}\n')
|
||||||
|
|
||||||
|
changes = []; boards = []
|
||||||
|
for r, u, L, M, cat in rows:
|
||||||
|
if L != '페이지':
|
||||||
|
continue
|
||||||
|
newM, hasboard = expand(u, frozenset({norm(u)}), set())
|
||||||
|
if hasboard:
|
||||||
|
boards.append((r, cat, u, M, newM))
|
||||||
|
continue
|
||||||
|
if str(newM) != str(M):
|
||||||
|
changes.append((r, cat, M, newM, u))
|
||||||
|
|
||||||
|
changes.sort(key=lambda x: -(x[3] - (x[2] or 0)))
|
||||||
|
print(f'=== 수량 변경(증가/감소) 대상: {len(changes)}행 ===')
|
||||||
|
for r, cat, oldM, newM, u in changes:
|
||||||
|
print(f' 행{r} {cat[:24]:<24} M {oldM} → {newM} {u}')
|
||||||
|
if boards:
|
||||||
|
print(f'\n=== ⚠ 하위에 게시판 있어 합치지 않음(M 보류): {len(boards)}행 ===')
|
||||||
|
for r, cat, u, M, nm in boards:
|
||||||
|
print(f' 행{r} {cat[:24]:<24} (현 M={M}, 재귀={nm}) {u}')
|
||||||
|
|
||||||
|
if write and changes:
|
||||||
|
bak = XLSX.replace('.xlsx', '_backup_recurse전.xlsx')
|
||||||
|
shutil.copy(XLSX, bak)
|
||||||
|
for r, cat, oldM, newM, u in changes:
|
||||||
|
ws.cell(r, 13).value = newM
|
||||||
|
wb.save(XLSX)
|
||||||
|
print(f'\n저장 완료({len(changes)}행 M갱신). 백업: {bak}')
|
||||||
|
elif not write:
|
||||||
|
print('\n(DRY — 적용하려면 --write)')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
164
_스크립트/_gongju_split_all.py
Normal file
164
_스크립트/_gongju_split_all.py
Normal file
@ -0,0 +1,164 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""335~443행 일반 다단계 탭 분해. _split_plan.json 의 29개 잡을 적용.
|
||||||
|
- 각 잡: 소스행의 가장 깊은 카테고리열(base)을 찾아, 자식을 base+1, 손자를 base+2에 배치
|
||||||
|
- 잡은 행번호 내림차순(아래→위)으로 처리해 미처리 잡 인덱스 불변
|
||||||
|
- 검증된 시프트 로직(값+_style+hyperlink, 병합 +delta 시프트, base열 신규 병합)
|
||||||
|
- 순번(B)은 마지막에 305~끝 일괄 재번호
|
||||||
|
--write 로 적용.
|
||||||
|
"""
|
||||||
|
import sys, json, shutil
|
||||||
|
from copy import copy
|
||||||
|
import openpyxl
|
||||||
|
from openpyxl.utils.cell import range_boundaries
|
||||||
|
|
||||||
|
XLSX = r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\2.공주시\충청남도_공주시.xlsx'
|
||||||
|
PLAN = r'D:\01.프로젝트\DB수집\_스크립트\_split_plan.json'
|
||||||
|
B_START_ROW = 305
|
||||||
|
B_START_VAL = 311
|
||||||
|
CATCOLS = [4, 5, 6, 7, 8, 9, 10] # D..J
|
||||||
|
|
||||||
|
|
||||||
|
def resolve_path(ws, row):
|
||||||
|
"""소스행의 카테고리 경로(병합 반영). {col: value}"""
|
||||||
|
path = {}
|
||||||
|
for c in CATCOLS:
|
||||||
|
v = ws.cell(row, c).value
|
||||||
|
if v is None:
|
||||||
|
# 병합 top-left 찾기
|
||||||
|
for mr in ws.merged_cells.ranges:
|
||||||
|
if mr.min_col <= c <= mr.max_col and mr.min_row <= row <= mr.max_row:
|
||||||
|
v = ws.cell(mr.min_row, mr.min_col).value
|
||||||
|
break
|
||||||
|
if v is not None:
|
||||||
|
path[c] = v
|
||||||
|
return path
|
||||||
|
|
||||||
|
|
||||||
|
def flatten_leaves(job):
|
||||||
|
"""잡 children 트리 -> 잎 리스트 [(lvl1name, lvl2name|None, url, L, M)]"""
|
||||||
|
out = []
|
||||||
|
for ch in job['children']:
|
||||||
|
if ch.get('children'):
|
||||||
|
for gc in ch['children']:
|
||||||
|
out.append((ch['name'], gc['name'], gc['url'], gc.get('L', '페이지'), gc.get('M', 1)))
|
||||||
|
else:
|
||||||
|
out.append((ch['name'], None, ch['url'], ch.get('L', '페이지'), ch.get('M', 1)))
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def apply_job(ws, source, leaves, tmpl, last_data):
|
||||||
|
N = len(leaves)
|
||||||
|
delta = N - 1
|
||||||
|
base = max(resolve_path(ws, source).keys()) # 가장 깊은 카테고리열
|
||||||
|
base_val = ws.cell(source, base).value
|
||||||
|
if base_val is None:
|
||||||
|
p = resolve_path(ws, source); base_val = p[base]
|
||||||
|
|
||||||
|
# 1) 병합 해제(전체) — 시프트 위해
|
||||||
|
old_merges = [str(mr) for mr in list(ws.merged_cells.ranges)]
|
||||||
|
for mr in old_merges:
|
||||||
|
ws.unmerge_cells(mr)
|
||||||
|
|
||||||
|
# 2) 시프트 source+1..last_data -> +delta (아래에서 위로)
|
||||||
|
if delta > 0:
|
||||||
|
for sr in range(last_data, source, -1):
|
||||||
|
dr = sr + delta
|
||||||
|
for c in range(1, 28):
|
||||||
|
s = ws.cell(sr, c); d = ws.cell(dr, c)
|
||||||
|
d.value = s.value
|
||||||
|
if s.has_style:
|
||||||
|
d._style = copy(s._style)
|
||||||
|
if s.hyperlink is not None:
|
||||||
|
d.hyperlink = copy(s.hyperlink); d.hyperlink.ref = d.coordinate
|
||||||
|
s.hyperlink = None
|
||||||
|
else:
|
||||||
|
d.hyperlink = None
|
||||||
|
|
||||||
|
# 3) source..source+N-1 클리어 + 템플릿 스타일
|
||||||
|
for r in range(source, source + N):
|
||||||
|
for c in range(1, 28):
|
||||||
|
cell = ws.cell(r, c)
|
||||||
|
cell.value = None; cell.hyperlink = None
|
||||||
|
if c in tmpl:
|
||||||
|
cell._style = copy(tmpl[c])
|
||||||
|
|
||||||
|
# 4) 잎 기입
|
||||||
|
l1col, l2col = base + 1, base + 2
|
||||||
|
prev_l1 = None
|
||||||
|
for i, (l1, l2, url, L, M) in enumerate(leaves):
|
||||||
|
r = source + i
|
||||||
|
if l1 != prev_l1:
|
||||||
|
ws.cell(r, l1col).value = l1
|
||||||
|
prev_l1 = l1
|
||||||
|
if l2:
|
||||||
|
ws.cell(r, l2col).value = l2
|
||||||
|
kc = ws.cell(r, 11); kc.value = url; kc.hyperlink = url
|
||||||
|
ws.cell(r, 12).value = L
|
||||||
|
ws.cell(r, 13).value = M
|
||||||
|
ws.cell(r, 14).value = '어문'
|
||||||
|
ws.cell(r, 15).value = '미부착'
|
||||||
|
# base 값 최상단
|
||||||
|
ws.cell(source, base).value = base_val
|
||||||
|
|
||||||
|
# 5) 병합 재생성: 기존(>=source+1행 +delta) + 신규
|
||||||
|
def shift(rr):
|
||||||
|
return rr + delta if rr >= source + 1 else rr
|
||||||
|
for rng in old_merges:
|
||||||
|
c1, r1, c2, r2 = range_boundaries(rng)
|
||||||
|
ws.merge_cells(start_row=shift(r1), start_column=c1, end_row=shift(r2), end_column=c2)
|
||||||
|
# base열 신규 병합(잡 전체)
|
||||||
|
if N > 1:
|
||||||
|
ws.merge_cells(start_row=source, start_column=base, end_row=source + N - 1, end_column=base)
|
||||||
|
# l1col 손자그룹 병합(같은 l1 연속 run)
|
||||||
|
i = 0
|
||||||
|
while i < N:
|
||||||
|
j = i
|
||||||
|
while j + 1 < N and leaves[j + 1][0] == leaves[i][0]:
|
||||||
|
j += 1
|
||||||
|
if j > i: # 2개 이상 -> 병합
|
||||||
|
ws.merge_cells(start_row=source + i, start_column=l1col, end_row=source + j, end_column=l1col)
|
||||||
|
i = j + 1
|
||||||
|
|
||||||
|
return last_data + delta
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
write = '--write' in sys.argv
|
||||||
|
jobs = json.load(open(PLAN, encoding='utf-8'))
|
||||||
|
jobs.sort(key=lambda j: j['row'], reverse=True) # 내림차순(아래→위)
|
||||||
|
|
||||||
|
wb = openpyxl.load_workbook(XLSX)
|
||||||
|
ws = wb.active
|
||||||
|
# 템플릿 스타일(클린 데이터행 305)
|
||||||
|
tmpl = {c: copy(ws.cell(305, c)._style) for c in range(1, 28) if ws.cell(305, c).has_style}
|
||||||
|
|
||||||
|
# 현재 마지막 데이터행
|
||||||
|
last_data = 3
|
||||||
|
for r in range(3, ws.max_row + 1):
|
||||||
|
if ws.cell(r, 11).value:
|
||||||
|
last_data = r
|
||||||
|
|
||||||
|
print('잡 %d개, 시작 last_data=%d' % (len(jobs), last_data))
|
||||||
|
if not write:
|
||||||
|
for j in jobs:
|
||||||
|
print(' r%d %s -> 잎 %d' % (j['row'], j['label'], len(flatten_leaves(j))))
|
||||||
|
print('(DRY)')
|
||||||
|
return
|
||||||
|
|
||||||
|
bak = XLSX.replace('.xlsx', '_backup_분야별분해전.xlsx')
|
||||||
|
shutil.copy(XLSX, bak)
|
||||||
|
for j in jobs:
|
||||||
|
leaves = flatten_leaves(j)
|
||||||
|
last_data = apply_job(ws, j['row'], leaves, tmpl, last_data)
|
||||||
|
print(' 적용 r%d %s (+%d) last_data=%d' % (j['row'], j['label'][:24], len(leaves) - 1, last_data))
|
||||||
|
|
||||||
|
# 순번 B 305~last 일괄 재번호
|
||||||
|
for r in range(B_START_ROW, last_data + 1):
|
||||||
|
ws.cell(r, 2).value = B_START_VAL + (r - B_START_ROW)
|
||||||
|
|
||||||
|
wb.save(XLSX)
|
||||||
|
print('저장 완료. 백업:', bak, '| 최종 데이터행', last_data)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
154
_스크립트/_gongju_split_jangae.py
Normal file
154
_스크립트/_gongju_split_jangae.py
Normal file
@ -0,0 +1,154 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""공주시 305행(장애인 단일행, 순번311, sub06_01_02_01.do)을
|
||||||
|
크롤링한 탭 트리(1단계10→잎30)로 분해.
|
||||||
|
- 305행을 30행(305~334)으로 확장, 306~414행은 +29 시프트(335~443)
|
||||||
|
- 순번(B) 305~443 연속 재번호(311..), 305 이전(검수본·삭제갭) 불변
|
||||||
|
- 병합셀 재조정(>=306행 +29) + 장애인 F/G 병합 신규
|
||||||
|
- K열 하이퍼링크 유지/재설정
|
||||||
|
사용: --write 로 적용(미지정 시 계획만 출력)
|
||||||
|
"""
|
||||||
|
import sys, shutil
|
||||||
|
from copy import copy
|
||||||
|
import openpyxl
|
||||||
|
|
||||||
|
XLSX = r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\2.공주시\충청남도_공주시.xlsx'
|
||||||
|
SPLIT_ROW = 305 # 장애인 행
|
||||||
|
N = 30 # 새 행 수
|
||||||
|
SHIFT = N - 1 # 29
|
||||||
|
DATA_LAST = 414 # 현재 마지막 데이터행
|
||||||
|
B_START = 311 # 305행 순번
|
||||||
|
|
||||||
|
BASE = 'https://www.gongju.go.kr'
|
||||||
|
# (G, H, url, L, M) H=None 이면 단일(leaf)
|
||||||
|
ROWS = [
|
||||||
|
('장애인인권헌장', None, '/kr/sub06_01_02_01.do', '페이지', 1),
|
||||||
|
('장애인등록안내', '장애인이란', '/kr/sub06_01_02_03_01.do', '페이지', 1),
|
||||||
|
('장애인등록안내', '장애인등록/심사제도', '/kr/sub06_01_02_03_02.do', '페이지', 1),
|
||||||
|
('장애인등록안내', '장애인등록현황', '/kr/sub06_01_02_03_03.do', '페이지', 1),
|
||||||
|
('장애인등록안내', '장애인관련법률', '/kr/sub06_01_02_03_04.do', '페이지', 1),
|
||||||
|
('장애인복지카드안내', None, '/kr/sub06_01_02_04_01.do', '페이지', 1),
|
||||||
|
('장애인생활안정지원', '장애수당지급', '/kr/sub06_01_02_05_01.do', '페이지', 1),
|
||||||
|
('장애인생활안정지원', '장애인의료비지원', '/kr/sub06_01_02_05_02.do', '페이지', 1),
|
||||||
|
('장애인생활안정지원', '자립자금대여', '/kr/sub06_01_02_05_03.do', '페이지', 1),
|
||||||
|
('장애인생활안정지원', '재활보조기구 무료교부', '/kr/sub06_01_02_05_04.do', '페이지', 1),
|
||||||
|
('장애인생활안정지원', '장애아동수당', '/kr/sub06_01_02_05_05.do', '페이지', 1),
|
||||||
|
('장애인생활안정지원', '장애인일자리', '/kr/sub06_01_02_05_06.do', '페이지', 1),
|
||||||
|
('자동차관련시책', '장애인자동차표지 발급', '/kr/sub06_01_02_06_01.do', '페이지', 1),
|
||||||
|
('자동차관련시책', '고속도로통행료 할인', '/kr/sub06_01_02_06_02.do', '페이지', 1),
|
||||||
|
('자동차관련시책', '자동차특별소비세 면제', '/kr/sub06_01_02_06_03.do', '페이지', 1),
|
||||||
|
('세금감면시책', '소득세공제', '/kr/sub06_01_02_07_01.do', '페이지', 1),
|
||||||
|
('세금감면시책', '상속세공제', '/kr/sub06_01_02_07_02.do', '페이지', 1),
|
||||||
|
('각종요금할인', '전화요금할인', '/kr/sub06_01_02_08_01.do', '페이지', 1),
|
||||||
|
('각종요금할인', 'TV수신료면제', '/kr/sub06_01_02_08_02.do', '페이지', 1),
|
||||||
|
('각종요금할인', '이동통신 요금할인', '/kr/sub06_01_02_08_03.do', '페이지', 1),
|
||||||
|
('각종요금할인', '교통요금 할인', '/kr/sub06_01_02_08_04.do', '페이지', 1),
|
||||||
|
('각종요금할인', '공공시설 이용요금 감면', '/kr/sub06_01_02_08_05.do', '페이지', 1),
|
||||||
|
('각종요금할인', '장애인 전기요금 감면', '/kr/sub06_01_02_08_06.do', '페이지', 1),
|
||||||
|
('각종요금할인', '초고속인터넷 요금할인', '/kr/sub06_01_02_08_07.do', '페이지', 1),
|
||||||
|
('장애인전화상담', '장애인전화상담소', '/kr/sub06_01_02_09_01.do', '페이지', 1),
|
||||||
|
('장애인전화상담', '이용안내', '/kr/sub06_01_02_09_02.do', '페이지', 1),
|
||||||
|
('장애인 전동휠체어 급속충전기 설치장소', None, '/kr/sub06_01_02_10.do', '페이지', 1),
|
||||||
|
('한눈에 보는 장애인복지 서비스', '시설소개', '/kr/sub06_01_02_11_01.do', '페이지', 1),
|
||||||
|
('한눈에 보는 장애인복지 서비스', '사업홍보', 'https://www.gongju.go.kr/bbs/BBSMSTR_000000001671/list.do?mno=sub06_01_02_11_02', '게시판', 96),
|
||||||
|
('한눈에 보는 장애인복지 서비스', '채용정보', 'https://www.gongju.go.kr/bbs/BBSMSTR_000000001672/list.do?mno=sub06_01_02_11_03', '게시판', 11),
|
||||||
|
]
|
||||||
|
assert len(ROWS) == N
|
||||||
|
|
||||||
|
|
||||||
|
def full(u):
|
||||||
|
return u if u.startswith('http') else BASE + u
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
write = '--write' in sys.argv
|
||||||
|
wb = openpyxl.load_workbook(XLSX)
|
||||||
|
ws = wb.active
|
||||||
|
|
||||||
|
if write:
|
||||||
|
# 0) 모든 병합 먼저 해제(시프트 시 MergedCell 쓰기불가 방지)
|
||||||
|
old_merges = [str(mr) for mr in list(ws.merged_cells.ranges)]
|
||||||
|
for mr in old_merges:
|
||||||
|
ws.unmerge_cells(mr)
|
||||||
|
|
||||||
|
# 1) 시프트: source 414..306 -> dest +29 (아래에서 위로)
|
||||||
|
for sr in range(DATA_LAST, SPLIT_ROW, -1): # 414..306
|
||||||
|
dr = sr + SHIFT
|
||||||
|
for c in range(1, 28):
|
||||||
|
src = ws.cell(sr, c)
|
||||||
|
dst = ws.cell(dr, c)
|
||||||
|
dst.value = src.value
|
||||||
|
if src.has_style:
|
||||||
|
dst._style = copy(src._style)
|
||||||
|
# hyperlink 이동
|
||||||
|
if src.hyperlink is not None:
|
||||||
|
dst.hyperlink = copy(src.hyperlink)
|
||||||
|
dst.hyperlink.ref = dst.coordinate
|
||||||
|
src.hyperlink = None
|
||||||
|
else:
|
||||||
|
dst.hyperlink = None
|
||||||
|
|
||||||
|
# 2) 305~334 영역 클리어 + 스타일 템플릿(원래 305행) 적용
|
||||||
|
# 원래 305행 스타일은 시프트 안했으므로 그대로 305에 남아있음(템플릿)
|
||||||
|
tmpl = {c: copy(ws.cell(SPLIT_ROW, c)._style) for c in range(1, 28) if ws.cell(SPLIT_ROW, c).has_style}
|
||||||
|
for r in range(SPLIT_ROW, SPLIT_ROW + N):
|
||||||
|
for c in range(1, 28):
|
||||||
|
cell = ws.cell(r, c)
|
||||||
|
cell.value = None
|
||||||
|
cell.hyperlink = None
|
||||||
|
if c in tmpl:
|
||||||
|
cell._style = copy(tmpl[c])
|
||||||
|
|
||||||
|
# 3) 30행 데이터 기입 (G는 그룹 최상단에만)
|
||||||
|
prev_g = None
|
||||||
|
for i, (g, h, u, L, M) in enumerate(ROWS):
|
||||||
|
r = SPLIT_ROW + i
|
||||||
|
if g != prev_g:
|
||||||
|
ws.cell(r, 7).value = g # G 소 (그룹 최상단)
|
||||||
|
prev_g = g
|
||||||
|
if h:
|
||||||
|
ws.cell(r, 8).value = h # H 세부
|
||||||
|
fu = full(u)
|
||||||
|
kc = ws.cell(r, 11)
|
||||||
|
kc.value = fu
|
||||||
|
kc.hyperlink = fu
|
||||||
|
ws.cell(r, 12).value = L # L
|
||||||
|
ws.cell(r, 13).value = M # M
|
||||||
|
ws.cell(r, 14).value = '어문' # N
|
||||||
|
ws.cell(r, 15).value = '미부착' # O
|
||||||
|
# F=장애인 (병합 최상단)
|
||||||
|
ws.cell(SPLIT_ROW, 6).value = '장애인'
|
||||||
|
|
||||||
|
# 4) 순번(B) 305~443 연속 재번호
|
||||||
|
for r in range(SPLIT_ROW, DATA_LAST + SHIFT + 1):
|
||||||
|
ws.cell(r, 2).value = B_START + (r - SPLIT_ROW)
|
||||||
|
|
||||||
|
# 5) 병합셀 재조정 (0단계서 해제한 old_merges를 시프트해 재생성)
|
||||||
|
from openpyxl.utils.cell import range_boundaries
|
||||||
|
def shift(rr):
|
||||||
|
return rr + SHIFT if rr >= SPLIT_ROW + 1 else rr
|
||||||
|
for rng in old_merges:
|
||||||
|
c1, r1, c2, r2 = range_boundaries(rng)
|
||||||
|
nr1, nr2 = shift(r1), shift(r2)
|
||||||
|
ws.merge_cells(start_row=nr1, start_column=c1, end_row=nr2, end_column=c2)
|
||||||
|
# 신규 병합: F305:F334
|
||||||
|
ws.merge_cells(start_row=305, start_column=6, end_row=334, end_column=6)
|
||||||
|
# G 병합(다중 H 그룹)
|
||||||
|
gmerges = [(306, 309), (311, 316), (317, 319), (320, 321), (322, 328), (329, 330), (332, 334)]
|
||||||
|
for a, b in gmerges:
|
||||||
|
ws.merge_cells(start_row=a, start_column=7, end_row=b, end_column=7)
|
||||||
|
|
||||||
|
bak = XLSX.replace('.xlsx', '_backup_장애인분해전.xlsx')
|
||||||
|
shutil.copy(XLSX, bak)
|
||||||
|
wb.save(XLSX)
|
||||||
|
print('저장 완료. 백업:', bak)
|
||||||
|
else:
|
||||||
|
print('=== 계획(DRY) ===')
|
||||||
|
print('305행(장애인) → 30행 확장, 306~414 → +29 시프트')
|
||||||
|
print('순번 305~443 = 311..%d' % (B_START + (DATA_LAST + SHIFT - SPLIT_ROW)))
|
||||||
|
print('F305:F334 병합, G병합 7개')
|
||||||
|
for i, (g, h, u, L, M) in enumerate(ROWS):
|
||||||
|
print(' r%d G=%s H=%s L=%s M=%s' % (SPLIT_ROW + i, g, h or '', L, M))
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
707
_스크립트/_jeonbuk_phase1_all.py
Normal file
707
_스크립트/_jeonbuk_phase1_all.py
Normal file
@ -0,0 +1,707 @@
|
|||||||
|
"""전북특별자치도 14개 시·군 + 제주특별자치도 2개 시 Phase 1 일괄 처리.
|
||||||
|
|
||||||
|
매뉴얼: D:\\01.프로젝트\\DB수집\\사이트맵_수집_매뉴얼.md
|
||||||
|
출력: 각 폴더의 {기관명}.xlsx (D~K열)
|
||||||
|
|
||||||
|
파서 매핑은 _probe_jeonbuk*.py 탐색 결과 기반.
|
||||||
|
"""
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import re
|
||||||
|
import shutil
|
||||||
|
import ssl
|
||||||
|
import sys
|
||||||
|
import warnings
|
||||||
|
from copy import copy
|
||||||
|
from urllib.parse import urljoin
|
||||||
|
|
||||||
|
import openpyxl
|
||||||
|
import requests
|
||||||
|
from requests.adapters import HTTPAdapter
|
||||||
|
from urllib3.util.ssl_ import create_urllib3_context
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
from openpyxl.styles import Alignment, Font
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
ROOT = r'D:\01.프로젝트\DB수집'
|
||||||
|
TEMPLATE = ROOT + r'\자료_취합_예시.xlsx'
|
||||||
|
|
||||||
|
|
||||||
|
class WeakSSLAdapter(HTTPAdapter):
|
||||||
|
def init_poolmanager(self, *args, **kwargs):
|
||||||
|
ctx = create_urllib3_context()
|
||||||
|
ctx.set_ciphers('DEFAULT@SECLEVEL=0')
|
||||||
|
ctx.options |= 0x4
|
||||||
|
ctx.check_hostname = False
|
||||||
|
ctx.verify_mode = ssl.CERT_NONE
|
||||||
|
kwargs['ssl_context'] = ctx
|
||||||
|
return super().init_poolmanager(*args, **kwargs)
|
||||||
|
|
||||||
|
|
||||||
|
def make_session(weak_ssl=False):
|
||||||
|
s = requests.Session()
|
||||||
|
s.headers.update(H)
|
||||||
|
if weak_ssl:
|
||||||
|
s.mount('https://', WeakSSLAdapter())
|
||||||
|
return s
|
||||||
|
|
||||||
|
|
||||||
|
def fetch_html(url, session=None, timeout=25):
|
||||||
|
s = session or requests.Session()
|
||||||
|
if not session:
|
||||||
|
s.headers.update(H)
|
||||||
|
r = s.get(url, timeout=timeout, verify=False, allow_redirects=True)
|
||||||
|
meta = re.search(rb'<meta[^>]*charset=["\']?\s*([\w-]+)', r.content[:4096], re.I)
|
||||||
|
r.encoding = meta.group(1).decode('ascii', errors='ignore') if meta else r.apparent_encoding
|
||||||
|
return r.text
|
||||||
|
|
||||||
|
|
||||||
|
def clean_text(s):
|
||||||
|
return re.sub(r'\s+', ' ', s or '').strip().replace('\xa0', '').lstrip('-').strip()
|
||||||
|
|
||||||
|
|
||||||
|
def extract_href(a):
|
||||||
|
if a is None:
|
||||||
|
return ''
|
||||||
|
href = (a.get('href') or '').strip()
|
||||||
|
if not href or href.startswith('#') or href.lower().startswith('javascript:'):
|
||||||
|
return ''
|
||||||
|
return href
|
||||||
|
|
||||||
|
|
||||||
|
COLS = 'DEFGHIJ'
|
||||||
|
|
||||||
|
|
||||||
|
def rows_from_paths(tmp):
|
||||||
|
"""[{'path':[(text,href)...], 'href':..}] → D~J dict 행 리스트."""
|
||||||
|
rows = []
|
||||||
|
for item in tmp:
|
||||||
|
p = item['path']
|
||||||
|
row = {c: '' for c in COLS}
|
||||||
|
row['href'] = item['href']
|
||||||
|
for i, (t, _) in enumerate(p):
|
||||||
|
row[COLS[i] if i < len(COLS) else 'J'] = t
|
||||||
|
rows.append(row)
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def walk_ul(ul, base_path, out, recursive_li=True):
|
||||||
|
"""ul > li > a (+ 중첩 ul) 재귀. 망가진 마크업(li 안 li)도 허용."""
|
||||||
|
for li in ul.find_all('li', recursive=False):
|
||||||
|
a = li.find('a', recursive=False)
|
||||||
|
if not a:
|
||||||
|
# div>a 형태(전주) 허용
|
||||||
|
d = li.find('div', recursive=False)
|
||||||
|
a = d.find('a', recursive=False) if d else None
|
||||||
|
if not a:
|
||||||
|
continue
|
||||||
|
text = clean_text(a.get_text())
|
||||||
|
href = extract_href(a)
|
||||||
|
path = base_path + [(text, href)]
|
||||||
|
if not text:
|
||||||
|
continue
|
||||||
|
# 자식 ul (정상) 또는 li 직접 중첩(망가진 마크업)
|
||||||
|
child_uls = li.find_all('ul', recursive=False)
|
||||||
|
child_lis = [c for c in li.find_all('li', recursive=False)]
|
||||||
|
if child_uls:
|
||||||
|
out.append({'path': list(path), 'href': href})
|
||||||
|
for cul in child_uls:
|
||||||
|
walk_ul(cul, path, out)
|
||||||
|
elif recursive_li and child_lis:
|
||||||
|
out.append({'path': list(path), 'href': href})
|
||||||
|
for cli in child_lis:
|
||||||
|
# cli를 단일 li로 감싼 가짜 ul처럼 처리
|
||||||
|
sub = a.find_parent() # not used
|
||||||
|
_walk_single_li(cli, path, out)
|
||||||
|
else:
|
||||||
|
out.append({'path': list(path), 'href': href})
|
||||||
|
|
||||||
|
|
||||||
|
def _walk_single_li(li, base_path, out):
|
||||||
|
a = li.find('a', recursive=False)
|
||||||
|
if not a:
|
||||||
|
d = li.find('div', recursive=False)
|
||||||
|
a = d.find('a', recursive=False) if d else None
|
||||||
|
if not a:
|
||||||
|
return
|
||||||
|
text = clean_text(a.get_text())
|
||||||
|
href = extract_href(a)
|
||||||
|
path = base_path + [(text, href)]
|
||||||
|
if not text:
|
||||||
|
return
|
||||||
|
child_uls = li.find_all('ul', recursive=False)
|
||||||
|
child_lis = li.find_all('li', recursive=False)
|
||||||
|
if child_uls:
|
||||||
|
out.append({'path': list(path), 'href': href})
|
||||||
|
for cul in child_uls:
|
||||||
|
walk_ul(cul, path, out)
|
||||||
|
elif child_lis:
|
||||||
|
out.append({'path': list(path), 'href': href})
|
||||||
|
for cli in child_lis:
|
||||||
|
_walk_single_li(cli, path, out)
|
||||||
|
else:
|
||||||
|
out.append({'path': list(path), 'href': href})
|
||||||
|
|
||||||
|
|
||||||
|
# ================================================================
|
||||||
|
# 파서들
|
||||||
|
# ================================================================
|
||||||
|
|
||||||
|
def parse_menu_div(soup, base):
|
||||||
|
"""고창/김제/임실: div.sitemap > div.menuN > h4>a (D) + div > ul > li>a (E) + ul (F) 재귀."""
|
||||||
|
rows = []
|
||||||
|
sm = soup.select_one('div.sitemap')
|
||||||
|
if not sm:
|
||||||
|
return rows
|
||||||
|
for block in sm.find_all('div', recursive=False):
|
||||||
|
h4 = block.find('h4')
|
||||||
|
D = clean_text(h4.get_text()) if h4 else ''
|
||||||
|
if not D:
|
||||||
|
continue
|
||||||
|
inner = block.find('div', recursive=False)
|
||||||
|
ul = inner.find('ul', recursive=False) if inner else block.find('ul', recursive=False)
|
||||||
|
if not ul:
|
||||||
|
rows.append({'D': D, 'href': '', **{c: '' for c in 'EFGHIJ'}})
|
||||||
|
continue
|
||||||
|
tmp = []
|
||||||
|
walk_ul(ul, [], tmp)
|
||||||
|
for r in rows_from_paths(tmp):
|
||||||
|
rows.append({'D': D, 'E': r.get('D', ''), 'F': r.get('E', ''),
|
||||||
|
'G': r.get('F', ''), 'H': r.get('G', ''), 'I': r.get('H', ''),
|
||||||
|
'J': r.get('I', ''), 'href': r.get('href', '')})
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def parse_jeongeup(soup, base):
|
||||||
|
"""정읍: div.sitemap > div.st_mapNN > p.tit>a (D) + ul > li > b>a (E) + ul > li>a (F)."""
|
||||||
|
rows = []
|
||||||
|
sm = soup.select_one('div.sitemap')
|
||||||
|
if not sm:
|
||||||
|
return rows
|
||||||
|
for block in sm.find_all('div', recursive=False):
|
||||||
|
ptit = block.find('p', class_='tit')
|
||||||
|
D = clean_text(ptit.get_text()) if ptit else ''
|
||||||
|
if not D:
|
||||||
|
continue
|
||||||
|
ul = block.find('ul', recursive=False)
|
||||||
|
if not ul:
|
||||||
|
rows.append({'D': D, 'href': '', **{c: '' for c in 'EFGHIJ'}})
|
||||||
|
continue
|
||||||
|
for li in ul.find_all('li', recursive=False):
|
||||||
|
b = li.find('b', recursive=False)
|
||||||
|
b_a = b.find('a') if b else li.find('a', recursive=False)
|
||||||
|
E = clean_text(b_a.get_text()) if b_a else ''
|
||||||
|
E_href = extract_href(b_a) if b_a else ''
|
||||||
|
sub = li.find('ul', recursive=False)
|
||||||
|
if not sub:
|
||||||
|
rows.append({'D': D, 'E': E, 'href': E_href, **{c: '' for c in 'FGHIJ'}})
|
||||||
|
continue
|
||||||
|
rows.append({'D': D, 'E': E, 'href': E_href, **{c: '' for c in 'FGHIJ'}})
|
||||||
|
tmp = []
|
||||||
|
walk_ul(sub, [], tmp)
|
||||||
|
for r in rows_from_paths(tmp):
|
||||||
|
rows.append({'D': D, 'E': E, 'F': r.get('D', ''), 'G': r.get('E', ''),
|
||||||
|
'H': r.get('F', ''), 'I': r.get('G', ''), 'J': r.get('H', ''),
|
||||||
|
'href': r.get('href', '')})
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def parse_group_sitemap(soup, base):
|
||||||
|
"""남원/익산: div.sitemap_group (여러개) > h4.title (D) + ul.sitemap_2dep > li>a (E) + ul (F) 재귀."""
|
||||||
|
rows = []
|
||||||
|
groups = soup.select('div.sitemap_group')
|
||||||
|
for g in groups:
|
||||||
|
h4 = g.find('h4', class_='title') or g.find('h4')
|
||||||
|
D = clean_text(h4.get_text()) if h4 else ''
|
||||||
|
if not D:
|
||||||
|
continue
|
||||||
|
ul = g.find('ul', class_='sitemap_2dep') or g.find('ul', recursive=False)
|
||||||
|
if not ul:
|
||||||
|
rows.append({'D': D, 'href': '', **{c: '' for c in 'EFGHIJ'}})
|
||||||
|
continue
|
||||||
|
tmp = []
|
||||||
|
walk_ul(ul, [], tmp)
|
||||||
|
for r in rows_from_paths(tmp):
|
||||||
|
rows.append({'D': D, 'E': r.get('D', ''), 'F': r.get('E', ''),
|
||||||
|
'G': r.get('F', ''), 'H': r.get('G', ''), 'I': r.get('H', ''),
|
||||||
|
'J': r.get('I', ''), 'href': r.get('href', '')})
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def parse_namwon(soup, base):
|
||||||
|
"""남원: div.sitemap > (h4 (D) + ul (E/F 재귀)) 형제 반복."""
|
||||||
|
rows = []
|
||||||
|
sm = soup.select_one('div.sitemap')
|
||||||
|
if not sm:
|
||||||
|
return rows
|
||||||
|
curD = ''
|
||||||
|
for child in sm.find_all(['h4', 'ul'], recursive=False):
|
||||||
|
if child.name == 'h4':
|
||||||
|
curD = clean_text(child.get_text())
|
||||||
|
elif child.name == 'ul' and curD:
|
||||||
|
tmp = []
|
||||||
|
walk_ul(child, [], tmp)
|
||||||
|
for r in rows_from_paths(tmp):
|
||||||
|
rows.append({'D': curD, 'E': r.get('D', ''), 'F': r.get('E', ''),
|
||||||
|
'G': r.get('F', ''), 'H': r.get('G', ''), 'I': r.get('H', ''),
|
||||||
|
'J': r.get('I', ''), 'href': r.get('href', '')})
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def parse_muju(soup, base):
|
||||||
|
"""무주: div#sitemap > div.sitemapN > (h4.sNN>a (D) + ul>li>a (E)) 형제 반복."""
|
||||||
|
rows = []
|
||||||
|
cont = soup.select_one('div#sitemap')
|
||||||
|
if not cont:
|
||||||
|
return rows
|
||||||
|
for box in cont.find_all('div', recursive=False):
|
||||||
|
curD = ''
|
||||||
|
for child in box.find_all(['h4', 'ul'], recursive=False):
|
||||||
|
if child.name == 'h4':
|
||||||
|
a = child.find('a')
|
||||||
|
curD = clean_text(a.get_text() if a else child.get_text())
|
||||||
|
elif child.name == 'ul' and curD:
|
||||||
|
for li in child.find_all('li', recursive=False):
|
||||||
|
a = li.find('a', recursive=False)
|
||||||
|
if not a:
|
||||||
|
continue
|
||||||
|
E = clean_text(a.get_text())
|
||||||
|
rows.append({'D': curD, 'E': E, 'href': extract_href(a),
|
||||||
|
**{c: '' for c in 'FGHIJ'}})
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def parse_buan(soup, base):
|
||||||
|
"""부안: nav#onmenu > ul > li > div.depth_box > div.depth_boxcon > strong (D) + ul>li>a (E) + ul (F)."""
|
||||||
|
rows = []
|
||||||
|
nav = soup.select_one('nav#onmenu')
|
||||||
|
if not nav:
|
||||||
|
return rows
|
||||||
|
top = nav.find('ul')
|
||||||
|
if not top:
|
||||||
|
return rows
|
||||||
|
for li in top.find_all('li', recursive=False):
|
||||||
|
box = li.find('div', class_='depth_boxcon')
|
||||||
|
if not box:
|
||||||
|
continue
|
||||||
|
strong = box.find('strong')
|
||||||
|
a0 = li.find('a', recursive=False)
|
||||||
|
D = clean_text(strong.get_text()) if strong else clean_text(a0.get_text() if a0 else '')
|
||||||
|
if not D:
|
||||||
|
continue
|
||||||
|
ul = box.find('ul')
|
||||||
|
if not ul:
|
||||||
|
continue
|
||||||
|
tmp = []
|
||||||
|
walk_ul(ul, [], tmp)
|
||||||
|
for r in rows_from_paths(tmp):
|
||||||
|
rows.append({'D': D, 'E': r.get('D', ''), 'F': r.get('E', ''),
|
||||||
|
'G': r.get('F', ''), 'H': r.get('G', ''), 'I': r.get('H', ''),
|
||||||
|
'J': r.get('I', ''), 'href': r.get('href', '')})
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def parse_sunchang(soup, base):
|
||||||
|
"""순창: ul.gnb > li > a (D) + div.box ul.gnb_2dep > li>a (E) + ul.gnb_3dep (F) + ul.gnb_4dep (G)."""
|
||||||
|
rows = []
|
||||||
|
gnb = soup.select_one('ul.gnb')
|
||||||
|
if not gnb:
|
||||||
|
return rows
|
||||||
|
for li in gnb.find_all('li', recursive=False):
|
||||||
|
a0 = li.find('a', recursive=False)
|
||||||
|
D = clean_text(a0.get_text()) if a0 else ''
|
||||||
|
if not D:
|
||||||
|
continue
|
||||||
|
ul2 = li.find('ul', class_='gnb_2dep')
|
||||||
|
if not ul2:
|
||||||
|
rows.append({'D': D, 'href': extract_href(a0), **{c: '' for c in 'EFGHIJ'}})
|
||||||
|
continue
|
||||||
|
for li2 in ul2.find_all('li', recursive=False):
|
||||||
|
a2 = li2.find('a', recursive=False)
|
||||||
|
if not a2:
|
||||||
|
continue
|
||||||
|
E = clean_text(a2.get_text())
|
||||||
|
E_href = extract_href(a2)
|
||||||
|
ul3 = li2.find('ul', class_='gnb_3dep')
|
||||||
|
if not ul3:
|
||||||
|
rows.append({'D': D, 'E': E, 'href': E_href, **{c: '' for c in 'FGHIJ'}})
|
||||||
|
continue
|
||||||
|
rows.append({'D': D, 'E': E, 'href': E_href, **{c: '' for c in 'FGHIJ'}})
|
||||||
|
for li3 in ul3.find_all('li', recursive=False):
|
||||||
|
a3 = li3.find('a', recursive=False)
|
||||||
|
if not a3:
|
||||||
|
continue
|
||||||
|
F = clean_text(a3.get_text())
|
||||||
|
F_href = extract_href(a3)
|
||||||
|
ul4 = li3.find('ul', class_='gnb_4dep')
|
||||||
|
if not ul4:
|
||||||
|
rows.append({'D': D, 'E': E, 'F': F, 'href': F_href, **{c: '' for c in 'GHIJ'}})
|
||||||
|
continue
|
||||||
|
rows.append({'D': D, 'E': E, 'F': F, 'href': F_href, **{c: '' for c in 'GHIJ'}})
|
||||||
|
for li4 in ul4.find_all('li', recursive=False):
|
||||||
|
a4 = li4.find('a', recursive=False)
|
||||||
|
if not a4:
|
||||||
|
continue
|
||||||
|
rows.append({'D': D, 'E': E, 'F': F, 'G': clean_text(a4.get_text()),
|
||||||
|
'href': extract_href(a4), **{c: '' for c in 'HIJ'}})
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def parse_jangsu(soup, base):
|
||||||
|
"""장수: div.sitemap > ul.siteMapList > li.sml_1depth > a.sml_1depthBtn (D) + ul.sml_2depthList > li>a (E) + ul 재귀."""
|
||||||
|
rows = []
|
||||||
|
ul = soup.select_one('ul.siteMapList')
|
||||||
|
if not ul:
|
||||||
|
return rows
|
||||||
|
for li in ul.find_all('li', recursive=False):
|
||||||
|
a0 = li.find('a', recursive=False)
|
||||||
|
D = clean_text(a0.get_text()) if a0 else ''
|
||||||
|
if not D:
|
||||||
|
continue
|
||||||
|
ul2 = li.find('ul', recursive=False)
|
||||||
|
if not ul2:
|
||||||
|
rows.append({'D': D, 'href': extract_href(a0), **{c: '' for c in 'EFGHIJ'}})
|
||||||
|
continue
|
||||||
|
tmp = []
|
||||||
|
walk_ul(ul2, [], tmp)
|
||||||
|
for r in rows_from_paths(tmp):
|
||||||
|
rows.append({'D': D, 'E': r.get('D', ''), 'F': r.get('E', ''),
|
||||||
|
'G': r.get('F', ''), 'H': r.get('G', ''), 'I': r.get('H', ''),
|
||||||
|
'J': r.get('I', ''), 'href': r.get('href', '')})
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def parse_jeonju(soup, base):
|
||||||
|
"""전주: div.sitemap_Warp (여러개) > h4.title_h4 (D) + ul > li > div>a (E) + ul/li (F) 망가진 마크업."""
|
||||||
|
rows = []
|
||||||
|
warps = soup.select('div.sitemap_Warp')
|
||||||
|
for w in warps:
|
||||||
|
h4 = w.find('h4', class_='title_h4') or w.find('h4')
|
||||||
|
D = clean_text(h4.get_text()) if h4 else ''
|
||||||
|
if not D:
|
||||||
|
continue
|
||||||
|
ul = w.find('ul', recursive=False)
|
||||||
|
if not ul:
|
||||||
|
continue
|
||||||
|
tmp = []
|
||||||
|
walk_ul(ul, [], tmp)
|
||||||
|
for r in rows_from_paths(tmp):
|
||||||
|
rows.append({'D': D, 'E': r.get('D', ''), 'F': r.get('E', ''),
|
||||||
|
'G': r.get('F', ''), 'H': r.get('G', ''), 'I': r.get('H', ''),
|
||||||
|
'J': r.get('I', ''), 'href': r.get('href', '')})
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def parse_jinan(soup, base):
|
||||||
|
"""진안: div.sitemap > dl > dt (D) + dd > ul > li>a (E) + ul (F) 재귀."""
|
||||||
|
rows = []
|
||||||
|
sm = soup.select_one('div.sitemap')
|
||||||
|
if not sm:
|
||||||
|
return rows
|
||||||
|
for dl in sm.find_all('dl', recursive=False):
|
||||||
|
dt = dl.find('dt')
|
||||||
|
D = clean_text(dt.get_text()) if dt else ''
|
||||||
|
if not D:
|
||||||
|
continue
|
||||||
|
dd = dl.find('dd')
|
||||||
|
ul = dd.find('ul', recursive=False) if dd else None
|
||||||
|
if not ul:
|
||||||
|
rows.append({'D': D, 'href': '', **{c: '' for c in 'EFGHIJ'}})
|
||||||
|
continue
|
||||||
|
tmp = []
|
||||||
|
walk_ul(ul, [], tmp)
|
||||||
|
for r in rows_from_paths(tmp):
|
||||||
|
rows.append({'D': D, 'E': r.get('D', ''), 'F': r.get('E', ''),
|
||||||
|
'G': r.get('F', ''), 'H': r.get('G', ''), 'I': r.get('H', ''),
|
||||||
|
'J': r.get('I', ''), 'href': r.get('href', '')})
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def parse_box_heading(soup, base):
|
||||||
|
"""서귀포/제주시: div.sitemap > div.sitemapBox|sitemap_menu > h3|h4 (D) + ul > li>a (E)."""
|
||||||
|
rows = []
|
||||||
|
sm = soup.select_one('div.sitemap')
|
||||||
|
if not sm:
|
||||||
|
return rows
|
||||||
|
for box in sm.find_all('div', recursive=False):
|
||||||
|
h = box.find(['h3', 'h4'])
|
||||||
|
D = clean_text(h.get_text()) if h else ''
|
||||||
|
if not D:
|
||||||
|
continue
|
||||||
|
ul = box.find('ul')
|
||||||
|
if not ul:
|
||||||
|
rows.append({'D': D, 'href': '', **{c: '' for c in 'EFGHIJ'}})
|
||||||
|
continue
|
||||||
|
tmp = []
|
||||||
|
walk_ul(ul, [], tmp)
|
||||||
|
for r in rows_from_paths(tmp):
|
||||||
|
rows.append({'D': D, 'E': r.get('D', ''), 'F': r.get('E', ''),
|
||||||
|
'G': r.get('F', ''), 'H': r.get('G', ''), 'I': r.get('H', ''),
|
||||||
|
'J': r.get('I', ''), 'href': r.get('href', '')})
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def parse_wanju_json(soup, base):
|
||||||
|
"""완주: Playwright로 추출한 _wanju_menu.json 로드."""
|
||||||
|
rows = []
|
||||||
|
data = json.load(open(os.path.join(os.path.dirname(os.path.abspath(__file__)), '_wanju_menu.json'), encoding='utf-8'))
|
||||||
|
for r in data:
|
||||||
|
rows.append({'D': r.get('D', ''), 'E': r.get('E', ''), 'F': r.get('F', ''),
|
||||||
|
'href': r.get('href', ''), 'G': '', 'H': '', 'I': '', 'J': ''})
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
# ================================================================
|
||||||
|
# 사이트 설정
|
||||||
|
# ================================================================
|
||||||
|
def jb(i, name, folder, base, sitemap, parser, domain, weak=False, fetch_kind='html'):
|
||||||
|
return {'idx': i, 'name': name, 'base': base, 'sitemap': sitemap,
|
||||||
|
'sheet': f'{i:02d}_{name}', 'parser': parser, 'domain': domain,
|
||||||
|
'folder': fr'{ROOT}\작업파일\광역_사이트맵\전북특별자치도\{i}.{name}', 'weak_ssl': weak,
|
||||||
|
'fetch_kind': fetch_kind}
|
||||||
|
|
||||||
|
|
||||||
|
def jj(i, name, base, sitemap, parser, domain):
|
||||||
|
return {'idx': i, 'name': name, 'base': base, 'sitemap': sitemap,
|
||||||
|
'sheet': f'{i:02d}_{name}', 'parser': parser, 'domain': domain,
|
||||||
|
'folder': fr'{ROOT}\작업파일\광역_사이트맵\제주특별자치도\{i}.{name}', 'weak_ssl': False,
|
||||||
|
'fetch_kind': 'html'}
|
||||||
|
|
||||||
|
|
||||||
|
SITES = [
|
||||||
|
jb(1, '고창군', '', 'https://www.gochang.go.kr',
|
||||||
|
'https://www.gochang.go.kr/index.gochang?menuCd=DOM_000000106006000000', parse_menu_div, 'gochang.go.kr'),
|
||||||
|
jb(2, '군산시', '', 'https://www.gunsan.go.kr',
|
||||||
|
'https://www.gunsan.go.kr/main', parse_gunsan if False else None, 'gunsan.go.kr'),
|
||||||
|
jb(3, '김제시', '', 'https://www.gimje.go.kr',
|
||||||
|
'https://www.gimje.go.kr/index.gimje?menuCd=DOM_000000107002000000', parse_menu_div, 'gimje.go.kr'),
|
||||||
|
jb(4, '남원시', '', 'https://www.namwon.go.kr',
|
||||||
|
'https://www.namwon.go.kr/index.do?menuUid=ff8080818f2717db018f277767500088', parse_namwon, 'namwon.go.kr'),
|
||||||
|
jb(5, '무주군', '', 'https://www.muju.go.kr',
|
||||||
|
'https://www.muju.go.kr/index.9is?contentUid=ff8080816db80238016dc8e98fa10ef6', parse_muju, 'muju.go.kr'),
|
||||||
|
jb(6, '부안군', '', 'https://www.buan.go.kr',
|
||||||
|
'https://www.buan.go.kr/index.buan?contentsSid=1', parse_buan, 'buan.go.kr'),
|
||||||
|
jb(7, '순창군', '', 'https://www.sunchang.go.kr',
|
||||||
|
'https://www.sunchang.go.kr/', parse_sunchang, 'sunchang.go.kr'),
|
||||||
|
jb(8, '완주군', '', 'https://www.wanju.go.kr',
|
||||||
|
'JSON', parse_wanju_json, 'wanju.go.kr', fetch_kind='json'),
|
||||||
|
jb(9, '익산시', '', 'https://www.iksan.go.kr',
|
||||||
|
'https://www.iksan.go.kr/index.do?menuUid=ff80808199f0d11c019a041b8e35174a', parse_group_sitemap, 'iksan.go.kr'),
|
||||||
|
jb(10, '임실군', '', 'https://www.imsil.go.kr',
|
||||||
|
'https://www.imsil.go.kr/index.imsil?menuCd=DOM_000000107003000000', parse_menu_div, 'imsil.go.kr'),
|
||||||
|
jb(11, '장수군', '', 'https://www.jangsu.go.kr',
|
||||||
|
'https://www.jangsu.go.kr/index.jangsu?menuCd=DOM_000000107001000000', parse_jangsu, 'jangsu.go.kr'),
|
||||||
|
jb(12, '전주시', '', 'https://www.jeonju.go.kr',
|
||||||
|
'https://www.jeonju.go.kr/index.9is?contentUid=ff8080818c7c2e8e018c7fd67b9a03b6', parse_jeonju, 'jeonju.go.kr', fetch_kind='html5'),
|
||||||
|
jb(13, '정읍시', '', 'https://www.jeongeup.go.kr',
|
||||||
|
'https://www.jeongeup.go.kr/index.jeongeup?menuCd=DOM_000000106002000000', parse_jeongeup, 'jeongeup.go.kr'),
|
||||||
|
jb(14, '진안군', '', 'https://www.jinan.go.kr',
|
||||||
|
'https://www.jinan.go.kr/index.jinan?menuCd=DOM_000000110002000000', parse_jinan, 'jinan.go.kr'),
|
||||||
|
jj(1, '서귀포시', 'https://www.seogwipo.go.kr',
|
||||||
|
'https://www.seogwipo.go.kr/help/sitemap.htm', parse_box_heading, 'seogwipo.go.kr'),
|
||||||
|
jj(2, '제주시', 'https://www.jejusi.go.kr',
|
||||||
|
'https://www.jejusi.go.kr/guide/sitemap.do', parse_box_heading, 'jejusi.go.kr'),
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def parse_gunsan(soup, base):
|
||||||
|
"""군산: div#all_pcmenu > div.allmenubox > a.Bmenu (D) + div.allmw ul.sub_pcmenu > li>a (E) + ul.dep3 (F) 재귀."""
|
||||||
|
rows = []
|
||||||
|
pc = soup.select_one('div#all_pcmenu')
|
||||||
|
if not pc:
|
||||||
|
return rows
|
||||||
|
for box in pc.select('div.allmenubox'):
|
||||||
|
bm = box.find('a', class_='Bmenu') or box.find('a')
|
||||||
|
D = clean_text(bm.get_text()) if bm else ''
|
||||||
|
if not D:
|
||||||
|
continue
|
||||||
|
ul = box.find('ul', class_='sub_pcmenu')
|
||||||
|
if not ul:
|
||||||
|
rows.append({'D': D, 'href': extract_href(bm), **{c: '' for c in 'EFGHIJ'}})
|
||||||
|
continue
|
||||||
|
# ★ ul.dep4 = 모바일 아코디언 잔재(F형제 전체를 '- '접두로 복제). 진짜 자식 아님 → 제거.
|
||||||
|
# 안 지우면 각 F의 G자식으로 재귀돼 카르테시안 폭발(2026-05-31 군산 208행 버그 수정).
|
||||||
|
for d4 in ul.select('ul.dep4'):
|
||||||
|
d4.decompose()
|
||||||
|
tmp = []
|
||||||
|
walk_ul(ul, [], tmp)
|
||||||
|
for r in rows_from_paths(tmp):
|
||||||
|
rows.append({'D': D, 'E': r.get('D', ''), 'F': r.get('E', ''),
|
||||||
|
'G': r.get('F', ''), 'H': r.get('G', ''), 'I': r.get('H', ''),
|
||||||
|
'J': r.get('I', ''), 'href': r.get('href', '')})
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
# 군산 파서 바인딩(전방참조 해결)
|
||||||
|
for _s in SITES:
|
||||||
|
if _s['name'] == '군산시':
|
||||||
|
_s['parser'] = parse_gunsan
|
||||||
|
|
||||||
|
|
||||||
|
# ================================================================
|
||||||
|
# 엑셀 생성 (충북 스크립트와 동일 로직)
|
||||||
|
# ================================================================
|
||||||
|
def write_excel(site, raw_rows):
|
||||||
|
name = site['name']
|
||||||
|
base = site['base']
|
||||||
|
domain = site['domain']
|
||||||
|
output = f"{site['folder']}\\{os.path.basename(os.path.dirname(site['folder']))}_{name}.xlsx"
|
||||||
|
|
||||||
|
def abs_url(href):
|
||||||
|
if not href:
|
||||||
|
return ''
|
||||||
|
href = href.strip()
|
||||||
|
if href.startswith(('javascript:', '#')):
|
||||||
|
return ''
|
||||||
|
if href.startswith(('http://', 'https://')):
|
||||||
|
return href
|
||||||
|
return urljoin(base + '/', href)
|
||||||
|
|
||||||
|
def is_external(url):
|
||||||
|
return url.startswith(('http://', 'https://')) and domain not in url
|
||||||
|
|
||||||
|
final_rows = []
|
||||||
|
i = 0
|
||||||
|
removed = 0
|
||||||
|
while i < len(raw_rows):
|
||||||
|
row = raw_rows[i]
|
||||||
|
if (i + 1 < len(raw_rows)
|
||||||
|
and row.get('G', '') == ''
|
||||||
|
and raw_rows[i + 1].get('D') == row.get('D')
|
||||||
|
and raw_rows[i + 1].get('E') == row.get('E')
|
||||||
|
and raw_rows[i + 1].get('F') == row.get('F')
|
||||||
|
and raw_rows[i + 1].get('G', '') != ''
|
||||||
|
and raw_rows[i + 1].get('href') == row.get('href')):
|
||||||
|
removed += 1
|
||||||
|
i += 1
|
||||||
|
continue
|
||||||
|
final_rows.append(row)
|
||||||
|
i += 1
|
||||||
|
|
||||||
|
print(f' [{name}] 원본 {len(raw_rows)} → 중복 {removed} → 최종 {len(final_rows)}')
|
||||||
|
if not final_rows:
|
||||||
|
print(f' [{name}] !! 행 0개 — 파서 점검 필요. 엑셀 생성 스킵.')
|
||||||
|
return False
|
||||||
|
|
||||||
|
shutil.copy(TEMPLATE, output)
|
||||||
|
wb = openpyxl.load_workbook(output)
|
||||||
|
ws = wb.active
|
||||||
|
ws.title = site['sheet']
|
||||||
|
|
||||||
|
HEADER = {'B1:R1', 'S1:W1', 'Y1:AA1'}
|
||||||
|
for rng in [str(m) for m in ws.merged_cells.ranges if str(m) not in HEADER]:
|
||||||
|
ws.unmerge_cells(rng)
|
||||||
|
for row in ws.iter_rows(min_row=3, max_row=ws.max_row, min_col=1, max_col=ws.max_column):
|
||||||
|
for cell in row:
|
||||||
|
cell.value = None
|
||||||
|
|
||||||
|
START = 3
|
||||||
|
template_r = 3
|
||||||
|
cur_max = ws.max_row
|
||||||
|
for idx, item in enumerate(final_rows, start=START):
|
||||||
|
if idx > cur_max:
|
||||||
|
for c in range(1, ws.max_column + 1):
|
||||||
|
srcc = ws.cell(template_r, c)
|
||||||
|
tgt = ws.cell(idx, c)
|
||||||
|
if srcc.has_style:
|
||||||
|
tgt.font = copy(srcc.font)
|
||||||
|
tgt.fill = copy(srcc.fill)
|
||||||
|
tgt.border = copy(srcc.border)
|
||||||
|
tgt.alignment = copy(srcc.alignment)
|
||||||
|
tgt.number_format = srcc.number_format
|
||||||
|
tgt.protection = copy(srcc.protection)
|
||||||
|
url = abs_url(item.get('href', ''))
|
||||||
|
ws.cell(idx, 2).value = idx - 2
|
||||||
|
ws.cell(idx, 3).value = name
|
||||||
|
ws.cell(idx, 4).value = item.get('D', '')
|
||||||
|
ws.cell(idx, 5).value = item.get('E', '')
|
||||||
|
ws.cell(idx, 6).value = item.get('F', '')
|
||||||
|
ws.cell(idx, 7).value = item.get('G', '')
|
||||||
|
ws.cell(idx, 8).value = item.get('H', '')
|
||||||
|
ws.cell(idx, 9).value = item.get('I', '')
|
||||||
|
ws.cell(idx, 10).value = item.get('J', '')
|
||||||
|
ws.cell(idx, 11).value = url
|
||||||
|
if is_external(url):
|
||||||
|
ws.cell(idx, 19).value = '외부링크'
|
||||||
|
|
||||||
|
END = START + len(final_rows) - 1
|
||||||
|
center = Alignment(horizontal='center', vertical='center', wrap_text=True)
|
||||||
|
left = Alignment(horizontal='left', vertical='center', wrap_text=False)
|
||||||
|
|
||||||
|
def merge_runs(col_letter, col_idx, group_cols=()):
|
||||||
|
runs = []
|
||||||
|
cur_val = ws.cell(START, col_idx).value
|
||||||
|
cur_grp = tuple(ws.cell(START, g).value for g in group_cols)
|
||||||
|
run_start = START
|
||||||
|
for r in range(START + 1, END + 1):
|
||||||
|
v = ws.cell(r, col_idx).value
|
||||||
|
g = tuple(ws.cell(r, gg).value for gg in group_cols)
|
||||||
|
if v == cur_val and g == cur_grp:
|
||||||
|
continue
|
||||||
|
if cur_val not in (None, '') and r - 1 > run_start:
|
||||||
|
runs.append((run_start, r - 1))
|
||||||
|
cur_val, cur_grp, run_start = v, g, r
|
||||||
|
if cur_val not in (None, '') and END > run_start:
|
||||||
|
runs.append((run_start, END))
|
||||||
|
for s, e in runs:
|
||||||
|
ws.merge_cells(f'{col_letter}{s}:{col_letter}{e}')
|
||||||
|
ws.cell(s, col_idx).alignment = center
|
||||||
|
return len(runs)
|
||||||
|
|
||||||
|
n_f = merge_runs('F', 6, group_cols=(4, 5))
|
||||||
|
n_e = merge_runs('E', 5, group_cols=(4,))
|
||||||
|
n_d = merge_runs('D', 4)
|
||||||
|
for r in range(START, END + 1):
|
||||||
|
for c in (4, 5, 6, 7):
|
||||||
|
if ws.cell(r, c).value is not None:
|
||||||
|
ws.cell(r, c).alignment = center
|
||||||
|
|
||||||
|
for r in range(1, END + 1):
|
||||||
|
ws.row_dimensions[r].height = 15
|
||||||
|
|
||||||
|
link_n = 0
|
||||||
|
for r in range(START, END + 1):
|
||||||
|
cell = ws.cell(r, 11)
|
||||||
|
u = cell.value
|
||||||
|
if u and isinstance(u, str) and u.startswith(('http://', 'https://')):
|
||||||
|
cell.hyperlink = u
|
||||||
|
old = cell.font
|
||||||
|
cell.font = Font(name=old.name or '맑은 고딕', size=old.size or 11,
|
||||||
|
bold=old.bold, italic=old.italic,
|
||||||
|
color='0000FF', underline='single')
|
||||||
|
cell.alignment = left
|
||||||
|
link_n += 1
|
||||||
|
|
||||||
|
wb.save(output)
|
||||||
|
ext_n = sum(1 for r in range(START, END + 1) if ws.cell(r, 19).value == '외부링크')
|
||||||
|
print(f' [{name}] 병합 D:{n_d} E:{n_e} F:{n_f} | 외부링크 {ext_n} | K링크 {link_n} → {output}')
|
||||||
|
return True
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
targets = sys.argv[1:] if len(sys.argv) > 1 else [s['name'] for s in SITES]
|
||||||
|
for site in SITES:
|
||||||
|
if site['name'] not in targets:
|
||||||
|
continue
|
||||||
|
print(f"\n{'='*60}\n[{site['idx']}.{site['name']}] {site['sitemap']}\n{'='*60}")
|
||||||
|
try:
|
||||||
|
if site['fetch_kind'] == 'json':
|
||||||
|
raw_rows = site['parser'](None, site['base'])
|
||||||
|
else:
|
||||||
|
sess = make_session(weak_ssl=site.get('weak_ssl', False))
|
||||||
|
html = fetch_html(site['sitemap'], session=sess)
|
||||||
|
engine = 'html5lib' if site.get('fetch_kind') == 'html5' else 'html.parser'
|
||||||
|
soup = BeautifulSoup(html, engine)
|
||||||
|
raw_rows = site['parser'](soup, site['base'])
|
||||||
|
write_excel(site, raw_rows)
|
||||||
|
except Exception as e:
|
||||||
|
print(f' [{site["name"]}] !! 실패: {e}')
|
||||||
|
import traceback
|
||||||
|
traceback.print_exc()
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
360
_스크립트/_jeonbuk_phase234_all.py
Normal file
360
_스크립트/_jeonbuk_phase234_all.py
Normal file
@ -0,0 +1,360 @@
|
|||||||
|
"""전북특별자치도 14개 시·군 + 제주특별자치도 2개 시 Phase 2~4 일괄 처리.
|
||||||
|
|
||||||
|
L(형태)·M(건수)·N(저작물유형)·O(공공누리)·P(부착위치)·Q(링크여부) 자동 채움.
|
||||||
|
매뉴얼: D:\\01.프로젝트\\DB수집\\사이트맵_수집_매뉴얼.md
|
||||||
|
KOGL(O열) 권위 판정은 별도 _recheck 단계에서 (feedback_kogl_image_rule).
|
||||||
|
"""
|
||||||
|
import re
|
||||||
|
import ssl
|
||||||
|
import sys
|
||||||
|
import time
|
||||||
|
import warnings
|
||||||
|
from urllib.parse import urljoin
|
||||||
|
from concurrent.futures import ThreadPoolExecutor, as_completed
|
||||||
|
|
||||||
|
import openpyxl
|
||||||
|
import requests
|
||||||
|
from requests.adapters import HTTPAdapter
|
||||||
|
from urllib3.util.ssl_ import create_urllib3_context
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
try:
|
||||||
|
sys.stdout.reconfigure(line_buffering=True)
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
ROOT = r'D:\01.프로젝트\DB수집'
|
||||||
|
|
||||||
|
|
||||||
|
class WeakSSLAdapter(HTTPAdapter):
|
||||||
|
def init_poolmanager(self, *args, **kwargs):
|
||||||
|
ctx = create_urllib3_context()
|
||||||
|
ctx.set_ciphers('DEFAULT@SECLEVEL=0')
|
||||||
|
ctx.options |= 0x4
|
||||||
|
ctx.check_hostname = False
|
||||||
|
ctx.verify_mode = ssl.CERT_NONE
|
||||||
|
kwargs['ssl_context'] = ctx
|
||||||
|
return super().init_poolmanager(*args, **kwargs)
|
||||||
|
|
||||||
|
|
||||||
|
TOTAL_PAT = re.compile(r'총\s*(?:게시물\s*)?(\d[\d,]*)\s*(?:건|개|page|페이지)', re.I)
|
||||||
|
TOTAL_PAT_LOOSE = re.compile(r'(?:전체|총)\s*[:\-]?\s*(\d[\d,]*)\s*건', re.I)
|
||||||
|
KOGL_IMG_PAT = re.compile(r'(?:new_)?img_open(?:type|code)(\d{1,2})\.(?:png|jpe?g|gif)', re.I)
|
||||||
|
KOGL_LINK_PAT = re.compile(r'kogl\.or\.kr/info/licenseType(\d)', re.I)
|
||||||
|
YOUTUBE_PAT = re.compile(r'(?:youtube\.com|youtu\.be)', re.I)
|
||||||
|
VIDEO_EXT = re.compile(r'\.(mp4|webm|mov|avi)(?:\?|$)', re.I)
|
||||||
|
DETAIL_PAT = re.compile(
|
||||||
|
r'(mode=V|view\.do|view\.9is|/view\b|bbtSn=|dataUid=|dataSid=|seqRepeat=|'
|
||||||
|
r'nttId=|nttNo=|articleNo=|boardSeq=|bbsSeq=|not_ancmt|menukey=)', re.I)
|
||||||
|
|
||||||
|
BODY_SEL = ['#main-contents', '#content', '#contents', '.contents', '#txt',
|
||||||
|
'main', '#container', '#sub']
|
||||||
|
|
||||||
|
|
||||||
|
def site(i, prov, name, weak=False):
|
||||||
|
folder = fr'{ROOT}\작업파일\광역_사이트맵\{prov}\{i}.{name}'
|
||||||
|
return name, {'xlsx': fr'{folder}\{prov}_{name}.xlsx', 'body_sel': BODY_SEL, 'weak_ssl': weak}
|
||||||
|
|
||||||
|
|
||||||
|
SITES = dict([
|
||||||
|
site(1, '전북특별자치도', '고창군'),
|
||||||
|
site(2, '전북특별자치도', '군산시'),
|
||||||
|
site(3, '전북특별자치도', '김제시'),
|
||||||
|
site(4, '전북특별자치도', '남원시'),
|
||||||
|
site(5, '전북특별자치도', '무주군'),
|
||||||
|
site(6, '전북특별자치도', '부안군'),
|
||||||
|
site(7, '전북특별자치도', '순창군'),
|
||||||
|
site(8, '전북특별자치도', '완주군'),
|
||||||
|
site(9, '전북특별자치도', '익산시'),
|
||||||
|
site(10, '전북특별자치도', '임실군'),
|
||||||
|
site(11, '전북특별자치도', '장수군'),
|
||||||
|
site(12, '전북특별자치도', '전주시'),
|
||||||
|
site(13, '전북특별자치도', '정읍시'),
|
||||||
|
site(14, '전북특별자치도', '진안군'),
|
||||||
|
site(1, '제주특별자치도', '서귀포시'),
|
||||||
|
site(2, '제주특별자치도', '제주시'),
|
||||||
|
])
|
||||||
|
|
||||||
|
|
||||||
|
def make_session(weak_ssl=False):
|
||||||
|
s = requests.Session()
|
||||||
|
s.headers.update(H)
|
||||||
|
if weak_ssl:
|
||||||
|
s.mount('https://', WeakSSLAdapter())
|
||||||
|
return s
|
||||||
|
|
||||||
|
|
||||||
|
def fetch(session, url, timeout=5):
|
||||||
|
try:
|
||||||
|
r = session.get(url, timeout=timeout, verify=False, allow_redirects=True)
|
||||||
|
meta = re.search(rb'<meta[^>]*charset=["\']?\s*([\w-]+)', r.content[:4096], re.I)
|
||||||
|
r.encoding = meta.group(1).decode('ascii', errors='ignore') if meta else r.apparent_encoding
|
||||||
|
if r.status_code == 200:
|
||||||
|
return BeautifulSoup(r.text, 'html.parser')
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def get_body(soup, selectors):
|
||||||
|
for sel in selectors:
|
||||||
|
el = soup.select_one(sel)
|
||||||
|
if el:
|
||||||
|
return el
|
||||||
|
return soup
|
||||||
|
|
||||||
|
|
||||||
|
def detect_form(body):
|
||||||
|
has_paging = bool(body.select('.pagination, .paging, nav.paging, .page_nav, .board_paging'))
|
||||||
|
text_inputs = [i for i in body.find_all('input')
|
||||||
|
if (i.get('type') or 'text').lower() in ('text', 'search')]
|
||||||
|
has_search = len(text_inputs) >= 1
|
||||||
|
txt = body.get_text(' ', strip=True)
|
||||||
|
m = TOTAL_PAT.search(txt) or TOTAL_PAT_LOOSE.search(txt)
|
||||||
|
total = None
|
||||||
|
if m:
|
||||||
|
digits = m.group(1).replace(',', '')
|
||||||
|
if digits.isdigit():
|
||||||
|
total = int(digits)
|
||||||
|
is_board = has_paging or has_search or (total is not None)
|
||||||
|
if is_board:
|
||||||
|
return '게시판', total if total is not None else 0
|
||||||
|
return '페이지', 1
|
||||||
|
|
||||||
|
|
||||||
|
def extract_detail_urls(body, base_url, limit=5):
|
||||||
|
urls = []
|
||||||
|
seen = set()
|
||||||
|
for a in body.find_all('a', href=True):
|
||||||
|
h = a['href']
|
||||||
|
if not h or h.startswith('#'):
|
||||||
|
continue
|
||||||
|
if DETAIL_PAT.search(h):
|
||||||
|
full = urljoin(base_url, h)
|
||||||
|
if full not in seen:
|
||||||
|
seen.add(full)
|
||||||
|
urls.append(full)
|
||||||
|
if len(urls) >= limit:
|
||||||
|
break
|
||||||
|
return urls
|
||||||
|
|
||||||
|
|
||||||
|
def detect_media(body):
|
||||||
|
has_text = len(body.get_text(strip=True)) > 30
|
||||||
|
has_image = False
|
||||||
|
for img in body.find_all('img'):
|
||||||
|
src = img.get('src', '')
|
||||||
|
if KOGL_IMG_PAT.search(src):
|
||||||
|
continue
|
||||||
|
if not src:
|
||||||
|
continue
|
||||||
|
has_image = True
|
||||||
|
break
|
||||||
|
has_video = False
|
||||||
|
for iframe in body.find_all('iframe'):
|
||||||
|
if YOUTUBE_PAT.search(iframe.get('src', '')):
|
||||||
|
has_video = True
|
||||||
|
break
|
||||||
|
if not has_video:
|
||||||
|
for a in body.find_all('a', href=True):
|
||||||
|
if YOUTUBE_PAT.search(a['href']):
|
||||||
|
has_video = True
|
||||||
|
break
|
||||||
|
if not has_video and body.find_all('video'):
|
||||||
|
has_video = True
|
||||||
|
if not has_video and VIDEO_EXT.search(str(body)):
|
||||||
|
has_video = True
|
||||||
|
return has_image, has_video, has_text
|
||||||
|
|
||||||
|
|
||||||
|
def n_string(has_text, has_image, has_video):
|
||||||
|
parts = []
|
||||||
|
if has_text: parts.append('어문')
|
||||||
|
if has_image: parts.append('이미지')
|
||||||
|
if has_video: parts.append('영상')
|
||||||
|
return ','.join(parts) if parts else '없음'
|
||||||
|
|
||||||
|
|
||||||
|
def img_has_valid_anchor(img):
|
||||||
|
p = img.parent
|
||||||
|
while p is not None:
|
||||||
|
if p.name == 'a':
|
||||||
|
href = p.get('href', '')
|
||||||
|
if href and not href.startswith('#') and not href.lower().startswith('javascript:'):
|
||||||
|
return True
|
||||||
|
return False
|
||||||
|
p = p.parent
|
||||||
|
return False
|
||||||
|
|
||||||
|
|
||||||
|
def detect_kogl(body):
|
||||||
|
types = set()
|
||||||
|
q_any_y = False
|
||||||
|
q_any_n = False
|
||||||
|
for a in body.find_all('a', href=True):
|
||||||
|
m = KOGL_LINK_PAT.search(a['href'])
|
||||||
|
if m:
|
||||||
|
types.add(int(m.group(1)))
|
||||||
|
q_any_y = True
|
||||||
|
for img in body.find_all('img'):
|
||||||
|
src = img.get('src', '')
|
||||||
|
m = KOGL_IMG_PAT.search(src)
|
||||||
|
if m:
|
||||||
|
types.add(int(m.group(1)))
|
||||||
|
if img_has_valid_anchor(img):
|
||||||
|
q_any_y = True
|
||||||
|
else:
|
||||||
|
q_any_n = True
|
||||||
|
for el in body.find_all(style=True):
|
||||||
|
m = KOGL_IMG_PAT.search(el.get('style', ''))
|
||||||
|
if m:
|
||||||
|
types.add(int(m.group(1)))
|
||||||
|
q_any_n = True
|
||||||
|
if not types:
|
||||||
|
return set(), None
|
||||||
|
return types, ('Y' if q_any_y else 'N')
|
||||||
|
|
||||||
|
|
||||||
|
def process_row(session, url, body_selectors):
|
||||||
|
out = {'L': '', 'M': '', 'N': '', 'O': '', 'P': '', 'Q': '', 'note': ''}
|
||||||
|
soup = fetch(session, url)
|
||||||
|
if soup is None:
|
||||||
|
out['note'] = '접근 실패'
|
||||||
|
return out
|
||||||
|
body = get_body(soup, body_selectors)
|
||||||
|
|
||||||
|
form, count = detect_form(body)
|
||||||
|
out['L'] = form
|
||||||
|
out['M'] = count if form == '게시판' else 1
|
||||||
|
|
||||||
|
has_img, has_vid, has_txt = detect_media(body)
|
||||||
|
types_main, q_main = detect_kogl(body)
|
||||||
|
P = '게시판' if types_main else ''
|
||||||
|
types_all = set(types_main)
|
||||||
|
q_flags = []
|
||||||
|
if q_main:
|
||||||
|
q_flags.append(q_main)
|
||||||
|
|
||||||
|
if form == '게시판':
|
||||||
|
detail_urls = extract_detail_urls(body, url, limit=2)
|
||||||
|
for du in detail_urls:
|
||||||
|
d_soup = fetch(session, du, timeout=5)
|
||||||
|
if not d_soup:
|
||||||
|
continue
|
||||||
|
d_body = get_body(d_soup, body_selectors)
|
||||||
|
di, dv, dt = detect_media(d_body)
|
||||||
|
has_img = has_img or di
|
||||||
|
has_vid = has_vid or dv
|
||||||
|
has_txt = has_txt or dt
|
||||||
|
dt_types, dt_q = detect_kogl(d_body)
|
||||||
|
if dt_types and not types_main and not P:
|
||||||
|
P = '게시물'
|
||||||
|
types_all |= dt_types
|
||||||
|
if dt_q:
|
||||||
|
q_flags.append(dt_q)
|
||||||
|
|
||||||
|
out['N'] = n_string(has_txt, has_img, has_vid)
|
||||||
|
|
||||||
|
if not types_all:
|
||||||
|
out['O'] = '미부착'
|
||||||
|
else:
|
||||||
|
sorted_types = sorted(types_all)
|
||||||
|
out['O'] = ','.join(f'{n}유형' for n in sorted_types)
|
||||||
|
out['P'] = P if P else '게시판'
|
||||||
|
out['Q'] = 'Y' if 'Y' in q_flags else 'N'
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def run_site(name, xlsx, body_selectors, weak_ssl=False, workers=14):
|
||||||
|
print(f'\n{"="*60}\n[{name}] {xlsx}\n{"="*60}')
|
||||||
|
wb = openpyxl.load_workbook(xlsx)
|
||||||
|
ws = wb.active
|
||||||
|
START = 3
|
||||||
|
END = START - 1
|
||||||
|
for r in range(START, ws.max_row + 1):
|
||||||
|
if ws.cell(r, 2).value is None:
|
||||||
|
break
|
||||||
|
END = r
|
||||||
|
|
||||||
|
tasks = []
|
||||||
|
for r in range(START, END + 1):
|
||||||
|
url = ws.cell(r, 11).value
|
||||||
|
is_ext = (ws.cell(r, 19).value == '외부링크')
|
||||||
|
tasks.append((r, url, is_ext))
|
||||||
|
n_ext = sum(1 for t in tasks if t[2])
|
||||||
|
print(f' 처리 대상: 총 {len(tasks)}행 (외부링크 {n_ext})')
|
||||||
|
|
||||||
|
t0 = time.time()
|
||||||
|
results = {}
|
||||||
|
session = make_session(weak_ssl=weak_ssl)
|
||||||
|
|
||||||
|
def worker(task):
|
||||||
|
row, url, is_ext = task
|
||||||
|
if is_ext:
|
||||||
|
return row, {'L': '사이트', 'M': '', 'N': '', 'O': '', 'P': '', 'Q': '', 'note': ''}
|
||||||
|
if not url or not isinstance(url, str):
|
||||||
|
return row, {'L': '', 'M': '', 'N': '', 'O': '', 'P': '', 'Q': '', 'note': 'URL 없음'}
|
||||||
|
return row, process_row(session, url, body_selectors)
|
||||||
|
|
||||||
|
done = 0
|
||||||
|
with ThreadPoolExecutor(max_workers=workers) as ex:
|
||||||
|
futs = [ex.submit(worker, t) for t in tasks]
|
||||||
|
for fut in as_completed(futs):
|
||||||
|
row, res = fut.result()
|
||||||
|
results[row] = res
|
||||||
|
done += 1
|
||||||
|
if done % 50 == 0 or done == len(tasks):
|
||||||
|
print(f' 진행 {done}/{len(tasks)} ({time.time()-t0:.0f}s)')
|
||||||
|
|
||||||
|
for r in range(START, END + 1):
|
||||||
|
res = results.get(r, {})
|
||||||
|
if not res:
|
||||||
|
continue
|
||||||
|
if res.get('L'): ws.cell(r, 12).value = res['L']
|
||||||
|
if res.get('M') != '': ws.cell(r, 13).value = res['M']
|
||||||
|
if res.get('N'): ws.cell(r, 14).value = res['N']
|
||||||
|
if res.get('O'): ws.cell(r, 15).value = res['O']
|
||||||
|
if res.get('P'): ws.cell(r, 16).value = res['P']
|
||||||
|
if res.get('Q'): ws.cell(r, 17).value = res['Q']
|
||||||
|
if res.get('note'):
|
||||||
|
existing = ws.cell(r, 19).value
|
||||||
|
if not existing:
|
||||||
|
ws.cell(r, 19).value = res['note']
|
||||||
|
|
||||||
|
wb.save(xlsx)
|
||||||
|
forms = {}
|
||||||
|
attach = {'미부착': 0, '부착': 0, '기타': 0}
|
||||||
|
q_dist = {'Y': 0, 'N': 0, '': 0}
|
||||||
|
for r, res in results.items():
|
||||||
|
forms[res.get('L', '')] = forms.get(res.get('L', ''), 0) + 1
|
||||||
|
o = res.get('O', '')
|
||||||
|
if o == '미부착': attach['미부착'] += 1
|
||||||
|
elif o and '유형' in o: attach['부착'] += 1
|
||||||
|
else: attach['기타'] += 1
|
||||||
|
q = res.get('Q', '')
|
||||||
|
q_dist[q] = q_dist.get(q, 0) + 1
|
||||||
|
print(f' [{name}] L분포 {forms} | 부착 {attach} | Q {q_dist} | 시간 {time.time()-t0:.0f}s')
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
targets = sys.argv[1:] if len(sys.argv) > 1 else list(SITES.keys())
|
||||||
|
total_t0 = time.time()
|
||||||
|
for name in targets:
|
||||||
|
if name not in SITES:
|
||||||
|
print(f' 알 수 없음: {name}')
|
||||||
|
continue
|
||||||
|
cfg = SITES[name]
|
||||||
|
try:
|
||||||
|
run_site(name, cfg['xlsx'], cfg['body_sel'], weak_ssl=cfg.get('weak_ssl', False))
|
||||||
|
except Exception as e:
|
||||||
|
print(f' [{name}] 실패: {e}')
|
||||||
|
import traceback
|
||||||
|
traceback.print_exc()
|
||||||
|
print(f'\n총 소요: {time.time()-total_t0:.0f}s')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
30
_스크립트/_mk_batch2_ntools.py
Normal file
30
_스크립트/_mk_batch2_ntools.py
Normal file
@ -0,0 +1,30 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""공공기관2 몽타주 N 도구 생성(nshot/napply/nbatch 경로 적응)."""
|
||||||
|
import io, os
|
||||||
|
D = r"D:\01.프로젝트\DB수집\_스크립트"
|
||||||
|
|
||||||
|
def transform(src, dst, repls):
|
||||||
|
s = io.open(os.path.join(D, src), encoding="utf-8").read()
|
||||||
|
for a, b in repls:
|
||||||
|
if a not in s:
|
||||||
|
print(" [warn] 못찾음:", repr(a[:50]))
|
||||||
|
s = s.replace(a, b)
|
||||||
|
io.open(os.path.join(D, dst), "w", encoding="utf-8").write(s)
|
||||||
|
print("생성:", dst)
|
||||||
|
|
||||||
|
# 번호폴더 구조 대응: xlsx 경로를 glob로
|
||||||
|
XLSX_OLD = "xlsx = os.path.join(OUTDIR, f'{name}.xlsx')"
|
||||||
|
XLSX_NEW = "xlsx = __import__('glob').glob(os.path.join(OUTDIR, f'*.{name}', f'{name}.xlsx'))[0]"
|
||||||
|
OUTDIR_OLD = r"OUTDIR = r'D:\01.프로젝트\DB수집\공공기관'"
|
||||||
|
OUTDIR_NEW = r"OUTDIR = r'D:\01.프로젝트\DB수집\작업파일\공공기관2'"
|
||||||
|
|
||||||
|
transform("_공공기관_nshot.py", "_공공기관2_nshot.py", [(OUTDIR_OLD, OUTDIR_NEW), (XLSX_OLD, XLSX_NEW)])
|
||||||
|
transform("_공공기관_napply.py", "_공공기관2_napply.py", [(OUTDIR_OLD, OUTDIR_NEW), (XLSX_OLD, XLSX_NEW)])
|
||||||
|
|
||||||
|
# nbatch: 모듈/probe 경로
|
||||||
|
transform("_공공기관_nbatch.py", "_공공기관2_nbatch.py", [
|
||||||
|
("_공공기관_nshot", "_공공기관2_nshot"),
|
||||||
|
("_공공기관_napply", "_공공기관2_napply"),
|
||||||
|
("_공공기관_probe.json", "_공공기관2_probe.json"),
|
||||||
|
])
|
||||||
|
print("done")
|
||||||
39
_스크립트/_mk_batch2_tools.py
Normal file
39
_스크립트/_mk_batch2_tools.py
Normal file
@ -0,0 +1,39 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""공공기관2 배치용 probe/phase1/phase234 스크립트 생성 (경로만 치환)."""
|
||||||
|
import io, os
|
||||||
|
|
||||||
|
D = r"D:\01.프로젝트\DB수집\_스크립트"
|
||||||
|
|
||||||
|
def transform(src_name, dst_name, repls):
|
||||||
|
src = io.open(os.path.join(D, src_name), encoding="utf-8").read()
|
||||||
|
for a, b in repls:
|
||||||
|
if a not in src:
|
||||||
|
print(" [warn] 패턴 못찾음:", repr(a[:60]))
|
||||||
|
src = src.replace(a, b)
|
||||||
|
io.open(os.path.join(D, dst_name), "w", encoding="utf-8").write(src)
|
||||||
|
print("생성:", dst_name)
|
||||||
|
|
||||||
|
# probe: STATUS + 출력 json
|
||||||
|
transform("_공공기관_probe.py", "_공공기관2_probe.py", [
|
||||||
|
(r"작업파일\공공기관_작업현황.xlsx", r"작업파일\공공기관2\공공기관2_작업현황.xlsx"),
|
||||||
|
("_공공기관_probe.json", "_공공기관2_probe.json"),
|
||||||
|
])
|
||||||
|
|
||||||
|
# phase1: OUTDIR(번호폴더구조) + PROBE json + 출력경로 번호폴더
|
||||||
|
transform("_공공기관_phase1.py", "_공공기관2_phase1.py", [
|
||||||
|
(r"OUTDIR = r'D:\01.프로젝트\DB수집\공공기관'",
|
||||||
|
r"OUTDIR = r'D:\01.프로젝트\DB수집\작업파일\공공기관2'"),
|
||||||
|
("_공공기관_probe.json", "_공공기관2_probe.json"),
|
||||||
|
("output = os.path.join(OUTDIR, f'{name}.xlsx')",
|
||||||
|
"output = os.path.join(OUTDIR, f'{num}.{name}', f'{name}.xlsx')"),
|
||||||
|
])
|
||||||
|
|
||||||
|
# phase234: OUTDIR + PROBE json + 입력경로 번호폴더
|
||||||
|
transform("_공공기관_phase234.py", "_공공기관2_phase234.py", [
|
||||||
|
(r"OUTDIR = r'D:\01.프로젝트\DB수집\공공기관'",
|
||||||
|
r"OUTDIR = r'D:\01.프로젝트\DB수집\작업파일\공공기관2'"),
|
||||||
|
("_공공기관_probe.json", "_공공기관2_probe.json"),
|
||||||
|
("xlsx = os.path.join(OUTDIR, f'{name}.xlsx')",
|
||||||
|
"xlsx = os.path.join(OUTDIR, f'{num}.{name}', f'{name}.xlsx')"),
|
||||||
|
])
|
||||||
|
print("done")
|
||||||
179
_스크립트/_plan_splits.py
Normal file
179
_스크립트/_plan_splits.py
Normal file
@ -0,0 +1,179 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""335~443행 분해 계획 산출(읽기전용).
|
||||||
|
- 미분해 F섹션(type1 탭 URL이 시트에 1개=자기자신만 존재) → 그 type1 탭들로 분해
|
||||||
|
- 이미 분해된 잎 행의 깊은 중첩(type3가 별도 URL, 시트에 없음) → 그 type3로 분해
|
||||||
|
- type3가 #nav 앵커면 분해 안 함(M=앵커수)
|
||||||
|
재귀로 깊은 단계까지. 결과: 분해 잡 목록 + 총 증가행수.
|
||||||
|
"""
|
||||||
|
import re, json, warnings
|
||||||
|
from urllib.parse import urlparse
|
||||||
|
import requests
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
import openpyxl
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
XLSX = r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\2.공주시\충청남도_공주시.xlsx'
|
||||||
|
BASE = 'https://www.gongju.go.kr'
|
||||||
|
H = {'User-Agent': 'Mozilla/5.0 Chrome/120 Safari/537.36'}
|
||||||
|
LO, HI = 335, 443
|
||||||
|
_cache = {}
|
||||||
|
|
||||||
|
|
||||||
|
def full(h):
|
||||||
|
return h if h.startswith('http') else BASE + h
|
||||||
|
|
||||||
|
|
||||||
|
def norm(u):
|
||||||
|
"""시트 멤버십 비교용 정규화: host 제거, path(+bbs id)만"""
|
||||||
|
p = urlparse(u if u.startswith('http') else BASE + u)
|
||||||
|
path = p.path
|
||||||
|
return path
|
||||||
|
|
||||||
|
|
||||||
|
def get(u):
|
||||||
|
u = full(u)
|
||||||
|
if u in _cache:
|
||||||
|
return _cache[u]
|
||||||
|
try:
|
||||||
|
x = requests.get(u, headers=H, timeout=25, verify=False)
|
||||||
|
x.encoding = x.apparent_encoding or 'utf-8'
|
||||||
|
html = x.text
|
||||||
|
except Exception:
|
||||||
|
html = ''
|
||||||
|
_cache[u] = html
|
||||||
|
return html
|
||||||
|
|
||||||
|
|
||||||
|
def tab_groups(html):
|
||||||
|
s = BeautifulSoup(html, 'html.parser')
|
||||||
|
t1 = []
|
||||||
|
t3 = []
|
||||||
|
for ul in s.find_all('ul'):
|
||||||
|
cls = ' '.join(ul.get('class') or [])
|
||||||
|
if 'tab-ul' not in cls.lower():
|
||||||
|
continue
|
||||||
|
items = [(a.get_text(strip=True), a.get('href', '')) for a in ul.find_all('a')
|
||||||
|
if a.get_text(strip=True) and a.get('href')]
|
||||||
|
if not items:
|
||||||
|
continue
|
||||||
|
if 'type1' in cls:
|
||||||
|
t1 = items
|
||||||
|
elif not t3:
|
||||||
|
t3 = items
|
||||||
|
return t1, t3
|
||||||
|
|
||||||
|
|
||||||
|
def board_count(html):
|
||||||
|
s = BeautifulSoup(html, 'html.parser')
|
||||||
|
el = s.select_one('.program--count strong')
|
||||||
|
if el:
|
||||||
|
m = re.sub(r'[^0-9]', '', el.get_text())
|
||||||
|
return int(m) if m else None
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def is_board(u):
|
||||||
|
return bool(re.search(r'list\.do|/bbs/|BBSMSTR', u))
|
||||||
|
|
||||||
|
|
||||||
|
def leaf_info(name, u):
|
||||||
|
"""잎 1개의 (L,M) 결정 + 그 잎의 #nav 앵커수 반영"""
|
||||||
|
fu = full(u)
|
||||||
|
html = get(fu)
|
||||||
|
if is_board(u):
|
||||||
|
c = board_count(html)
|
||||||
|
return {'name': name, 'url': fu, 'L': '게시판', 'M': c if c is not None else 0}
|
||||||
|
# 페이지: 자기 type3가 #nav면 M=앵커수
|
||||||
|
_, t3 = tab_groups(html)
|
||||||
|
navs = [h for t, h in t3 if h.startswith('#')]
|
||||||
|
if len(navs) >= 2:
|
||||||
|
return {'name': name, 'url': fu, 'L': '페이지', 'M': len(navs)}
|
||||||
|
return {'name': name, 'url': fu, 'L': '페이지', 'M': 1}
|
||||||
|
|
||||||
|
|
||||||
|
def url_type3(u):
|
||||||
|
"""페이지의 type3 별도URL 탭들(없으면 [])"""
|
||||||
|
if is_board(u):
|
||||||
|
return []
|
||||||
|
_, t3 = tab_groups(get(full(u)))
|
||||||
|
return [(t, full(h)) for t, h in t3 if not h.startswith('#')]
|
||||||
|
|
||||||
|
|
||||||
|
def expand_nested(name, u, sheet_urls, depth=0):
|
||||||
|
"""깊이 2 평탄 분해(공주 탭 최대 2단계). 자식이 url-type3를 가지면 1단계만 더 펼침."""
|
||||||
|
grand = url_type3(u) if depth == 0 else []
|
||||||
|
if len(grand) >= 2:
|
||||||
|
return {'name': name, 'url': full(u),
|
||||||
|
'children': [leaf_info(t, h) for t, h in grand]}
|
||||||
|
return leaf_info(name, u)
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
wb = openpyxl.load_workbook(XLSX)
|
||||||
|
ws = wb.active
|
||||||
|
sheet_urls = set()
|
||||||
|
for r in range(3, 445):
|
||||||
|
u = ws.cell(r, 11).value
|
||||||
|
if isinstance(u, str) and u.startswith('http'):
|
||||||
|
sheet_urls.add(norm(u))
|
||||||
|
|
||||||
|
jobs = []
|
||||||
|
for r in range(LO, HI + 1):
|
||||||
|
u = ws.cell(r, 11).value
|
||||||
|
if not (isinstance(u, str) and u.startswith('http')):
|
||||||
|
continue
|
||||||
|
html = get(u)
|
||||||
|
t1, t3 = tab_groups(html)
|
||||||
|
t1_urls = [(t, full(h)) for t, h in t1 if not h.startswith('#')]
|
||||||
|
present = sum(1 for t, h in t1_urls if norm(h) in sheet_urls)
|
||||||
|
url_t3 = [(t, full(h)) for t, h in t3 if not h.startswith('#')]
|
||||||
|
|
||||||
|
label = (ws.cell(r, 6).value or '')
|
||||||
|
if ws.cell(r, 7).value:
|
||||||
|
label += ' > ' + str(ws.cell(r, 7).value)
|
||||||
|
if ws.cell(r, 8).value:
|
||||||
|
label += ' > ' + str(ws.cell(r, 8).value)
|
||||||
|
|
||||||
|
# 미분해 F섹션: type1 URL>=2 이고 시트에 자기 1개만 → type1 자식(각자 type3 1단계 더)
|
||||||
|
if len(t1_urls) >= 2 and present <= 1:
|
||||||
|
children = [expand_nested(t, h, sheet_urls, depth=0) for t, h in t1_urls]
|
||||||
|
jobs.append(('SECTION', r, label, children))
|
||||||
|
# 이미 분해된 G-잎의 깊은 type3 중첩 → type3 자식(잎)
|
||||||
|
elif len(url_t3) >= 2:
|
||||||
|
children = [leaf_info(t, h) for t, h in url_t3]
|
||||||
|
jobs.append(('NEST', r, label, children))
|
||||||
|
|
||||||
|
# 요약
|
||||||
|
def count_leaves(nodes):
|
||||||
|
n = 0
|
||||||
|
for nd in nodes:
|
||||||
|
if nd.get('children'):
|
||||||
|
n += count_leaves(nd['children'])
|
||||||
|
else:
|
||||||
|
n += 1
|
||||||
|
return n
|
||||||
|
|
||||||
|
total_add = 0
|
||||||
|
print('=== 분해 계획 (335~443) ===')
|
||||||
|
for kind, r, label, children in jobs:
|
||||||
|
leaves = count_leaves(children)
|
||||||
|
add = leaves - 1 # 기존 1행 대체
|
||||||
|
total_add += add
|
||||||
|
print('\n[%s] r%d %s → 잎 %d (기존1, +%d)' % (kind, r, label, leaves, add))
|
||||||
|
def pr(nodes, ind=1):
|
||||||
|
for nd in nodes:
|
||||||
|
tag = ('게시판%s' % nd['M']) if nd.get('L') == '게시판' else ('페이지M%s' % nd.get('M', ''))
|
||||||
|
if nd.get('children'):
|
||||||
|
print(' ' * ind + '▸ %s' % nd['name'])
|
||||||
|
pr(nd['children'], ind + 1)
|
||||||
|
else:
|
||||||
|
print(' ' * ind + '- %s [%s]' % (nd['name'], tag))
|
||||||
|
pr(children)
|
||||||
|
print('\n총 분해 잡 %d개, 총 증가 행수 +%d (최종 데이터행 ~ %d)' % (len(jobs), total_add, 443 + total_add))
|
||||||
|
json.dump([{'kind': k, 'row': r, 'label': l, 'children': c} for k, r, l, c in jobs],
|
||||||
|
open(r'D:\01.프로젝트\DB수집\_스크립트\_split_plan.json', 'w', encoding='utf-8'),
|
||||||
|
ensure_ascii=False, indent=1)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
107
_스크립트/_probe2.py
Normal file
107
_스크립트/_probe2.py
Normal file
@ -0,0 +1,107 @@
|
|||||||
|
"""Detailed structural analysis of representative sitemap pages."""
|
||||||
|
import warnings
|
||||||
|
from urllib.parse import urljoin
|
||||||
|
|
||||||
|
import requests
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
|
||||||
|
SAMPLES = {
|
||||||
|
# name: (sitemap_url, selector_for_container)
|
||||||
|
'논산시': ('https://nonsan.go.kr/kor/html/sub07/0701.html', '.sitemap'),
|
||||||
|
'당진시': ('https://www.dangjin.go.kr/kor/sitemap_11.do', '#sitemap'),
|
||||||
|
'보령시': ('https://www.brcn.go.kr/kor/sitemap_11.do', '#sitemap'),
|
||||||
|
'서천군': ('https://www.seocheon.go.kr/kor/sitemap_11.do', '#sitemap'),
|
||||||
|
'청양군': ('https://www.cheongyang.go.kr/kor/sitemap_11.do', '#sitemap'),
|
||||||
|
'태안군': ('https://www.taean.go.kr/kor/sitemap_11.do', '#sitemap'),
|
||||||
|
'아산시': ('https://www.asan.go.kr/main/sitemap.do', '.sitemap'),
|
||||||
|
'예산군': ('https://www.yesan.go.kr/kor/sitemap.do', 'ul.sitemap'),
|
||||||
|
'천안시': ('https://www.cheonan.go.kr/kor/sitemap.do', 'ul.sitemap'),
|
||||||
|
'홍성군': ('https://www.hongseong.go.kr/kor/sitemap.do', 'div.sitemap'),
|
||||||
|
'부여군': ('https://www.buyeo.go.kr/html/kr/html/sub07/0701.html', '.sitemap'),
|
||||||
|
'서산시': ('https://www.seosan.go.kr/www/contents.do?key=151', '.sitemap'),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def fetch(url):
|
||||||
|
r = requests.get(url, headers=H, timeout=20, verify=False, allow_redirects=True)
|
||||||
|
r.encoding = r.apparent_encoding
|
||||||
|
return r.url, r.text, r.status_code
|
||||||
|
|
||||||
|
|
||||||
|
def dump_outline(el, indent=0, max_lines=80, lines=None):
|
||||||
|
if lines is None:
|
||||||
|
lines = []
|
||||||
|
if len(lines) >= max_lines:
|
||||||
|
return lines
|
||||||
|
name = el.name
|
||||||
|
cls = ' '.join(el.get('class', []))
|
||||||
|
eid = el.get('id', '')
|
||||||
|
label = f'{name}'
|
||||||
|
if eid:
|
||||||
|
label += f'#{eid}'
|
||||||
|
if cls:
|
||||||
|
label += f'.{cls.replace(" ", ".")}'
|
||||||
|
if name == 'a':
|
||||||
|
txt = el.get_text(strip=True)[:40]
|
||||||
|
href = el.get('href', '')[:60]
|
||||||
|
lines.append(' ' * indent + f'{label} "{txt}" → {href}')
|
||||||
|
else:
|
||||||
|
lines.append(' ' * indent + label)
|
||||||
|
for c in el.find_all(recursive=False):
|
||||||
|
if c.name in ('script', 'style'):
|
||||||
|
continue
|
||||||
|
dump_outline(c, indent + 1, max_lines, lines)
|
||||||
|
if len(lines) >= max_lines:
|
||||||
|
return lines
|
||||||
|
return lines
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
for name, (url, sel) in SAMPLES.items():
|
||||||
|
print(f'\n{"="*70}\n{name}: {url}\n{"="*70}')
|
||||||
|
try:
|
||||||
|
final, html, code = fetch(url)
|
||||||
|
except Exception as e:
|
||||||
|
print(f' ERR: {e}')
|
||||||
|
continue
|
||||||
|
if code != 200:
|
||||||
|
print(f' HTTP {code}')
|
||||||
|
# Try alternative locations
|
||||||
|
for alt in [url.replace('sub07', 'sub06'), url.replace('contents.do?key=151', 'sitemap.do')]:
|
||||||
|
try:
|
||||||
|
final, html, code = fetch(alt)
|
||||||
|
if code == 200:
|
||||||
|
print(f' 대체 OK: {alt}')
|
||||||
|
break
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
else:
|
||||||
|
continue
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
el = soup.select_one(sel)
|
||||||
|
if el is None:
|
||||||
|
print(f' selector "{sel}" 매칭 실패')
|
||||||
|
# Find any container with many anchors
|
||||||
|
for d in soup.find_all(['div', 'section', 'main', 'ul']):
|
||||||
|
if len(d.find_all('a')) >= 100:
|
||||||
|
cls = ' '.join(d.get('class', []))
|
||||||
|
print(f' 대안 발견: {d.name}.{cls} id={d.get("id","")} (a={len(d.find_all("a"))})')
|
||||||
|
el = d
|
||||||
|
break
|
||||||
|
if el is None:
|
||||||
|
continue
|
||||||
|
print(f'\n 컨테이너: {el.name} class={el.get("class")} id={el.get("id")}')
|
||||||
|
print(f' anchors={len(el.find_all("a"))} dl={len(el.find_all("dl"))} ul={len(el.find_all("ul"))} li={len(el.find_all("li"))}')
|
||||||
|
print(' 구조 (max 80 lines):')
|
||||||
|
lines = dump_outline(el)
|
||||||
|
for l in lines:
|
||||||
|
print(' ', l)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
81
_스크립트/_probe3.py
Normal file
81
_스크립트/_probe3.py
Normal file
@ -0,0 +1,81 @@
|
|||||||
|
"""Find sitemap URLs for: 논산시, 아산시, 부여군, 서산시."""
|
||||||
|
import re
|
||||||
|
import warnings
|
||||||
|
from urllib.parse import urljoin
|
||||||
|
|
||||||
|
import requests
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
|
||||||
|
|
||||||
|
def fetch(url):
|
||||||
|
try:
|
||||||
|
r = requests.get(url, headers=H, timeout=15, verify=False, allow_redirects=True)
|
||||||
|
r.encoding = r.apparent_encoding
|
||||||
|
return r.status_code, r.url, r.text
|
||||||
|
except Exception as e:
|
||||||
|
return 0, str(e), ''
|
||||||
|
|
||||||
|
|
||||||
|
def find_links_to(html, base, keyword='사이트맵'):
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
found = []
|
||||||
|
for a in soup.find_all('a', href=True):
|
||||||
|
txt = a.get_text(strip=True)
|
||||||
|
href = a['href']
|
||||||
|
if keyword in txt or 'sitemap' in href.lower() or 'allMenu' in href:
|
||||||
|
full = urljoin(base, href)
|
||||||
|
if full != base:
|
||||||
|
found.append((txt, full))
|
||||||
|
return found
|
||||||
|
|
||||||
|
|
||||||
|
CASES = [
|
||||||
|
('논산시', 'https://nonsan.go.kr/'),
|
||||||
|
('아산시', 'https://www.asan.go.kr/main/'),
|
||||||
|
('부여군', 'https://www.buyeo.go.kr/html/kr/'),
|
||||||
|
('서산시', 'https://www.seosan.go.kr/www/index.do'),
|
||||||
|
]
|
||||||
|
|
||||||
|
for name, base in CASES:
|
||||||
|
print(f'\n=== {name} {base} ===')
|
||||||
|
code, url, html = fetch(base)
|
||||||
|
if code != 200:
|
||||||
|
print(f' main 실패: {code}')
|
||||||
|
continue
|
||||||
|
links = find_links_to(html, url)
|
||||||
|
# Deduplicate
|
||||||
|
seen = set()
|
||||||
|
uniq = []
|
||||||
|
for t, u in links:
|
||||||
|
if u not in seen:
|
||||||
|
seen.add(u)
|
||||||
|
uniq.append((t, u))
|
||||||
|
print(f' 사이트맵/allMenu 링크 후보 ({len(uniq)}):')
|
||||||
|
for t, u in uniq[:10]:
|
||||||
|
print(f' "{t}" → {u}')
|
||||||
|
# For top candidate, fetch and analyze
|
||||||
|
for t, u in uniq[:3]:
|
||||||
|
c2, u2, h2 = fetch(u)
|
||||||
|
if c2 != 200:
|
||||||
|
print(f' [{u}] HTTP {c2}')
|
||||||
|
continue
|
||||||
|
soup = BeautifulSoup(h2, 'html.parser')
|
||||||
|
# Find best container
|
||||||
|
best = (0, None, None)
|
||||||
|
for sel in ['.sitemap_grep', '.sitemap', '#sitemap', '.allMenu', '#allMenu',
|
||||||
|
'div[class*=sitemap]', 'div[class*=allMenu]', 'ul.sitemap_list',
|
||||||
|
'ul.depth1_ul', 'ul.depth1-ul', '#gnb', 'div.menu_all']:
|
||||||
|
for el in soup.select(sel):
|
||||||
|
ac = len(el.find_all('a'))
|
||||||
|
if ac > best[0]:
|
||||||
|
best = (ac, sel, el)
|
||||||
|
if best[1]:
|
||||||
|
cls = ' '.join(best[2].get('class', []))
|
||||||
|
print(f' ★ {u2}: best={best[1]!r} cls={cls!r} a={best[0]}')
|
||||||
|
else:
|
||||||
|
print(f' [{u2}]: no container')
|
||||||
99
_스크립트/_probe4.py
Normal file
99
_스크립트/_probe4.py
Normal file
@ -0,0 +1,99 @@
|
|||||||
|
"""Brute-force common sitemap paths for 아산시, 부여군, 서산시."""
|
||||||
|
import warnings
|
||||||
|
from urllib.parse import urlparse
|
||||||
|
|
||||||
|
import requests
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
|
||||||
|
|
||||||
|
def fetch(url):
|
||||||
|
try:
|
||||||
|
r = requests.get(url, headers=H, timeout=10, verify=False, allow_redirects=True)
|
||||||
|
r.encoding = r.apparent_encoding
|
||||||
|
return r.status_code, r.url, r.text
|
||||||
|
except Exception as e:
|
||||||
|
return 0, str(e), ''
|
||||||
|
|
||||||
|
|
||||||
|
def analyze(html):
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
best = (0, '', '')
|
||||||
|
for sel in ['.sitemap_grep', '.sitemap', '#sitemap', '.allMenu', '#allMenu',
|
||||||
|
'div[class*=sitemap]', 'div[class*=allMenu]', 'ul.sitemap_list',
|
||||||
|
'ul.depth1_ul', 'ul.depth1-ul', '#gnb', 'div.menu_all',
|
||||||
|
'div[class*=menu_all]']:
|
||||||
|
for el in soup.select(sel):
|
||||||
|
ac = len(el.find_all('a'))
|
||||||
|
if ac > best[0]:
|
||||||
|
best = (ac, sel, ' '.join(el.get('class', [])))
|
||||||
|
return best
|
||||||
|
|
||||||
|
|
||||||
|
PATHS = [
|
||||||
|
'/main/sitemap.do',
|
||||||
|
'/main/sub01_01.do',
|
||||||
|
'/main/sub.do?key=121',
|
||||||
|
'/main/sub.do?key=151',
|
||||||
|
'/sitemap.do',
|
||||||
|
'/main/contents.do?key=121',
|
||||||
|
'/main/contents.do?key=151',
|
||||||
|
'/main/sitemap.html',
|
||||||
|
'/main/sitemap',
|
||||||
|
'/main/menu.do',
|
||||||
|
'/main/allMenu.do',
|
||||||
|
'/main/totalMenu.do',
|
||||||
|
'/sub01_01.do',
|
||||||
|
]
|
||||||
|
|
||||||
|
ORIGINS = {
|
||||||
|
'아산시': 'https://www.asan.go.kr',
|
||||||
|
'부여군': 'https://www.buyeo.go.kr',
|
||||||
|
'서산시': 'https://www.seosan.go.kr',
|
||||||
|
}
|
||||||
|
|
||||||
|
# 부여군 prefix
|
||||||
|
BUYEO_PATHS = [
|
||||||
|
'/html/kr/sitemap.html',
|
||||||
|
'/html/kr/html/guide/0701.html',
|
||||||
|
'/html/kr/html/sub06/0701.html',
|
||||||
|
'/html/kr/html/sub07/0701.html',
|
||||||
|
'/html/kr/html/sub01/0101.html',
|
||||||
|
'/html/kr/sub.do?key=151',
|
||||||
|
]
|
||||||
|
|
||||||
|
# 서산시 prefix
|
||||||
|
SEOSAN_PATHS = [
|
||||||
|
'/www/sitemap.do',
|
||||||
|
'/www/sub.do?key=121',
|
||||||
|
'/www/sub.do?key=151',
|
||||||
|
'/www/menu_all.do',
|
||||||
|
'/www/allMenu.do',
|
||||||
|
'/www/contents.do?key=121',
|
||||||
|
'/www/contents.do?key=151',
|
||||||
|
'/www/contents.do?key=141',
|
||||||
|
'/www/contents.do?key=131',
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def try_paths(name, origin, paths):
|
||||||
|
print(f'\n=== {name} ===')
|
||||||
|
for p in paths:
|
||||||
|
url = origin + p
|
||||||
|
code, real, html = fetch(url)
|
||||||
|
if code != 200:
|
||||||
|
continue
|
||||||
|
info = analyze(html)
|
||||||
|
if info[0] >= 50:
|
||||||
|
print(f' ★ {url}: a={info[0]} sel={info[1]!r} cls={info[2]!r}')
|
||||||
|
else:
|
||||||
|
print(f' {url}: a={info[0]} (인덱싱 부족)')
|
||||||
|
|
||||||
|
|
||||||
|
try_paths('아산시', ORIGINS['아산시'], PATHS)
|
||||||
|
try_paths('부여군', ORIGINS['부여군'], BUYEO_PATHS + PATHS)
|
||||||
|
try_paths('서산시', ORIGINS['서산시'], SEOSAN_PATHS)
|
||||||
37
_스크립트/_probe5.py
Normal file
37
_스크립트/_probe5.py
Normal file
@ -0,0 +1,37 @@
|
|||||||
|
"""Inspect main pages of 아산시, 부여군, 서산시 — grep for 사이트맵 keyword anywhere."""
|
||||||
|
import re
|
||||||
|
import warnings
|
||||||
|
from urllib.parse import urljoin, urlparse
|
||||||
|
|
||||||
|
import requests
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
|
||||||
|
PAGES = {
|
||||||
|
'아산시': 'https://www.asan.go.kr/main/',
|
||||||
|
'부여군': 'https://www.buyeo.go.kr/html/kr/',
|
||||||
|
'서산시': 'https://www.seosan.go.kr/www/index.do',
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def fetch(url):
|
||||||
|
r = requests.get(url, headers=H, timeout=15, verify=False, allow_redirects=True)
|
||||||
|
r.encoding = r.apparent_encoding
|
||||||
|
return r.status_code, r.url, r.text
|
||||||
|
|
||||||
|
|
||||||
|
for name, url in PAGES.items():
|
||||||
|
print(f'\n=== {name} === {url}')
|
||||||
|
code, real, html = fetch(url)
|
||||||
|
print(f' HTTP {code}, len {len(html)}')
|
||||||
|
# Search for 사이트맵 or sitemap or allMenu
|
||||||
|
for keyword in ['사이트맵', 'sitemap', 'allMenu', '전체메뉴', 'totalMenu']:
|
||||||
|
pat = re.compile(r'[\'"][^\'"]*' + keyword + r'[^\'"]*[\'"]', re.I)
|
||||||
|
matches = pat.findall(html)[:10]
|
||||||
|
if matches:
|
||||||
|
print(f' "{keyword}" 매칭:')
|
||||||
|
for m in matches:
|
||||||
|
print(f' {m}')
|
||||||
57
_스크립트/_probe6.py
Normal file
57
_스크립트/_probe6.py
Normal file
@ -0,0 +1,57 @@
|
|||||||
|
"""Look for sitemap-like containers directly in main pages."""
|
||||||
|
import re
|
||||||
|
import warnings
|
||||||
|
from urllib.parse import urljoin
|
||||||
|
|
||||||
|
import requests
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
|
||||||
|
|
||||||
|
def fetch(url):
|
||||||
|
r = requests.get(url, headers=H, timeout=15, verify=False, allow_redirects=True)
|
||||||
|
r.encoding = r.apparent_encoding
|
||||||
|
return r.url, r.text
|
||||||
|
|
||||||
|
|
||||||
|
# 부여군 — find anchor with text 사이트맵
|
||||||
|
print('\n=== 부여군: extract sitemap link ===')
|
||||||
|
real, html = fetch('https://www.buyeo.go.kr/html/kr/')
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
for a in soup.find_all('a', href=True):
|
||||||
|
txt = a.get_text(strip=True)
|
||||||
|
if '사이트맵' in txt:
|
||||||
|
full = urljoin(real, a['href'])
|
||||||
|
print(f' 사이트맵 → {full}')
|
||||||
|
# Also look in JS for sitemap URL
|
||||||
|
for s in re.findall(r"location\.(?:href|replace)\s*=\s*['\"]([^'\"]+)['\"]", html):
|
||||||
|
if 'sitemap' in s.lower():
|
||||||
|
print(f' JS sitemap → {s}')
|
||||||
|
|
||||||
|
|
||||||
|
# 아산시 — gnb-menu inside the page
|
||||||
|
print('\n=== 아산시: scan for inline gnb menu ===')
|
||||||
|
real, html = fetch('https://www.asan.go.kr/main/')
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
# Find any container with many anchors that looks menu-like
|
||||||
|
for el in soup.select('nav, [class*=gnb], [class*=menu]'):
|
||||||
|
a_count = len(el.find_all('a'))
|
||||||
|
if a_count >= 50:
|
||||||
|
cls = ' '.join(el.get('class', []))
|
||||||
|
eid = el.get('id', '')
|
||||||
|
print(f' {el.name}#{eid}.{cls} a={a_count}')
|
||||||
|
|
||||||
|
# 서산시 — same approach
|
||||||
|
print('\n=== 서산시: scan for inline menu ===')
|
||||||
|
real, html = fetch('https://www.seosan.go.kr/www/index.do')
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
for el in soup.select('nav, [class*=gnb], [class*=menu], [class*=allMenu]'):
|
||||||
|
a_count = len(el.find_all('a'))
|
||||||
|
if a_count >= 30:
|
||||||
|
cls = ' '.join(el.get('class', []))
|
||||||
|
eid = el.get('id', '')
|
||||||
|
print(f' {el.name}#{eid}.{cls} a={a_count}')
|
||||||
76
_스크립트/_probe7.py
Normal file
76
_스크립트/_probe7.py
Normal file
@ -0,0 +1,76 @@
|
|||||||
|
"""Inspect specific selectors found in probe6 for 아산시, 서산시, 부여군."""
|
||||||
|
import warnings
|
||||||
|
from urllib.parse import urljoin
|
||||||
|
|
||||||
|
import requests
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
|
||||||
|
|
||||||
|
def fetch(url):
|
||||||
|
r = requests.get(url, headers=H, timeout=15, verify=False, allow_redirects=True)
|
||||||
|
r.encoding = r.apparent_encoding
|
||||||
|
return r.url, r.text
|
||||||
|
|
||||||
|
|
||||||
|
def outline(el, depth=0, max_lines=120, lines=None):
|
||||||
|
if lines is None:
|
||||||
|
lines = []
|
||||||
|
if len(lines) >= max_lines:
|
||||||
|
return lines
|
||||||
|
name = el.name
|
||||||
|
cls = ' '.join(el.get('class', []))
|
||||||
|
eid = el.get('id', '')
|
||||||
|
label = name
|
||||||
|
if eid:
|
||||||
|
label += f'#{eid}'
|
||||||
|
if cls:
|
||||||
|
label += '.' + cls.replace(' ', '.')
|
||||||
|
if name == 'a':
|
||||||
|
txt = el.get_text(strip=True)[:50]
|
||||||
|
href = el.get('href', '')[:80]
|
||||||
|
lines.append(' ' * depth + f'{label} "{txt}" → {href}')
|
||||||
|
else:
|
||||||
|
lines.append(' ' * depth + label)
|
||||||
|
for c in el.find_all(recursive=False):
|
||||||
|
if c.name in ('script', 'style'):
|
||||||
|
continue
|
||||||
|
outline(c, depth + 1, max_lines, lines)
|
||||||
|
if len(lines) >= max_lines:
|
||||||
|
return lines
|
||||||
|
return lines
|
||||||
|
|
||||||
|
|
||||||
|
# 부여군: find .pc_sitemap
|
||||||
|
print('\n=== 부여군: .pc_sitemap 또는 inline sitemap ===')
|
||||||
|
real, html = fetch('https://www.buyeo.go.kr/html/kr/')
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
for sel in ['.pc_sitemap', '#pc_sitemap', '.sitemap_grep', '.allMenu', '.allmenu',
|
||||||
|
'div[class*=sitemap]', 'div[class*=gnb]']:
|
||||||
|
for el in soup.select(sel):
|
||||||
|
ac = len(el.find_all('a'))
|
||||||
|
if ac >= 30:
|
||||||
|
cls = ' '.join(el.get('class', []))
|
||||||
|
print(f' ★ sel={sel} {el.name}.{cls} id={el.get("id","")} a={ac}')
|
||||||
|
|
||||||
|
# 아산시: mobile-nav.krds-gnb-mobile
|
||||||
|
print('\n=== 아산시: mobile-nav ===')
|
||||||
|
real, html = fetch('https://www.asan.go.kr/main/')
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
el = soup.select_one('nav#mobile-nav')
|
||||||
|
if el:
|
||||||
|
for line in outline(el, max_lines=80):
|
||||||
|
print(' ', line)
|
||||||
|
|
||||||
|
# 서산시: top_menu
|
||||||
|
print('\n=== 서산시: ul#top_menu ===')
|
||||||
|
real, html = fetch('https://www.seosan.go.kr/www/index.do')
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
el = soup.select_one('ul#top_menu') or soup.select_one('div.menu_wrap')
|
||||||
|
if el:
|
||||||
|
for line in outline(el, max_lines=80):
|
||||||
|
print(' ', line)
|
||||||
85
_스크립트/_probe8.py
Normal file
85
_스크립트/_probe8.py
Normal file
@ -0,0 +1,85 @@
|
|||||||
|
"""Find sitemap containers in 부여군 main page and 아산시 desktop GNB."""
|
||||||
|
import warnings
|
||||||
|
import requests
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
|
||||||
|
|
||||||
|
def fetch(url):
|
||||||
|
r = requests.get(url, headers=H, timeout=15, verify=False, allow_redirects=True)
|
||||||
|
r.encoding = r.apparent_encoding
|
||||||
|
return r.url, r.text
|
||||||
|
|
||||||
|
|
||||||
|
def outline(el, depth=0, max_lines=120, lines=None):
|
||||||
|
if lines is None:
|
||||||
|
lines = []
|
||||||
|
if len(lines) >= max_lines:
|
||||||
|
return lines
|
||||||
|
name = el.name
|
||||||
|
cls = ' '.join(el.get('class', []))
|
||||||
|
eid = el.get('id', '')
|
||||||
|
label = name
|
||||||
|
if eid:
|
||||||
|
label += f'#{eid}'
|
||||||
|
if cls:
|
||||||
|
label += '.' + cls.replace(' ', '.')
|
||||||
|
if name == 'a':
|
||||||
|
txt = el.get_text(strip=True)[:50]
|
||||||
|
href = el.get('href', '')[:80]
|
||||||
|
lines.append(' ' * depth + f'{label} "{txt}" → {href}')
|
||||||
|
else:
|
||||||
|
lines.append(' ' * depth + label)
|
||||||
|
for c in el.find_all(recursive=False):
|
||||||
|
if c.name in ('script', 'style'):
|
||||||
|
continue
|
||||||
|
outline(c, depth + 1, max_lines, lines)
|
||||||
|
if len(lines) >= max_lines:
|
||||||
|
return lines
|
||||||
|
return lines
|
||||||
|
|
||||||
|
|
||||||
|
# 부여군 — scan ALL container types for menu-like content
|
||||||
|
print('=== 부여군: scan all containers ===')
|
||||||
|
real, html = fetch('https://www.buyeo.go.kr/html/kr/')
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
# Find divs with most anchors
|
||||||
|
candidates = []
|
||||||
|
for d in soup.find_all(['div', 'nav', 'ul']):
|
||||||
|
a_count = len(d.find_all('a'))
|
||||||
|
if a_count >= 100:
|
||||||
|
cls = ' '.join(d.get('class', []))
|
||||||
|
eid = d.get('id', '')
|
||||||
|
candidates.append((a_count, d.name, eid, cls, d))
|
||||||
|
candidates.sort(reverse=True, key=lambda x: x[0])
|
||||||
|
for ac, name, eid, cls, _ in candidates[:10]:
|
||||||
|
print(f' {name}#{eid}.{cls[:60]} a={ac}')
|
||||||
|
|
||||||
|
# Outline top container
|
||||||
|
if candidates:
|
||||||
|
print('\n Top container outline:')
|
||||||
|
for line in outline(candidates[0][4], max_lines=60):
|
||||||
|
print(' ', line)
|
||||||
|
|
||||||
|
# 아산시 — desktop GNB
|
||||||
|
print('\n\n=== 아산시: desktop GNB ===')
|
||||||
|
real, html = fetch('https://www.asan.go.kr/main/')
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
# Look for desktop GNB sub-lists
|
||||||
|
for el in soup.select('[id^=mGnb-anchor], .gnb-sub-list, .submenu-wrap, .gnb-wrap'):
|
||||||
|
ac = len(el.find_all('a'))
|
||||||
|
if ac >= 30:
|
||||||
|
cls = ' '.join(el.get('class', []))
|
||||||
|
eid = el.get('id', '')
|
||||||
|
print(f' {el.name}#{eid}.{cls} a={ac}')
|
||||||
|
|
||||||
|
# Try the main desktop nav specifically
|
||||||
|
el = soup.select_one('nav.krds-gnb:not(#mobile-nav)')
|
||||||
|
if el:
|
||||||
|
print('\n Desktop nav outline:')
|
||||||
|
for line in outline(el, max_lines=80):
|
||||||
|
print(' ', line)
|
||||||
110
_스크립트/_probe9.py
Normal file
110
_스크립트/_probe9.py
Normal file
@ -0,0 +1,110 @@
|
|||||||
|
"""Detailed structural look at 논산시, 아산시, 부여군 main sitemaps."""
|
||||||
|
import warnings
|
||||||
|
import requests
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
|
||||||
|
|
||||||
|
def fetch(url):
|
||||||
|
r = requests.get(url, headers=H, timeout=15, verify=False, allow_redirects=True)
|
||||||
|
r.encoding = r.apparent_encoding
|
||||||
|
return r.url, r.text
|
||||||
|
|
||||||
|
|
||||||
|
def outline(el, depth=0, max_lines=120, lines=None):
|
||||||
|
if lines is None:
|
||||||
|
lines = []
|
||||||
|
if len(lines) >= max_lines:
|
||||||
|
return lines
|
||||||
|
name = el.name
|
||||||
|
cls = ' '.join(el.get('class', []))
|
||||||
|
eid = el.get('id', '')
|
||||||
|
label = name
|
||||||
|
if eid:
|
||||||
|
label += f'#{eid}'
|
||||||
|
if cls:
|
||||||
|
label += '.' + cls.replace(' ', '.')
|
||||||
|
if name == 'a':
|
||||||
|
txt = el.get_text(strip=True)[:50]
|
||||||
|
href = el.get('href', '')[:80]
|
||||||
|
lines.append(' ' * depth + f'{label} "{txt}" → {href}')
|
||||||
|
elif name == 'button':
|
||||||
|
txt = el.get_text(strip=True)[:50]
|
||||||
|
lines.append(' ' * depth + f'{label} BTN "{txt}"')
|
||||||
|
else:
|
||||||
|
lines.append(' ' * depth + label)
|
||||||
|
for c in el.find_all(recursive=False):
|
||||||
|
if c.name in ('script', 'style'):
|
||||||
|
continue
|
||||||
|
outline(c, depth + 1, max_lines, lines)
|
||||||
|
if len(lines) >= max_lines:
|
||||||
|
return lines
|
||||||
|
return lines
|
||||||
|
|
||||||
|
|
||||||
|
# 논산시 - look at div.sitemap.type1
|
||||||
|
print('=== 논산시 div.sitemap.type1 ===')
|
||||||
|
real, html = fetch('https://nonsan.go.kr/kor/html/sub07/0701.html')
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
el = soup.select_one('div.sitemap.type1') or soup.select_one('div.sitemap')
|
||||||
|
if el:
|
||||||
|
cls = ' '.join(el.get('class', []))
|
||||||
|
print(f' Container: {el.name}.{cls} a={len(el.find_all("a"))}')
|
||||||
|
for line in outline(el, max_lines=70):
|
||||||
|
print(' ', line)
|
||||||
|
|
||||||
|
# 부여군 - try /html/kr/sitemap_pop.html or similar
|
||||||
|
print('\n\n=== 부여군 alternate sitemap probes ===')
|
||||||
|
for path in [
|
||||||
|
'/html/kr/sitemap_pop.html',
|
||||||
|
'/html/kr/sitemap.html',
|
||||||
|
'/html/kr/html/sub06/0601.html',
|
||||||
|
'/html/kr/html/sub05/0501.html',
|
||||||
|
'/html/kr/html/sub04/0401.html',
|
||||||
|
'/html/kr/html/sub09/0901.html',
|
||||||
|
'/html/kr/popup/sitemap.html',
|
||||||
|
'/html/kr/include/sitemap.html',
|
||||||
|
]:
|
||||||
|
url = f'https://www.buyeo.go.kr{path}'
|
||||||
|
try:
|
||||||
|
r = requests.get(url, headers=H, timeout=8, verify=False, allow_redirects=True)
|
||||||
|
r.encoding = r.apparent_encoding
|
||||||
|
if r.status_code == 200:
|
||||||
|
s = BeautifulSoup(r.text, 'html.parser')
|
||||||
|
best_count = 0
|
||||||
|
best_sel = ''
|
||||||
|
for sel in ['.sitemap', '#sitemap', '.sitemap_grep', '.allMenu', '.pc_sitemap', 'div[class*=sitemap]']:
|
||||||
|
for el in s.select(sel):
|
||||||
|
ac = len(el.find_all('a'))
|
||||||
|
if ac > best_count:
|
||||||
|
best_count = ac
|
||||||
|
best_sel = sel
|
||||||
|
print(f' {path}: HTTP 200, best_sel={best_sel}, a={best_count}')
|
||||||
|
except Exception as e:
|
||||||
|
pass
|
||||||
|
|
||||||
|
# 부여군 main page - look for inline pc_sitemap or modal
|
||||||
|
print('\n\n=== 부여군 main page: look for inline allMenu / pc_sitemap modal ===')
|
||||||
|
real, html = fetch('https://www.buyeo.go.kr/html/kr/')
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
# Search for elements with id or class containing "sitemap"
|
||||||
|
for el in soup.select('[class*=sitemap], [id*=sitemap], [id*=allMenu], [class*=allMenu], [id*=allmenu], [class*=allmenu]'):
|
||||||
|
cls = ' '.join(el.get('class', []))
|
||||||
|
eid = el.get('id', '')
|
||||||
|
ac = len(el.find_all('a'))
|
||||||
|
print(f' {el.name}#{eid}.{cls} a={ac}')
|
||||||
|
|
||||||
|
# 아산시 - look at the .gnb-menu inline structure (use sectioned anchors mGnb-anchor1..6)
|
||||||
|
print('\n\n=== 아산시: all gnb-sub-list sections combined ===')
|
||||||
|
real, html = fetch('https://www.asan.go.kr/main/')
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
all_a = []
|
||||||
|
for sec in soup.select('div[id^=mGnb-anchor]'):
|
||||||
|
eid = sec.get('id', '')
|
||||||
|
section_a = sec.find_all('a', href=True)
|
||||||
|
print(f' {eid}: a={len(section_a)}')
|
||||||
|
for a in section_a[:3]:
|
||||||
|
print(f' "{a.get_text(strip=True)[:40]}" → {a["href"][:80]}')
|
||||||
35
_스크립트/_probe_brcn.py
Normal file
35
_스크립트/_probe_brcn.py
Normal file
@ -0,0 +1,35 @@
|
|||||||
|
"""Look at raw HTML around h4.site01 in 보령시 to find category name."""
|
||||||
|
import re
|
||||||
|
import warnings
|
||||||
|
import requests
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
|
||||||
|
for name, url in [('보령시', 'https://www.brcn.go.kr/kor/sitemap_11.do'),
|
||||||
|
('서천군', 'https://www.seocheon.go.kr/kor/sitemap_11.do')]:
|
||||||
|
print(f'\n=== {name} === {url}')
|
||||||
|
r = requests.get(url, headers=H, timeout=15, verify=False, allow_redirects=True)
|
||||||
|
r.encoding = r.apparent_encoding
|
||||||
|
html = r.text
|
||||||
|
# Find h4.siteNN occurrences
|
||||||
|
for m in re.finditer(r'<h4[^>]*site\d+[^>]*>([\s\S]*?)</h4>', html):
|
||||||
|
inner = re.sub(r'\s+', ' ', m.group(1)).strip()[:200]
|
||||||
|
full = re.sub(r'\s+', ' ', m.group(0)).strip()[:200]
|
||||||
|
print(f' h4: {full}')
|
||||||
|
# Find gnb anchors
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
gnb = soup.select_one('#gnb, #tm, nav#topmenu, nav.gnb, .gnb_wrap, .menu_wrap')
|
||||||
|
if gnb:
|
||||||
|
a_top = [a.get_text(strip=True) for a in gnb.select('> ul > li > a') or gnb.select('ul > li > a.th_1st, ul > li > a.depth1_ti, ul > li > a.first, ul > li > a')]
|
||||||
|
print(f' GNB top items: {a_top[:10]}')
|
||||||
|
# Find the top-level menu anchors (any way)
|
||||||
|
print(' Possible main category anchors:')
|
||||||
|
for cls_pattern in ['ov', 'th_1st', 'depth1_ti', 'first', 'gnb-main-trigger']:
|
||||||
|
for a in soup.find_all('a', class_=cls_pattern):
|
||||||
|
t = a.get_text(strip=True)
|
||||||
|
if t and len(t) < 30:
|
||||||
|
print(f' .{cls_pattern}: "{t}"')
|
||||||
|
break
|
||||||
157
_스크립트/_probe_chungbuk.py
Normal file
157
_스크립트/_probe_chungbuk.py
Normal file
@ -0,0 +1,157 @@
|
|||||||
|
"""충청북도 11개 시·군 사이트맵 URL/패턴 탐지."""
|
||||||
|
import re
|
||||||
|
import warnings
|
||||||
|
from urllib.parse import urljoin, urlparse
|
||||||
|
from concurrent.futures import ThreadPoolExecutor, as_completed
|
||||||
|
|
||||||
|
import requests
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
|
||||||
|
CITIES = [
|
||||||
|
('괴산군', 'https://www.goesan.go.kr/www/index.do'),
|
||||||
|
('단양군', 'https://www.danyang.go.kr/dy21/1'),
|
||||||
|
('보은군', 'https://www.boeun.go.kr/www/index.do'),
|
||||||
|
('영동군', 'https://www.yd21.go.kr/'),
|
||||||
|
('옥천군', 'https://www.oc.go.kr/www/'),
|
||||||
|
('음성군', 'https://www.eumseong.go.kr/www/index.do'),
|
||||||
|
('제천시', 'https://www.jecheon.go.kr/www/index.do'),
|
||||||
|
('증평군', 'https://www.jp.go.kr/kor.do'),
|
||||||
|
('진천군', 'https://www.jincheon.go.kr/home/intro.do'),
|
||||||
|
('청주시', 'https://www.cheongju.go.kr/www/index.do'),
|
||||||
|
('충주시', 'https://www.chungju.go.kr/www/index.do'),
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def fetch(url, timeout=12):
|
||||||
|
try:
|
||||||
|
r = requests.get(url, headers=H, timeout=timeout, verify=False, allow_redirects=True)
|
||||||
|
r.encoding = r.apparent_encoding
|
||||||
|
if r.status_code == 200:
|
||||||
|
return r.url, r.text
|
||||||
|
except Exception as e:
|
||||||
|
return None, f'ERR: {e}'
|
||||||
|
return None, f'HTTP {r.status_code}'
|
||||||
|
|
||||||
|
|
||||||
|
def find_sitemap_anchor(html, base):
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
found = []
|
||||||
|
for a in soup.find_all('a', href=True):
|
||||||
|
txt = a.get_text(strip=True)
|
||||||
|
href = a['href']
|
||||||
|
if not (txt and href):
|
||||||
|
continue
|
||||||
|
if '사이트맵' in txt or '전체메뉴' in txt or 'sitemap' in href.lower() or 'allMenu' in href:
|
||||||
|
full = urljoin(base, href)
|
||||||
|
found.append((txt, full))
|
||||||
|
seen = set()
|
||||||
|
uniq = []
|
||||||
|
for t, u in found:
|
||||||
|
if u not in seen and u != base:
|
||||||
|
seen.add(u)
|
||||||
|
uniq.append((t, u))
|
||||||
|
return uniq
|
||||||
|
|
||||||
|
|
||||||
|
def analyze_sitemap_page(html):
|
||||||
|
"""Score candidate containers by anchor count and class hints."""
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
candidates = []
|
||||||
|
# Common containers
|
||||||
|
for sel in [
|
||||||
|
'div.sitemap.type1', 'div.sitemap.type2', 'div.sitemap',
|
||||||
|
'div.sitemap_grep', 'div.amThum', 'ul.sitemap_list',
|
||||||
|
'#sitemap', '#contents ul.sitemap', 'ul.sitemap',
|
||||||
|
'ul.depth1_ul', 'ul.depth1-ul', 'ul.depth1',
|
||||||
|
'ul.top_menu', '#gnb', 'nav#gnb', 'nav.gnb',
|
||||||
|
'.allmenu', '.allMenu', '#allMenu',
|
||||||
|
'div.menu_all', 'div.totalMenu',
|
||||||
|
]:
|
||||||
|
for el in soup.select(sel):
|
||||||
|
ac = len(el.find_all('a'))
|
||||||
|
if ac < 30:
|
||||||
|
continue
|
||||||
|
cls = ' '.join(el.get('class', []))
|
||||||
|
eid = el.get('id', '')
|
||||||
|
candidates.append((ac, sel, f'{el.name}#{eid}.{cls}'))
|
||||||
|
# Fallback: search any container by id/class containing 'sitemap' or 'allMenu'
|
||||||
|
for el in soup.find_all(True, class_=True):
|
||||||
|
cls = ' '.join(el.get('class', []))
|
||||||
|
if re.search(r'\b(sitemap|allmenu|amthum|depth1)\b', cls, re.I):
|
||||||
|
ac = len(el.find_all('a'))
|
||||||
|
if 30 <= ac <= 2000:
|
||||||
|
candidates.append((ac, f'class~{cls[:30]}', f'{el.name}.{cls}'))
|
||||||
|
candidates.sort(reverse=True)
|
||||||
|
return candidates[:5]
|
||||||
|
|
||||||
|
|
||||||
|
def probe(name, base):
|
||||||
|
print(f'\n=== {name} === {base}')
|
||||||
|
real, html = fetch(base)
|
||||||
|
if not real:
|
||||||
|
print(f' 메인 실패: {html}')
|
||||||
|
return name, None
|
||||||
|
print(f' 메인 OK: {real}')
|
||||||
|
# 1) Find sitemap links from main page
|
||||||
|
links = find_sitemap_anchor(html, real)
|
||||||
|
print(f' 사이트맵 링크 후보: {len(links)}')
|
||||||
|
for t, u in links[:5]:
|
||||||
|
print(f' "{t}" → {u}')
|
||||||
|
# 2) Common paths to try
|
||||||
|
parsed = urlparse(base)
|
||||||
|
origin = f'{parsed.scheme}://{parsed.netloc}'
|
||||||
|
common = [
|
||||||
|
'/www/sitemap.do', '/www/sub.do?key=121', '/www/contents.do?key=121',
|
||||||
|
'/kor/sitemap.do', '/kor/sitemap_11.do', '/kor/sitemap_1.do',
|
||||||
|
'/sitemap.do', '/sitemap.html',
|
||||||
|
'/www/sitemap/', '/main/sitemap.do',
|
||||||
|
'/home/sitemap.do', '/dy21/sitemap.do',
|
||||||
|
'/www/cms/sitemap.do',
|
||||||
|
]
|
||||||
|
candidates = [u for _, u in links] + [origin + p for p in common]
|
||||||
|
seen = set()
|
||||||
|
best = None
|
||||||
|
for c in candidates:
|
||||||
|
if c in seen:
|
||||||
|
continue
|
||||||
|
seen.add(c)
|
||||||
|
u2, h2 = fetch(c, timeout=10)
|
||||||
|
if not u2:
|
||||||
|
continue
|
||||||
|
info = analyze_sitemap_page(h2)
|
||||||
|
if info:
|
||||||
|
top = info[0]
|
||||||
|
if best is None or top[0] > best[1][0]:
|
||||||
|
best = (c, top, info)
|
||||||
|
if best:
|
||||||
|
url, top, info = best
|
||||||
|
print(f' ★ 사이트맵 URL: {url}')
|
||||||
|
for ac, sel, desc in info:
|
||||||
|
print(f' {sel}: a={ac} {desc[:80]}')
|
||||||
|
return name, {'url': url, 'best_sel': info[0][1], 'a_count': info[0][0]}
|
||||||
|
print(' 사이트맵 못 찾음')
|
||||||
|
return name, None
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
results = {}
|
||||||
|
with ThreadPoolExecutor(max_workers=6) as ex:
|
||||||
|
futs = {ex.submit(probe, n, b): n for n, b in CITIES}
|
||||||
|
for f in as_completed(futs):
|
||||||
|
n, info = f.result()
|
||||||
|
results[n] = info
|
||||||
|
print('\n=== 요약 ===')
|
||||||
|
for n, _ in CITIES:
|
||||||
|
info = results.get(n)
|
||||||
|
if info:
|
||||||
|
print(f' {n}: {info["url"]} [{info["best_sel"]}, a={info["a_count"]}]')
|
||||||
|
else:
|
||||||
|
print(f' {n}: NOT FOUND')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
135
_스크립트/_probe_chungbuk2.py
Normal file
135
_스크립트/_probe_chungbuk2.py
Normal file
@ -0,0 +1,135 @@
|
|||||||
|
"""충청북도 추가 분석:
|
||||||
|
1) depth1 패턴 구조 outline (괴산·청주·충주 샘플)
|
||||||
|
2) 단양·보은·영동·진천 사이트맵 재탐색
|
||||||
|
"""
|
||||||
|
import re
|
||||||
|
import warnings
|
||||||
|
import requests
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
|
||||||
|
|
||||||
|
def fetch(url, timeout=15):
|
||||||
|
try:
|
||||||
|
r = requests.get(url, headers=H, timeout=timeout, verify=False, allow_redirects=True)
|
||||||
|
r.encoding = r.apparent_encoding
|
||||||
|
return r.status_code, r.url, r.text
|
||||||
|
except Exception as e:
|
||||||
|
return 0, str(e), ''
|
||||||
|
|
||||||
|
|
||||||
|
def outline(el, d=0, max_lines=60, lines=None):
|
||||||
|
if lines is None: lines = []
|
||||||
|
if len(lines) >= max_lines: return lines
|
||||||
|
name = el.name
|
||||||
|
cls = ' '.join(el.get('class', []))
|
||||||
|
eid = el.get('id', '')
|
||||||
|
lbl = name
|
||||||
|
if eid: lbl += f'#{eid}'
|
||||||
|
if cls: lbl += '.' + cls.replace(' ', '.')
|
||||||
|
if name == 'a':
|
||||||
|
t = el.get_text(strip=True)[:40]
|
||||||
|
h = el.get('href', '')[:80]
|
||||||
|
lines.append(' ' * d + f'{lbl} "{t}" → {h}')
|
||||||
|
else:
|
||||||
|
lines.append(' ' * d + lbl)
|
||||||
|
for c in el.find_all(recursive=False):
|
||||||
|
if c.name in ('script', 'style'): continue
|
||||||
|
outline(c, d+1, max_lines, lines)
|
||||||
|
if len(lines) >= max_lines: return lines
|
||||||
|
return lines
|
||||||
|
|
||||||
|
|
||||||
|
# (1) Depth1 패턴 outline
|
||||||
|
print('='*70)
|
||||||
|
print('Depth1 pattern — 괴산군')
|
||||||
|
print('='*70)
|
||||||
|
_, _, html = fetch('https://www.goesan.go.kr/www/sitemap.do?key=28')
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
for sel in ['div.depth1', 'div.depth.depth1', '#sitemap', '.sitemap']:
|
||||||
|
el = soup.select_one(sel)
|
||||||
|
if el:
|
||||||
|
print(f'\n>>> selector: {sel} (a={len(el.find_all("a"))})')
|
||||||
|
for line in outline(el, max_lines=50):
|
||||||
|
print(' ', line)
|
||||||
|
break
|
||||||
|
|
||||||
|
print('\n' + '='*70)
|
||||||
|
print('Depth1 pattern — 청주시')
|
||||||
|
print('='*70)
|
||||||
|
_, _, html = fetch('https://www.cheongju.go.kr/www/sitemap.do?key=589')
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
for sel in ['#sitemap div.sitemap', 'div#sitemap.sitemap', 'div.sitemap']:
|
||||||
|
el = soup.select_one(sel)
|
||||||
|
if el:
|
||||||
|
print(f'\n>>> selector: {sel} (a={len(el.find_all("a"))})')
|
||||||
|
for line in outline(el, max_lines=50):
|
||||||
|
print(' ', line)
|
||||||
|
break
|
||||||
|
|
||||||
|
print('\n' + '='*70)
|
||||||
|
print('Depth1 pattern — 충주시')
|
||||||
|
print('='*70)
|
||||||
|
_, _, html = fetch('https://www.chungju.go.kr/www/sub.do?key=692')
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
for sel in ['#sitemap']:
|
||||||
|
el = soup.select_one(sel)
|
||||||
|
if el:
|
||||||
|
print(f'\n>>> selector: {sel} (a={len(el.find_all("a"))})')
|
||||||
|
for line in outline(el, max_lines=50):
|
||||||
|
print(' ', line)
|
||||||
|
break
|
||||||
|
|
||||||
|
|
||||||
|
# (2) 실패한 사이트들 재탐색
|
||||||
|
print('\n\n' + '='*70)
|
||||||
|
print('실패 사이트 재탐색')
|
||||||
|
print('='*70)
|
||||||
|
for name, url in [
|
||||||
|
('단양군', 'https://www.danyang.go.kr/dy21/1'),
|
||||||
|
('보은군', 'https://www.boeun.go.kr/www/index.do'),
|
||||||
|
('영동군', 'https://www.yd21.go.kr/'),
|
||||||
|
('진천군', 'https://www.jincheon.go.kr/home/intro.do'),
|
||||||
|
]:
|
||||||
|
print(f'\n--- {name} {url} ---')
|
||||||
|
code, real, html = fetch(url)
|
||||||
|
if code != 200:
|
||||||
|
print(f' 메인 실패: {code} {real}')
|
||||||
|
continue
|
||||||
|
print(f' 메인 OK: {real}')
|
||||||
|
# Find sitemap link
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
cands = []
|
||||||
|
for a in soup.find_all('a', href=True):
|
||||||
|
txt = a.get_text(strip=True)
|
||||||
|
if '사이트맵' in txt or '전체메뉴' in txt or '누리집' in txt or 'sitemap' in (a.get('href','')+txt).lower():
|
||||||
|
cands.append((txt, a['href']))
|
||||||
|
# Deduplicate
|
||||||
|
seen = set()
|
||||||
|
uniq = []
|
||||||
|
from urllib.parse import urljoin
|
||||||
|
for t, h in cands:
|
||||||
|
full = urljoin(real, h)
|
||||||
|
if full not in seen and full != real:
|
||||||
|
seen.add(full)
|
||||||
|
uniq.append((t, full))
|
||||||
|
for t, u in uniq[:5]:
|
||||||
|
print(f' 사이트맵 후보: "{t}" → {u}')
|
||||||
|
# Test first candidate
|
||||||
|
if uniq:
|
||||||
|
c, c_url = uniq[0]
|
||||||
|
c2, r2, h2 = fetch(c_url)
|
||||||
|
if c2 == 200:
|
||||||
|
s2 = BeautifulSoup(h2, 'html.parser')
|
||||||
|
best = (0, '', '')
|
||||||
|
for sel in ['div.depth1', 'div.depth.depth1', 'ul.depth1_ul', 'ul.depth1-ul',
|
||||||
|
'#sitemap', 'div.sitemap_grep', 'ul.sitemap', 'div.sitemap',
|
||||||
|
'div.sitemap_11', 'div.amThum']:
|
||||||
|
for el in s2.select(sel):
|
||||||
|
ac = len(el.find_all('a'))
|
||||||
|
if ac > best[0]:
|
||||||
|
best = (ac, sel, ' '.join(el.get('class', [])))
|
||||||
|
print(f' 사이트맵 분석: best_sel={best[1]} cls={best[2]} a={best[0]}')
|
||||||
213
_스크립트/_probe_chungbuk3.py
Normal file
213
_스크립트/_probe_chungbuk3.py
Normal file
@ -0,0 +1,213 @@
|
|||||||
|
"""실패한 4개 사이트 (단양·보은·영동·진천) 심층 탐색."""
|
||||||
|
import re
|
||||||
|
import ssl
|
||||||
|
import warnings
|
||||||
|
import requests
|
||||||
|
from requests.adapters import HTTPAdapter
|
||||||
|
from urllib3.util.ssl_ import create_urllib3_context
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
|
||||||
|
|
||||||
|
class WeakSSLAdapter(HTTPAdapter):
|
||||||
|
"""레거시 SSL/TLS handshake를 허용하는 어댑터 (영동군 등 구형 SSL)."""
|
||||||
|
def init_poolmanager(self, *args, **kwargs):
|
||||||
|
ctx = create_urllib3_context()
|
||||||
|
ctx.set_ciphers('DEFAULT@SECLEVEL=0')
|
||||||
|
ctx.options |= 0x4 # ssl.OP_LEGACY_SERVER_CONNECT (Python 3.12+)
|
||||||
|
ctx.check_hostname = False
|
||||||
|
ctx.verify_mode = ssl.CERT_NONE
|
||||||
|
kwargs['ssl_context'] = ctx
|
||||||
|
return super().init_poolmanager(*args, **kwargs)
|
||||||
|
|
||||||
|
|
||||||
|
def make_session(weak_ssl=False):
|
||||||
|
s = requests.Session()
|
||||||
|
s.headers.update(H)
|
||||||
|
if weak_ssl:
|
||||||
|
s.mount('https://', WeakSSLAdapter())
|
||||||
|
return s
|
||||||
|
|
||||||
|
|
||||||
|
def fetch(url, session=None, timeout=20):
|
||||||
|
s = session or requests.Session()
|
||||||
|
if not session:
|
||||||
|
s.headers.update(H)
|
||||||
|
try:
|
||||||
|
r = s.get(url, timeout=timeout, verify=False, allow_redirects=True)
|
||||||
|
r.encoding = r.apparent_encoding
|
||||||
|
return r.status_code, r.url, r.text
|
||||||
|
except Exception as e:
|
||||||
|
return 0, str(e), ''
|
||||||
|
|
||||||
|
|
||||||
|
def outline(el, d=0, max_lines=40, lines=None):
|
||||||
|
if lines is None: lines = []
|
||||||
|
if len(lines) >= max_lines: return lines
|
||||||
|
name = el.name
|
||||||
|
cls = ' '.join(el.get('class', []))
|
||||||
|
eid = el.get('id', '')
|
||||||
|
lbl = name
|
||||||
|
if eid: lbl += f'#{eid}'
|
||||||
|
if cls: lbl += '.' + cls.replace(' ', '.')
|
||||||
|
if name == 'a':
|
||||||
|
t = el.get_text(strip=True)[:40]
|
||||||
|
h = el.get('href', '')[:80]
|
||||||
|
lines.append(' ' * d + f'{lbl} "{t}" → {h}')
|
||||||
|
else:
|
||||||
|
lines.append(' ' * d + lbl)
|
||||||
|
for c in el.find_all(recursive=False):
|
||||||
|
if c.name in ('script', 'style'): continue
|
||||||
|
outline(c, d+1, max_lines, lines)
|
||||||
|
if len(lines) >= max_lines: return lines
|
||||||
|
return lines
|
||||||
|
|
||||||
|
|
||||||
|
def analyze(html):
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
best = (0, '', '', None)
|
||||||
|
for sel in [
|
||||||
|
'div.depth.depth1', 'div.depth1', '#sitemap div.site_map_col', '#sitemap',
|
||||||
|
'div.sitemap_grep', 'ul.sitemap_list', 'div.sitemap_box',
|
||||||
|
'ul.depth1_ul', 'ul.depth1-ul', 'ul.depth1',
|
||||||
|
'ul.sitemap', 'div.sitemap', 'div.amThum', '.allMenu', 'div.menu_all',
|
||||||
|
'nav#gnb', 'nav.gnb', '#gnb',
|
||||||
|
]:
|
||||||
|
for el in soup.select(sel):
|
||||||
|
ac = len(el.find_all('a'))
|
||||||
|
if ac > best[0]:
|
||||||
|
cls = ' '.join(el.get('class', []))
|
||||||
|
best = (ac, sel, cls, el)
|
||||||
|
return best
|
||||||
|
|
||||||
|
|
||||||
|
# 단양군 - try /dy21/98
|
||||||
|
print('='*70)
|
||||||
|
print('단양군 — /dy21/98')
|
||||||
|
print('='*70)
|
||||||
|
sess = make_session()
|
||||||
|
code, real, html = fetch('https://www.danyang.go.kr/dy21/98', sess)
|
||||||
|
print(f' HTTP {code}, len {len(html)}')
|
||||||
|
if code == 200:
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
# Search for the sitemap container
|
||||||
|
print(f' All elements w/ a>=50:')
|
||||||
|
for el in soup.find_all(['div', 'ul', 'nav']):
|
||||||
|
ac = len(el.find_all('a'))
|
||||||
|
if 50 <= ac:
|
||||||
|
cls = ' '.join(el.get('class', []))[:50]
|
||||||
|
eid = el.get('id', '')
|
||||||
|
print(f' {el.name}#{eid}.{cls} a={ac}')
|
||||||
|
best = analyze(html)
|
||||||
|
if best[3]:
|
||||||
|
print(f' Best: {best[1]} cls={best[2]} a={best[0]}')
|
||||||
|
for line in outline(best[3], max_lines=40):
|
||||||
|
print(' ', line)
|
||||||
|
|
||||||
|
|
||||||
|
# 진천군 - try /home/main.do
|
||||||
|
print('\n' + '='*70)
|
||||||
|
print('진천군 — /home/main.do')
|
||||||
|
print('='*70)
|
||||||
|
code, real, html = fetch('https://www.jincheon.go.kr/home/main.do', sess)
|
||||||
|
print(f' HTTP {code}, len {len(html)}')
|
||||||
|
if code == 200:
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
cands = []
|
||||||
|
for a in soup.find_all('a', href=True):
|
||||||
|
txt = a.get_text(strip=True)
|
||||||
|
if '사이트맵' in txt or '전체메뉴' in txt or '누리집' in txt:
|
||||||
|
from urllib.parse import urljoin
|
||||||
|
cands.append((txt, urljoin(real, a['href'])))
|
||||||
|
print(' 사이트맵 후보:')
|
||||||
|
for t, u in cands[:8]:
|
||||||
|
print(f' "{t}" → {u}')
|
||||||
|
# Common paths
|
||||||
|
from urllib.parse import urlparse
|
||||||
|
parsed = urlparse(real)
|
||||||
|
origin = f'{parsed.scheme}://{parsed.netloc}'
|
||||||
|
test_urls = [u for _, u in cands] + [
|
||||||
|
origin + '/home/sitemap.do',
|
||||||
|
origin + '/home/contents.do?key=121',
|
||||||
|
origin + '/home/sub.do?key=121',
|
||||||
|
]
|
||||||
|
for url in test_urls:
|
||||||
|
c2, r2, h2 = fetch(url, sess)
|
||||||
|
if c2 != 200:
|
||||||
|
continue
|
||||||
|
best = analyze(h2)
|
||||||
|
if best[0] >= 50:
|
||||||
|
print(f' ★ {url}: {best[1]} cls={best[2]} a={best[0]}')
|
||||||
|
|
||||||
|
|
||||||
|
# 보은군 - retry
|
||||||
|
print('\n' + '='*70)
|
||||||
|
print('보은군 — retry')
|
||||||
|
print('='*70)
|
||||||
|
code, real, html = fetch('https://www.boeun.go.kr/www/index.do', sess)
|
||||||
|
print(f' HTTP {code}, len {len(html)}')
|
||||||
|
if code == 200:
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
cands = []
|
||||||
|
for a in soup.find_all('a', href=True):
|
||||||
|
txt = a.get_text(strip=True)
|
||||||
|
if '사이트맵' in txt or '전체메뉴' in txt or '누리집 지도' in txt:
|
||||||
|
from urllib.parse import urljoin
|
||||||
|
cands.append((txt, urljoin(real, a['href'])))
|
||||||
|
print(' 사이트맵 후보:')
|
||||||
|
for t, u in cands[:8]:
|
||||||
|
print(f' "{t}" → {u}')
|
||||||
|
# Try probable paths
|
||||||
|
from urllib.parse import urlparse
|
||||||
|
parsed = urlparse(real)
|
||||||
|
origin = f'{parsed.scheme}://{parsed.netloc}'
|
||||||
|
test_urls = [u for _, u in cands] + [
|
||||||
|
origin + '/www/sitemap.do',
|
||||||
|
origin + '/www/sub.do?key=121',
|
||||||
|
]
|
||||||
|
for url in set(test_urls):
|
||||||
|
c2, r2, h2 = fetch(url, sess)
|
||||||
|
if c2 != 200:
|
||||||
|
continue
|
||||||
|
best = analyze(h2)
|
||||||
|
if best[0] >= 50:
|
||||||
|
print(f' ★ {url}: {best[1]} cls={best[2]} a={best[0]}')
|
||||||
|
|
||||||
|
|
||||||
|
# 영동군 - try with weak SSL adapter
|
||||||
|
print('\n' + '='*70)
|
||||||
|
print('영동군 — weak SSL adapter')
|
||||||
|
print('='*70)
|
||||||
|
sess_weak = make_session(weak_ssl=True)
|
||||||
|
code, real, html = fetch('https://www.yd21.go.kr/', sess_weak)
|
||||||
|
print(f' HTTP {code}, len {len(html)}')
|
||||||
|
if code == 200:
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
cands = []
|
||||||
|
for a in soup.find_all('a', href=True):
|
||||||
|
txt = a.get_text(strip=True)
|
||||||
|
if '사이트맵' in txt or '전체메뉴' in txt or '누리집' in txt:
|
||||||
|
from urllib.parse import urljoin
|
||||||
|
cands.append((txt, urljoin(real, a['href'])))
|
||||||
|
print(' 사이트맵 후보:')
|
||||||
|
for t, u in cands[:8]:
|
||||||
|
print(f' "{t}" → {u}')
|
||||||
|
from urllib.parse import urlparse
|
||||||
|
parsed = urlparse(real)
|
||||||
|
origin = f'{parsed.scheme}://{parsed.netloc}'
|
||||||
|
test_urls = [u for _, u in cands] + [
|
||||||
|
origin + '/kor/sitemap.do',
|
||||||
|
origin + '/sitemap.html',
|
||||||
|
origin + '/sitemap.do',
|
||||||
|
origin + '/contents/contents.html?cid=2151',
|
||||||
|
]
|
||||||
|
for url in set(test_urls):
|
||||||
|
c2, r2, h2 = fetch(url, sess_weak)
|
||||||
|
if c2 != 200:
|
||||||
|
continue
|
||||||
|
best = analyze(h2)
|
||||||
|
if best[0] >= 50:
|
||||||
|
print(f' ★ {url}: {best[1]} cls={best[2]} a={best[0]}')
|
||||||
86
_스크립트/_probe_chungbuk4.py
Normal file
86
_스크립트/_probe_chungbuk4.py
Normal file
@ -0,0 +1,86 @@
|
|||||||
|
"""단양·진천 구조 outline + 보은군 alternative."""
|
||||||
|
import warnings
|
||||||
|
import socket
|
||||||
|
import requests
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
|
||||||
|
|
||||||
|
def fetch(url, timeout=20):
|
||||||
|
try:
|
||||||
|
r = requests.get(url, headers=H, timeout=timeout, verify=False, allow_redirects=True)
|
||||||
|
r.encoding = r.apparent_encoding
|
||||||
|
return r.status_code, r.url, r.text
|
||||||
|
except Exception as e:
|
||||||
|
return 0, str(e), ''
|
||||||
|
|
||||||
|
|
||||||
|
def outline(el, d=0, max_lines=80, lines=None):
|
||||||
|
if lines is None: lines = []
|
||||||
|
if len(lines) >= max_lines: return lines
|
||||||
|
name = el.name
|
||||||
|
cls = ' '.join(el.get('class', []))
|
||||||
|
eid = el.get('id', '')
|
||||||
|
lbl = name
|
||||||
|
if eid: lbl += f'#{eid}'
|
||||||
|
if cls: lbl += '.' + cls.replace(' ', '.')
|
||||||
|
if name == 'a':
|
||||||
|
t = el.get_text(strip=True)[:40]
|
||||||
|
h = el.get('href', '')[:80]
|
||||||
|
lines.append(' ' * d + f'{lbl} "{t}" → {h}')
|
||||||
|
else:
|
||||||
|
lines.append(' ' * d + lbl)
|
||||||
|
for c in el.find_all(recursive=False):
|
||||||
|
if c.name in ('script', 'style'): continue
|
||||||
|
outline(c, d+1, max_lines, lines)
|
||||||
|
if len(lines) >= max_lines: return lines
|
||||||
|
return lines
|
||||||
|
|
||||||
|
|
||||||
|
# 단양군 - outline #menu_sitemap
|
||||||
|
print('='*70)
|
||||||
|
print('단양군 — #menu_sitemap outline')
|
||||||
|
print('='*70)
|
||||||
|
_, _, html = fetch('https://www.danyang.go.kr/dy21/98')
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
el = soup.select_one('#menu_sitemap') or soup.select_one('#contents_sitemap')
|
||||||
|
if el:
|
||||||
|
print(f' Container a={len(el.find_all("a"))}')
|
||||||
|
for line in outline(el, max_lines=60):
|
||||||
|
print(' ', line)
|
||||||
|
|
||||||
|
|
||||||
|
# 진천군 - outline nav#gnb on sub.do?menukey=445
|
||||||
|
print('\n' + '='*70)
|
||||||
|
print('진천군 — sub.do?menukey=445 nav#gnb outline')
|
||||||
|
print('='*70)
|
||||||
|
_, _, html = fetch('https://www.jincheon.go.kr/home/sub.do?menukey=445')
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
el = soup.select_one('nav#gnb') or soup.select_one('#gnb')
|
||||||
|
if el:
|
||||||
|
print(f' Container a={len(el.find_all("a"))}')
|
||||||
|
for line in outline(el, max_lines=60):
|
||||||
|
print(' ', line)
|
||||||
|
|
||||||
|
|
||||||
|
# 보은군 — try alternate hosts
|
||||||
|
print('\n' + '='*70)
|
||||||
|
print('보은군 — DNS/ alternate hosts')
|
||||||
|
print('='*70)
|
||||||
|
for host in ['www.boeun.go.kr', 'boeun.go.kr', 'boeun.chungbuk.go.kr']:
|
||||||
|
try:
|
||||||
|
ip = socket.gethostbyname(host)
|
||||||
|
print(f' {host} → {ip}')
|
||||||
|
except Exception as e:
|
||||||
|
print(f' {host}: {e}')
|
||||||
|
|
||||||
|
# Try fetching via curl-style direct
|
||||||
|
for url in ['https://www.boeun.go.kr/www/index.do',
|
||||||
|
'http://www.boeun.go.kr/www/index.do',
|
||||||
|
'https://boeun.go.kr/www/index.do',
|
||||||
|
'https://www.boeun.go.kr/']:
|
||||||
|
code, real, _ = fetch(url, timeout=10)
|
||||||
|
print(f' {url} → HTTP {code} ({real[:60]})')
|
||||||
61
_스크립트/_probe_chungbuk5.py
Normal file
61
_스크립트/_probe_chungbuk5.py
Normal file
@ -0,0 +1,61 @@
|
|||||||
|
"""진천군: nav#gnb 더 깊이 들어가서 실제 메뉴 찾기. + 보은군 한번 더."""
|
||||||
|
import re
|
||||||
|
import warnings
|
||||||
|
import requests
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
|
||||||
|
|
||||||
|
def fetch(url, timeout=20):
|
||||||
|
try:
|
||||||
|
r = requests.get(url, headers=H, timeout=timeout, verify=False, allow_redirects=True)
|
||||||
|
r.encoding = r.apparent_encoding
|
||||||
|
return r.status_code, r.url, r.text
|
||||||
|
except Exception as e:
|
||||||
|
return 0, str(e), ''
|
||||||
|
|
||||||
|
|
||||||
|
# 진천군 - find all elements with menu-like classes
|
||||||
|
print('='*70)
|
||||||
|
print('진천군 — find menu containers')
|
||||||
|
print('='*70)
|
||||||
|
_, _, html = fetch('https://www.jincheon.go.kr/home/sub.do?menukey=445')
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
# Look for elements with 'depth' / 'gnb' / 'menu' / 'sitemap' in classes
|
||||||
|
print('All elements with 100+ anchors and depth/menu/sitemap in class:')
|
||||||
|
for el in soup.find_all(['div', 'ul', 'nav', 'section']):
|
||||||
|
cls = ' '.join(el.get('class', []))
|
||||||
|
eid = el.get('id', '')
|
||||||
|
ac = len(el.find_all('a'))
|
||||||
|
if ac >= 100 and re.search(r'depth|menu|sitemap|allmenu|total|all-menu|gnb', cls + ' ' + eid, re.I):
|
||||||
|
print(f' {el.name}#{eid}.{cls[:60]} a={ac}')
|
||||||
|
|
||||||
|
# Find all .depth* classes
|
||||||
|
print('\n.depth* classes with anchors:')
|
||||||
|
for el in soup.select('[class*=depth]'):
|
||||||
|
cls = ' '.join(el.get('class', []))
|
||||||
|
ac = len(el.find_all('a'))
|
||||||
|
if ac >= 50:
|
||||||
|
print(f' {el.name}.{cls[:80]} a={ac}')
|
||||||
|
|
||||||
|
|
||||||
|
# 보은군 한번 더
|
||||||
|
print('\n' + '='*70)
|
||||||
|
print('보은군 retry')
|
||||||
|
print('='*70)
|
||||||
|
import time
|
||||||
|
time.sleep(1)
|
||||||
|
code, real, html = fetch('https://www.boeun.go.kr/www/index.do')
|
||||||
|
print(f' {code} {real[:100]}')
|
||||||
|
|
||||||
|
if code == 200:
|
||||||
|
print(' 성공! 사이트맵 후보 탐색')
|
||||||
|
from urllib.parse import urljoin
|
||||||
|
s2 = BeautifulSoup(html, 'html.parser')
|
||||||
|
for a in s2.find_all('a', href=True)[:200]:
|
||||||
|
txt = a.get_text(strip=True)
|
||||||
|
if '사이트맵' in txt or '전체메뉴' in txt or '누리집 지도' in txt:
|
||||||
|
print(f' "{txt}" → {urljoin(real, a["href"])}')
|
||||||
39
_스크립트/_probe_chungbuk6.py
Normal file
39
_스크립트/_probe_chungbuk6.py
Normal file
@ -0,0 +1,39 @@
|
|||||||
|
"""진천군 div.sitemap outline."""
|
||||||
|
import warnings
|
||||||
|
import requests
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
|
||||||
|
r = requests.get('https://www.jincheon.go.kr/home/sub.do?menukey=445', headers=H, timeout=15, verify=False)
|
||||||
|
r.encoding = r.apparent_encoding
|
||||||
|
soup = BeautifulSoup(r.text, 'html.parser')
|
||||||
|
|
||||||
|
def outline(el, d=0, max_lines=70, lines=None):
|
||||||
|
if lines is None: lines = []
|
||||||
|
if len(lines) >= max_lines: return lines
|
||||||
|
name = el.name
|
||||||
|
cls = ' '.join(el.get('class', []))
|
||||||
|
eid = el.get('id', '')
|
||||||
|
lbl = name
|
||||||
|
if eid: lbl += f'#{eid}'
|
||||||
|
if cls: lbl += '.' + cls.replace(' ', '.')
|
||||||
|
if name == 'a':
|
||||||
|
t = el.get_text(strip=True)[:40]
|
||||||
|
h = el.get('href', '')[:80]
|
||||||
|
lines.append(' ' * d + f'{lbl} "{t}" → {h}')
|
||||||
|
else:
|
||||||
|
lines.append(' ' * d + lbl)
|
||||||
|
for c in el.find_all(recursive=False):
|
||||||
|
if c.name in ('script', 'style'): continue
|
||||||
|
outline(c, d+1, max_lines, lines)
|
||||||
|
if len(lines) >= max_lines: return lines
|
||||||
|
return lines
|
||||||
|
|
||||||
|
el = soup.select_one('div.sitemap') or soup.select_one('ul.gnb-wrap')
|
||||||
|
if el:
|
||||||
|
print(f'Container: {el.name}.{" ".join(el.get("class",[]))} a={len(el.find_all("a"))}')
|
||||||
|
for line in outline(el, max_lines=60):
|
||||||
|
print(' ', line)
|
||||||
74
_스크립트/_probe_eGov.py
Normal file
74
_스크립트/_probe_eGov.py
Normal file
@ -0,0 +1,74 @@
|
|||||||
|
"""Inspect e-Gov sitemap_grep structure — count amThum sections and h2 text."""
|
||||||
|
import warnings
|
||||||
|
import requests
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
|
||||||
|
SITES = {
|
||||||
|
'당진시': 'https://www.dangjin.go.kr/kor/sitemap_11.do',
|
||||||
|
'보령시': 'https://www.brcn.go.kr/kor/sitemap_11.do',
|
||||||
|
'서천군': 'https://www.seocheon.go.kr/kor/sitemap_11.do',
|
||||||
|
'청양군': 'https://www.cheongyang.go.kr/kor/sitemap_11.do',
|
||||||
|
'태안군': 'https://www.taean.go.kr/kor/sitemap_11.do',
|
||||||
|
'서산시(top_menu)': 'https://www.seosan.go.kr/www/index.do',
|
||||||
|
'예산군': 'https://www.yesan.go.kr/kor/sitemap.do',
|
||||||
|
'천안시': 'https://www.cheonan.go.kr/kor/sitemap.do',
|
||||||
|
'홍성군': 'https://www.hongseong.go.kr/kor/sitemap.do',
|
||||||
|
}
|
||||||
|
|
||||||
|
for name, url in SITES.items():
|
||||||
|
print(f'\n=== {name} ===')
|
||||||
|
try:
|
||||||
|
r = requests.get(url, headers=H, timeout=15, verify=False, allow_redirects=True)
|
||||||
|
r.encoding = r.apparent_encoding
|
||||||
|
except Exception as e:
|
||||||
|
print(f' ERR: {e}')
|
||||||
|
continue
|
||||||
|
soup = BeautifulSoup(r.text, 'html.parser')
|
||||||
|
# Find amThum sections (eGov sitemap_grep)
|
||||||
|
amthums = soup.select('div.amThum')
|
||||||
|
if amthums:
|
||||||
|
print(f' amThum 섹션: {len(amthums)}')
|
||||||
|
for at in amthums[:10]:
|
||||||
|
h2 = at.find('h2')
|
||||||
|
h2_text = h2.get_text(strip=True) if h2 else ''
|
||||||
|
grep = at.find('div', class_='sitemap_grep')
|
||||||
|
n_list = len(grep.find_all('ul', class_='sitemap_list')) if grep else 0
|
||||||
|
n_first = len(at.find_all('a', class_='first'))
|
||||||
|
print(f' "{h2_text}" — sitemap_list={n_list}, a.first={n_first}')
|
||||||
|
continue
|
||||||
|
# Holsung 패턴 — div.sitemap.type2.nN > dl > dt + dd
|
||||||
|
sm = soup.select_one('div.sitemap[class*=type2]')
|
||||||
|
if sm:
|
||||||
|
dls = sm.find_all('dl', recursive=False)
|
||||||
|
print(f' sitemap type2 dl: {len(dls)}')
|
||||||
|
for dl in dls[:10]:
|
||||||
|
dt = dl.find('dt')
|
||||||
|
dt_text = dt.get_text(strip=True) if dt else ''
|
||||||
|
dds = dl.find_all('dd', recursive=False)
|
||||||
|
print(f' "{dt_text}" — dd={len(dds)}')
|
||||||
|
continue
|
||||||
|
# Yesan 패턴 — ul.depth1_ul
|
||||||
|
dep1 = soup.select('ul.depth1_ul > li, ul.depth1-ul > li')
|
||||||
|
if dep1:
|
||||||
|
print(f' depth1 li: {len(dep1)}')
|
||||||
|
for li in dep1[:10]:
|
||||||
|
a = li.find(['a', 'button'], recursive=False) or li.find(['a', 'button'])
|
||||||
|
if a:
|
||||||
|
txt = a.get_text(strip=True)
|
||||||
|
print(f' "{txt}"')
|
||||||
|
continue
|
||||||
|
# Seosan top_menu pattern
|
||||||
|
tm = soup.select_one('ul.top_menu')
|
||||||
|
if tm:
|
||||||
|
deps = tm.find_all('li', class_='depth1', recursive=False)
|
||||||
|
print(f' top_menu li.depth1: {len(deps)}')
|
||||||
|
for li in deps[:10]:
|
||||||
|
a = li.find('a', class_='depth1_ti')
|
||||||
|
if a:
|
||||||
|
print(f' "{a.get_text(strip=True)}"')
|
||||||
|
continue
|
||||||
|
print(' Unknown pattern')
|
||||||
98
_스크립트/_probe_fix.py
Normal file
98
_스크립트/_probe_fix.py
Normal file
@ -0,0 +1,98 @@
|
|||||||
|
"""Targeted probes for 청양군, 아산시, 부여군 to fix parsers."""
|
||||||
|
import warnings
|
||||||
|
import requests
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
|
||||||
|
|
||||||
|
def fetch(url):
|
||||||
|
r = requests.get(url, headers=H, timeout=15, verify=False, allow_redirects=True)
|
||||||
|
r.encoding = r.apparent_encoding
|
||||||
|
return r.text
|
||||||
|
|
||||||
|
|
||||||
|
# 청양군 — look at where the actual sitemap content lives
|
||||||
|
print('=== 청양군 ===')
|
||||||
|
html = fetch('https://www.cheongyang.go.kr/kor/sitemap_11.do')
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
# Look for #contents or main content
|
||||||
|
for sel in ['#contents', '.contents', 'main', '#txt', '#mainSection']:
|
||||||
|
el = soup.select_one(sel)
|
||||||
|
if el:
|
||||||
|
a_count = len(el.find_all('a'))
|
||||||
|
print(f' {sel}: a={a_count}')
|
||||||
|
# Look for the right sitemap container
|
||||||
|
print(' All elements with 200+ anchors:')
|
||||||
|
for el in soup.find_all(['div', 'ul', 'section']):
|
||||||
|
cls = ' '.join(el.get('class', []))
|
||||||
|
eid = el.get('id', '')
|
||||||
|
ac = len(el.find_all('a'))
|
||||||
|
if 200 <= ac <= 800 and (cls or eid):
|
||||||
|
print(f' {el.name}#{eid}.{cls[:60]} a={ac}')
|
||||||
|
|
||||||
|
|
||||||
|
print('\n=== 아산시 ===')
|
||||||
|
html = fetch('https://www.asan.go.kr/main/')
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
# Inspect mGnb-anchor1 structure deeply
|
||||||
|
sec = soup.find('div', id='mGnb-anchor1')
|
||||||
|
if sec:
|
||||||
|
# outline
|
||||||
|
def outline(el, d=0, max_lines=50, lines=None):
|
||||||
|
if lines is None: lines = []
|
||||||
|
if len(lines) >= max_lines: return lines
|
||||||
|
name = el.name
|
||||||
|
cls = ' '.join(el.get('class', []))
|
||||||
|
eid = el.get('id', '')
|
||||||
|
lbl = name
|
||||||
|
if eid: lbl += f'#{eid}'
|
||||||
|
if cls: lbl += '.' + cls.replace(' ', '.')
|
||||||
|
if name == 'a':
|
||||||
|
t = el.get_text(strip=True)[:30]
|
||||||
|
h = el.get('href', '')[:60]
|
||||||
|
lines.append(' ' * d + f'{lbl} "{t}" → {h}')
|
||||||
|
elif name == 'button':
|
||||||
|
t = el.get_text(strip=True)[:30]
|
||||||
|
lines.append(' ' * d + f'{lbl} BTN "{t}"')
|
||||||
|
else:
|
||||||
|
lines.append(' ' * d + lbl)
|
||||||
|
for c in el.find_all(recursive=False):
|
||||||
|
if c.name in ('script', 'style'): continue
|
||||||
|
outline(c, d+1, max_lines, lines)
|
||||||
|
if len(lines) >= max_lines: return lines
|
||||||
|
return lines
|
||||||
|
for line in outline(sec, max_lines=70):
|
||||||
|
print(' ', line)
|
||||||
|
|
||||||
|
|
||||||
|
print('\n=== 부여군 ===')
|
||||||
|
html = fetch('https://www.buyeo.go.kr/html/kr/')
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
# Inspect nav#gnb structure
|
||||||
|
sec = soup.select_one('nav#gnb')
|
||||||
|
if sec:
|
||||||
|
def outline(el, d=0, max_lines=80, lines=None):
|
||||||
|
if lines is None: lines = []
|
||||||
|
if len(lines) >= max_lines: return lines
|
||||||
|
name = el.name
|
||||||
|
cls = ' '.join(el.get('class', []))
|
||||||
|
eid = el.get('id', '')
|
||||||
|
lbl = name
|
||||||
|
if eid: lbl += f'#{eid}'
|
||||||
|
if cls: lbl += '.' + cls.replace(' ', '.')
|
||||||
|
if name == 'a':
|
||||||
|
t = el.get_text(strip=True)[:30]
|
||||||
|
h = el.get('href', '')[:60]
|
||||||
|
lines.append(' ' * d + f'{lbl} "{t}" → {h}')
|
||||||
|
else:
|
||||||
|
lines.append(' ' * d + lbl)
|
||||||
|
for c in el.find_all(recursive=False):
|
||||||
|
if c.name in ('script', 'style'): continue
|
||||||
|
outline(c, d+1, max_lines, lines)
|
||||||
|
if len(lines) >= max_lines: return lines
|
||||||
|
return lines
|
||||||
|
for line in outline(sec, max_lines=80):
|
||||||
|
print(' ', line)
|
||||||
60
_스크립트/_probe_fix2.py
Normal file
60
_스크립트/_probe_fix2.py
Normal file
@ -0,0 +1,60 @@
|
|||||||
|
"""제천시·증평군 구조 재확인."""
|
||||||
|
import warnings
|
||||||
|
import requests
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
|
||||||
|
|
||||||
|
def outline(el, d=0, max_lines=70, lines=None):
|
||||||
|
if lines is None: lines = []
|
||||||
|
if len(lines) >= max_lines: return lines
|
||||||
|
name = el.name
|
||||||
|
cls = ' '.join(el.get('class', []))
|
||||||
|
eid = el.get('id', '')
|
||||||
|
lbl = name
|
||||||
|
if eid: lbl += f'#{eid}'
|
||||||
|
if cls: lbl += '.' + cls.replace(' ', '.')
|
||||||
|
if name == 'a':
|
||||||
|
t = el.get_text(strip=True)[:40]
|
||||||
|
h = el.get('href', '')[:80]
|
||||||
|
lines.append(' ' * d + f'{lbl} "{t}" → {h}')
|
||||||
|
else:
|
||||||
|
lines.append(' ' * d + lbl)
|
||||||
|
for c in el.find_all(recursive=False):
|
||||||
|
if c.name in ('script', 'style'): continue
|
||||||
|
outline(c, d+1, max_lines, lines)
|
||||||
|
if len(lines) >= max_lines: return lines
|
||||||
|
return lines
|
||||||
|
|
||||||
|
|
||||||
|
# 제천시
|
||||||
|
print('='*70)
|
||||||
|
print('제천시 — div.depth1 outline')
|
||||||
|
print('='*70)
|
||||||
|
r = requests.get('https://www.jecheon.go.kr/www/sitemap.do?key=553', headers=H, timeout=15, verify=False)
|
||||||
|
r.encoding = r.apparent_encoding
|
||||||
|
soup = BeautifulSoup(r.text, 'html.parser')
|
||||||
|
for sel in ['div.depth1', 'div.depth.depth1', '#sitemap']:
|
||||||
|
el = soup.select_one(sel)
|
||||||
|
if el:
|
||||||
|
print(f'\n>>> {sel} (a={len(el.find_all("a"))}, classes={el.get("class")})')
|
||||||
|
for line in outline(el, max_lines=50):
|
||||||
|
print(' ', line)
|
||||||
|
break
|
||||||
|
|
||||||
|
|
||||||
|
# 증평군 - ul.depth1_ul outline
|
||||||
|
print('\n' + '='*70)
|
||||||
|
print('증평군 — ul.depth1_ul outline')
|
||||||
|
print('='*70)
|
||||||
|
r = requests.get('https://www.jp.go.kr/kor/sitemap_11.do', headers=H, timeout=15, verify=False)
|
||||||
|
r.encoding = r.apparent_encoding
|
||||||
|
soup = BeautifulSoup(r.text, 'html.parser')
|
||||||
|
el = soup.select_one('ul.depth1_ul')
|
||||||
|
if el:
|
||||||
|
print(f' (a={len(el.find_all("a"))}, classes={el.get("class")})')
|
||||||
|
for line in outline(el, max_lines=60):
|
||||||
|
print(' ', line)
|
||||||
181
_스크립트/_probe_jeonbuk.py
Normal file
181
_스크립트/_probe_jeonbuk.py
Normal file
@ -0,0 +1,181 @@
|
|||||||
|
"""전북·제주 16개 시·군 사이트맵 URL 탐색 + 컨테이너 구조 덤프."""
|
||||||
|
import re
|
||||||
|
import ssl
|
||||||
|
import sys
|
||||||
|
import warnings
|
||||||
|
from urllib.parse import urljoin, urlparse
|
||||||
|
|
||||||
|
import requests
|
||||||
|
from requests.adapters import HTTPAdapter
|
||||||
|
from urllib3.util.ssl_ import create_urllib3_context
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
|
||||||
|
|
||||||
|
class WeakSSLAdapter(HTTPAdapter):
|
||||||
|
def init_poolmanager(self, *args, **kwargs):
|
||||||
|
ctx = create_urllib3_context()
|
||||||
|
ctx.set_ciphers('DEFAULT@SECLEVEL=0')
|
||||||
|
ctx.options |= 0x4
|
||||||
|
ctx.check_hostname = False
|
||||||
|
ctx.verify_mode = ssl.CERT_NONE
|
||||||
|
kwargs['ssl_context'] = ctx
|
||||||
|
return super().init_poolmanager(*args, **kwargs)
|
||||||
|
|
||||||
|
|
||||||
|
def make_session(weak=False):
|
||||||
|
s = requests.Session()
|
||||||
|
s.headers.update(H)
|
||||||
|
if weak:
|
||||||
|
s.mount('https://', WeakSSLAdapter())
|
||||||
|
return s
|
||||||
|
|
||||||
|
|
||||||
|
def fetch(session, url, timeout=20):
|
||||||
|
r = session.get(url, timeout=timeout, verify=False, allow_redirects=True)
|
||||||
|
meta = re.search(rb'<meta[^>]*charset=["\']?\s*([\w-]+)', r.content[:4096], re.I)
|
||||||
|
if meta:
|
||||||
|
r.encoding = meta.group(1).decode('ascii', errors='ignore')
|
||||||
|
else:
|
||||||
|
r.encoding = r.apparent_encoding
|
||||||
|
return r
|
||||||
|
|
||||||
|
|
||||||
|
SITES = [
|
||||||
|
('고창군', 'https://www.gochang.go.kr/index.gochang?contentsSid=3136'),
|
||||||
|
('군산시', 'https://www.gunsan.go.kr/main'),
|
||||||
|
('김제시', 'https://www.gimje.go.kr/index.gimje'),
|
||||||
|
('남원시', 'https://www.namwon.go.kr/index.do?menuUid=ff8080818e3beff0018e40e8f63e02d2'),
|
||||||
|
('무주군', 'https://www.muju.go.kr/index.9is'),
|
||||||
|
('부안군', 'https://www.buan.go.kr/index.buan?contentsSid=1'),
|
||||||
|
('순창군', 'https://www.sunchang.go.kr/'),
|
||||||
|
('완주군', 'https://www.wanju.go.kr/index.9is'),
|
||||||
|
('익산시', 'https://www.iksan.go.kr/index.do?menuUid=ff8080819a39930e019a4de8c1ae0afd'),
|
||||||
|
('임실군', 'https://www.imsil.go.kr/index.imsil'),
|
||||||
|
('장수군', 'https://www.jangsu.go.kr/index.jangsu'),
|
||||||
|
('전주시', 'https://www.jeonju.go.kr/index.9is'),
|
||||||
|
('정읍시', 'https://www.jeongeup.go.kr/index.jeongeup'),
|
||||||
|
('진안군', 'https://www.jinan.go.kr/index.jinan?contentsSid=1379'),
|
||||||
|
('서귀포시', 'https://www.seogwipo.go.kr/index.htm'),
|
||||||
|
('제주시', 'https://www.jejusi.go.kr/index.ac'),
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def find_sitemap_links(soup, base):
|
||||||
|
"""페이지에서 사이트맵으로 보이는 링크 후보 수집."""
|
||||||
|
cands = []
|
||||||
|
for a in soup.find_all('a', href=True):
|
||||||
|
txt = (a.get_text() or '').strip()
|
||||||
|
href = a['href']
|
||||||
|
onclick = a.get('onclick', '')
|
||||||
|
blob = f'{txt} {href} {onclick}'.lower()
|
||||||
|
if '사이트맵' in txt or 'sitemap' in blob:
|
||||||
|
url = href
|
||||||
|
if url.startswith('#') or url.lower().startswith('javascript:'):
|
||||||
|
# onclick에서 추출 시도
|
||||||
|
m = re.search(r"""['"]([^'"]*(?:sitemap|site_map)[^'"]*)['"]""", onclick, re.I)
|
||||||
|
if m:
|
||||||
|
url = m.group(1)
|
||||||
|
else:
|
||||||
|
continue
|
||||||
|
cands.append((txt, urljoin(base, url)))
|
||||||
|
# 중복 제거
|
||||||
|
seen = set(); out = []
|
||||||
|
for t, u in cands:
|
||||||
|
if u not in seen:
|
||||||
|
seen.add(u); out.append((t, u))
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def dump_structure(soup):
|
||||||
|
"""사이트맵 페이지 주요 컨테이너 후보 출력."""
|
||||||
|
# 흔한 사이트맵 컨테이너 셀렉터 후보
|
||||||
|
sels = [
|
||||||
|
'div.sitemap', 'div#sitemap', 'div.site_map', 'div#site_map',
|
||||||
|
'ul#menu_sitemap', 'div.depth.depth1', 'div.depth1',
|
||||||
|
'ul.depth1_ul', 'div.amThum', 'ul.sitemap', 'div.sitemap_wrap',
|
||||||
|
'div.contents_sitemap', 'div.sitemapWrap', 'div.site-map',
|
||||||
|
'div.allmenu', 'div#allmenu', 'div.gnb_all', 'div.total_menu',
|
||||||
|
]
|
||||||
|
found = []
|
||||||
|
for sel in sels:
|
||||||
|
els = soup.select(sel)
|
||||||
|
if els:
|
||||||
|
found.append((sel, len(els)))
|
||||||
|
print(f' 매칭 셀렉터: {found}')
|
||||||
|
# 사이트맵스러운 컨테이너 한 개 잡아서 자식 구조 덤프
|
||||||
|
target = None
|
||||||
|
for sel, _ in found:
|
||||||
|
target = soup.select_one(sel)
|
||||||
|
if target:
|
||||||
|
print(f' >>> {sel} 내부 구조:')
|
||||||
|
break
|
||||||
|
if not target:
|
||||||
|
# body에서 class에 sitemap/menu/depth 포함 div 찾기
|
||||||
|
for div in soup.find_all(['div', 'ul'], class_=True):
|
||||||
|
cls = ' '.join(div.get('class', []))
|
||||||
|
if re.search(r'sitemap|site_map|allmenu|depth1|total_menu', cls, re.I):
|
||||||
|
target = div
|
||||||
|
print(f' >>> <{div.name} class="{cls}"> 내부 구조:')
|
||||||
|
break
|
||||||
|
if not target:
|
||||||
|
print(' !! 사이트맵 컨테이너 미발견')
|
||||||
|
return
|
||||||
|
# 자식 1~3레벨 태그/클래스 요약
|
||||||
|
def summarize(el, depth=0, maxdepth=4):
|
||||||
|
if depth > maxdepth:
|
||||||
|
return
|
||||||
|
for child in el.find_all(recursive=False):
|
||||||
|
cls = '.'.join(child.get('class', []))
|
||||||
|
idv = child.get('id', '')
|
||||||
|
tag = child.name
|
||||||
|
label = tag + (f'#{idv}' if idv else '') + (f'.{cls}' if cls else '')
|
||||||
|
a = child.find('a', recursive=False)
|
||||||
|
atxt = (a.get_text().strip()[:20] if a else '')
|
||||||
|
print(' ' * (depth + 1) + f'{label}' + (f' a="{atxt}"' if atxt else ''))
|
||||||
|
if depth < 2:
|
||||||
|
summarize(child, depth + 1, maxdepth)
|
||||||
|
summarize(target)
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
targets = sys.argv[1:]
|
||||||
|
for name, url in SITES:
|
||||||
|
if targets and name not in targets:
|
||||||
|
continue
|
||||||
|
print(f'\n{"="*70}\n[{name}] {url}')
|
||||||
|
for weak in (False, True):
|
||||||
|
try:
|
||||||
|
sess = make_session(weak=weak)
|
||||||
|
r = fetch(sess, url)
|
||||||
|
base = f'{urlparse(r.url).scheme}://{urlparse(r.url).netloc}'
|
||||||
|
soup = BeautifulSoup(r.text, 'html.parser')
|
||||||
|
print(f' status={r.status_code} final={r.url} weak_ssl={weak}')
|
||||||
|
links = find_sitemap_links(soup, base)
|
||||||
|
print(f' 사이트맵 링크 후보: {links[:6]}')
|
||||||
|
# 가장 그럴듯한 후보 따라가기
|
||||||
|
if links:
|
||||||
|
smurl = links[0][1]
|
||||||
|
try:
|
||||||
|
r2 = fetch(sess, smurl)
|
||||||
|
soup2 = BeautifulSoup(r2.text, 'html.parser')
|
||||||
|
print(f' 사이트맵 페이지: {r2.url} (status {r2.status_code})')
|
||||||
|
dump_structure(soup2)
|
||||||
|
except Exception as e:
|
||||||
|
print(f' 사이트맵 페이지 fetch 실패: {e}')
|
||||||
|
else:
|
||||||
|
print(' >>> 인덱스 자체 구조 확인:')
|
||||||
|
dump_structure(soup)
|
||||||
|
break
|
||||||
|
except Exception as e:
|
||||||
|
if weak:
|
||||||
|
print(f' !! 실패(weak 포함): {type(e).__name__}: {e}')
|
||||||
|
else:
|
||||||
|
print(f' (일반 SSL 실패 → weak 재시도): {type(e).__name__}')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
84
_스크립트/_probe_jeonbuk2.py
Normal file
84
_스크립트/_probe_jeonbuk2.py
Normal file
@ -0,0 +1,84 @@
|
|||||||
|
"""전북 미해결 사이트 2차 정밀 탐색."""
|
||||||
|
import re, ssl, sys, warnings
|
||||||
|
from urllib.parse import urljoin, urlparse
|
||||||
|
import requests
|
||||||
|
from requests.adapters import HTTPAdapter
|
||||||
|
from urllib3.util.ssl_ import create_urllib3_context
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
|
||||||
|
def sess():
|
||||||
|
s = requests.Session(); s.headers.update({'User-Agent': UA}); return s
|
||||||
|
|
||||||
|
def fetch(s, url, t=20):
|
||||||
|
r = s.get(url, timeout=t, verify=False, allow_redirects=True)
|
||||||
|
meta = re.search(rb'<meta[^>]*charset=["\']?\s*([\w-]+)', r.content[:4096], re.I)
|
||||||
|
r.encoding = meta.group(1).decode('ascii','ignore') if meta else r.apparent_encoding
|
||||||
|
return r
|
||||||
|
|
||||||
|
# 사이트맵 링크를 못 찾은 사이트: 인덱스 HTML에서 sitemap/menuCd/전체메뉴 힌트 검색
|
||||||
|
NOLINK = {
|
||||||
|
'군산시': 'https://www.gunsan.go.kr/main',
|
||||||
|
'남원시': 'https://www.namwon.go.kr/index.do?menuUid=ff8080818e3beff0018e40e8f63e02d2',
|
||||||
|
'부안군': 'https://www.buan.go.kr/index.buan?contentsSid=1',
|
||||||
|
'완주군': 'https://www.wanju.go.kr/index.9is',
|
||||||
|
'임실군': 'https://www.imsil.go.kr/index.imsil',
|
||||||
|
'정읍시': 'https://www.jeongeup.go.kr/index.jeongeup',
|
||||||
|
}
|
||||||
|
|
||||||
|
def probe_links(name, url):
|
||||||
|
print(f'\n{"="*70}\n[{name}] {url}')
|
||||||
|
s = sess()
|
||||||
|
try:
|
||||||
|
r = fetch(s, url)
|
||||||
|
except Exception as e:
|
||||||
|
print(' fetch fail', e); return
|
||||||
|
base = f'{urlparse(r.url).scheme}://{urlparse(r.url).netloc}'
|
||||||
|
soup = BeautifulSoup(r.text, 'html.parser')
|
||||||
|
# sitemap/전체메뉴 텍스트나 href를 가진 a 전부
|
||||||
|
hits = []
|
||||||
|
for a in soup.find_all('a'):
|
||||||
|
txt = (a.get_text() or '').strip()
|
||||||
|
href = a.get('href','') or ''
|
||||||
|
oc = a.get('onclick','') or ''
|
||||||
|
blob = f'{txt}|{href}|{oc}'
|
||||||
|
if re.search(r'사이트맵|전체메뉴|sitemap|site_map|allmenu', blob, re.I):
|
||||||
|
hits.append((txt[:20], href[:90], oc[:90]))
|
||||||
|
for h in hits[:15]:
|
||||||
|
print(' a:', h)
|
||||||
|
# menuCd 패턴 가진 href 중 sitemap 후보 (DOM_...02000000 류)
|
||||||
|
cds = set(re.findall(r'menuCd=DOM_\d+', r.text))
|
||||||
|
print(' menuCd 샘플:', list(cds)[:10])
|
||||||
|
|
||||||
|
for n, u in NOLINK.items():
|
||||||
|
probe_links(n, u)
|
||||||
|
|
||||||
|
# ---- 구조 깊이 확인이 필요한 사이트들 ----
|
||||||
|
DEEP = {
|
||||||
|
'고창군': 'https://www.gochang.go.kr/index.gochang?menuCd=DOM_000000106006000000',
|
||||||
|
'익산시': 'https://www.iksan.go.kr/index.do?menuUid=ff80808199f0d11c019a041b8e35174a',
|
||||||
|
'전주시': 'https://www.jeonju.go.kr/index.9is?contentUid=ff8080818c7c2e8e018c7fd67b9a03b6',
|
||||||
|
'순창군': 'https://www.sunchang.go.kr/',
|
||||||
|
'진안군': 'https://www.jinan.go.kr/index.jinan?menuCd=DOM_000000110002000000',
|
||||||
|
}
|
||||||
|
|
||||||
|
def deep(name, url, sel):
|
||||||
|
print(f'\n{"#"*70}\n[{name}] DEEP {url} sel={sel}')
|
||||||
|
s = sess()
|
||||||
|
try:
|
||||||
|
r = fetch(s, url)
|
||||||
|
except Exception as e:
|
||||||
|
print(' fail', e); return
|
||||||
|
soup = BeautifulSoup(r.text, 'html.parser')
|
||||||
|
cont = soup.select_one(sel)
|
||||||
|
if not cont:
|
||||||
|
print(' 컨테이너 없음'); return
|
||||||
|
# 첫 블록 하나만 골라 li/a href까지 자세히
|
||||||
|
print(cont.prettify()[:2500])
|
||||||
|
|
||||||
|
deep('고창군 첫메뉴', 'https://www.gochang.go.kr/index.gochang?menuCd=DOM_000000106006000000', 'div.sitemap div.menu1')
|
||||||
|
deep('익산시 group수', 'https://www.iksan.go.kr/index.do?menuUid=ff80808199f0d11c019a041b8e35174a', 'div.sitemap_group')
|
||||||
|
deep('전주시', 'https://www.jeonju.go.kr/index.9is?contentUid=ff8080818c7c2e8e018c7fd67b9a03b6', 'div.sitemap_Warp')
|
||||||
|
deep('순창군 gnb', 'https://www.sunchang.go.kr/', 'div.sitemap_box ul.gnb_list')
|
||||||
|
deep('진안군 첫dl', 'https://www.jinan.go.kr/index.jinan?menuCd=DOM_000000110002000000', 'div.sitemap dl')
|
||||||
99
_스크립트/_probe_jeonbuk3.py
Normal file
99
_스크립트/_probe_jeonbuk3.py
Normal file
@ -0,0 +1,99 @@
|
|||||||
|
"""전북 미해결 사이트 3차: 군산/남원/부안/완주/임실/순창 사이트맵 URL 확정."""
|
||||||
|
import re, sys, warnings
|
||||||
|
from urllib.parse import urljoin, urlparse
|
||||||
|
import requests
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
|
||||||
|
def sess():
|
||||||
|
s = requests.Session(); s.headers.update({'User-Agent': UA}); return s
|
||||||
|
|
||||||
|
def fetch(s, url, t=20):
|
||||||
|
r = s.get(url, timeout=t, verify=False, allow_redirects=True)
|
||||||
|
meta = re.search(rb'<meta[^>]*charset=["\']?\s*([\w-]+)', r.content[:4096], re.I)
|
||||||
|
r.encoding = meta.group(1).decode('ascii','ignore') if meta else r.apparent_encoding
|
||||||
|
return r
|
||||||
|
|
||||||
|
def show_container(soup, label):
|
||||||
|
for sel in ['div.sitemap','div#sitemap','div.sitemap_group','div.sitemap_Warp',
|
||||||
|
'div.sitemap_wrap','ul.siteMapList','div.allmenu','div#allmenu',
|
||||||
|
'div.contents','div.allmenubox','div.total_menu','div.site_map']:
|
||||||
|
els = soup.select(sel)
|
||||||
|
if els:
|
||||||
|
print(f' [{label}] sel={sel} x{len(els)}')
|
||||||
|
|
||||||
|
S = sess()
|
||||||
|
|
||||||
|
# 정읍 확정: 전체메뉴보기 menuCd
|
||||||
|
r = fetch(S, 'https://www.jeongeup.go.kr/index.jeongeup?menuCd=DOM_000000106002000000')
|
||||||
|
soup = BeautifulSoup(r.text, 'html.parser')
|
||||||
|
print('=== 정읍 사이트맵 ===', r.url)
|
||||||
|
show_container(soup, '정읍')
|
||||||
|
cont = soup.select_one('div.sitemap')
|
||||||
|
if cont:
|
||||||
|
blocks = cont.find_all('div', recursive=False)
|
||||||
|
print(' div.sitemap 직계 div:', len(blocks), [ '.'.join(b.get('class',[])) for b in blocks[:8]])
|
||||||
|
if blocks:
|
||||||
|
print(blocks[0].prettify()[:800])
|
||||||
|
|
||||||
|
# 부안/임실: index 페이지에서 "사이트맵/누리집지도/전체메뉴" 텍스트 가진 a 또는 그 주변 menuCd
|
||||||
|
for name, url in [('부안군','https://www.buan.go.kr/index.buan?contentsSid=1'),
|
||||||
|
('임실군','https://www.imsil.go.kr/index.imsil')]:
|
||||||
|
r = fetch(S, url)
|
||||||
|
soup = BeautifulSoup(r.text, 'html.parser')
|
||||||
|
print(f'\n=== {name} index — 사이트맵류 a ===')
|
||||||
|
for a in soup.find_all('a'):
|
||||||
|
t = (a.get_text() or '').strip()
|
||||||
|
if re.search(r'사이트맵|누리집지도|전체메뉴', t):
|
||||||
|
print(' a:', repr(t[:25]), '| href=', a.get('href'), '| onclick=', (a.get('onclick') or '')[:80])
|
||||||
|
# 부모/형제에 menuCd 있나
|
||||||
|
par = a.find_parent()
|
||||||
|
print(' parent menuCd:', re.findall(r'menuCd=DOM_\d+', str(par))[:3])
|
||||||
|
|
||||||
|
# 남원/완주: menuUid/contentUid 기반 — 사이트맵 링크 a 검색
|
||||||
|
for name, url, key in [('남원시','https://www.namwon.go.kr/index.do?menuUid=ff8080818e3beff0018e40e8f63e02d2','menuUid'),
|
||||||
|
('완주군','https://www.wanju.go.kr/index.9is','contentUid')]:
|
||||||
|
r = fetch(S, url)
|
||||||
|
soup = BeautifulSoup(r.text, 'html.parser')
|
||||||
|
print(f'\n=== {name} index — 사이트맵/전체메뉴 a ({key}) ===')
|
||||||
|
found = False
|
||||||
|
for a in soup.find_all('a'):
|
||||||
|
t = (a.get_text() or '').strip()
|
||||||
|
href = a.get('href') or ''
|
||||||
|
if re.search(r'사이트맵|누리집지도|전체메뉴|site_map|sitemap', t + href, re.I):
|
||||||
|
print(' a:', repr(t[:25]), '| href=', href[:100])
|
||||||
|
found = True
|
||||||
|
if not found:
|
||||||
|
print(' (없음) — 모든 a href에서 sitemap 토큰 검색:')
|
||||||
|
for a in soup.find_all('a', href=True):
|
||||||
|
if re.search(r'sitemap|site_map', a['href'], re.I):
|
||||||
|
print(' ', a['href'][:110], '|', (a.get_text() or '').strip()[:20])
|
||||||
|
|
||||||
|
# 군산: 사이트맵 페이지 추정 — /main 외 흔한 경로 시도
|
||||||
|
print('\n=== 군산 사이트맵 후보 ===')
|
||||||
|
for guess in ['https://www.gunsan.go.kr/sitemap','https://www.gunsan.go.kr/kor/sitemap.do',
|
||||||
|
'https://www.gunsan.go.kr/main?menuCd=','https://www.gunsan.go.kr/sitemap.do']:
|
||||||
|
try:
|
||||||
|
r = fetch(S, guess, t=10)
|
||||||
|
print(f' {guess} -> {r.status_code} {r.url}')
|
||||||
|
except Exception as e:
|
||||||
|
print(f' {guess} -> ERR {type(e).__name__}')
|
||||||
|
r = fetch(S, 'https://www.gunsan.go.kr/main')
|
||||||
|
soup = BeautifulSoup(r.text, 'html.parser')
|
||||||
|
print(' 군산 index 사이트맵류 a:')
|
||||||
|
for a in soup.find_all('a'):
|
||||||
|
t = (a.get_text() or '').strip()
|
||||||
|
href = a.get('href') or ''
|
||||||
|
if re.search(r'사이트맵|누리집지도|전체메뉴|site_map|sitemap', t + href, re.I):
|
||||||
|
print(' ', repr(t[:25]), '| href=', href[:100])
|
||||||
|
|
||||||
|
# 순창: 정적 사이트맵 페이지 탐색
|
||||||
|
print('\n=== 순창 사이트맵 후보 ===')
|
||||||
|
r = fetch(S, 'https://www.sunchang.go.kr/')
|
||||||
|
soup = BeautifulSoup(r.text, 'html.parser')
|
||||||
|
for a in soup.find_all('a'):
|
||||||
|
t = (a.get_text() or '').strip()
|
||||||
|
href = a.get('href') or ''
|
||||||
|
if re.search(r'사이트맵|누리집지도|전체메뉴|site_map|sitemap', t + href, re.I):
|
||||||
|
print(' a:', repr(t[:25]), '| href=', href[:110])
|
||||||
85
_스크립트/_probe_jeonbuk4.py
Normal file
85
_스크립트/_probe_jeonbuk4.py
Normal file
@ -0,0 +1,85 @@
|
|||||||
|
"""4차: 부안/완주/군산/순창 메뉴 소스 확정."""
|
||||||
|
import re, warnings
|
||||||
|
from urllib.parse import urljoin
|
||||||
|
import requests
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
def sess():
|
||||||
|
s = requests.Session(); s.headers.update({'User-Agent': UA}); return s
|
||||||
|
def fetch(s, url, t=20):
|
||||||
|
r = s.get(url, timeout=t, verify=False, allow_redirects=True)
|
||||||
|
meta = re.search(rb'<meta[^>]*charset=["\']?\s*([\w-]+)', r.content[:4096], re.I)
|
||||||
|
r.encoding = meta.group(1).decode('ascii','ignore') if meta else r.apparent_encoding
|
||||||
|
return r
|
||||||
|
S = sess()
|
||||||
|
|
||||||
|
def dump_menu_divs(soup, label):
|
||||||
|
print(f'\n--- {label}: class에 menu/gnb/sitemap/allmenu 포함 div·nav·ul (직계 a 텍스트) ---')
|
||||||
|
for el in soup.find_all(['div','nav','ul']):
|
||||||
|
cls = ' '.join(el.get('class', []))
|
||||||
|
idv = el.get('id','')
|
||||||
|
if re.search(r'menu|gnb|sitemap|site_map|allmenu|lnb|depth', cls + ' ' + idv, re.I):
|
||||||
|
n_a = len(el.find_all('a'))
|
||||||
|
n_li = len(el.find_all('li'))
|
||||||
|
if n_a >= 15: # 메가메뉴 후보
|
||||||
|
print(f' <{el.name} id="{idv}" class="{cls}"> a={n_a} li={n_li}')
|
||||||
|
|
||||||
|
# 부안: 사이트맵 menuCd 추정 — index의 모든 menuCd 중 *07*(정보공개류) 상위 노드 시도
|
||||||
|
print('=== 부안군 ===')
|
||||||
|
r = fetch(S, 'https://www.buan.go.kr/index.buan?contentsSid=1')
|
||||||
|
soup = BeautifulSoup(r.text, 'html.parser')
|
||||||
|
dump_menu_divs(soup, '부안 index')
|
||||||
|
# 누리집지도/사이트맵 후보 menuCd 시도
|
||||||
|
for cd in ['DOM_000000107003000000','DOM_000000108003000000','DOM_000000107002000000',
|
||||||
|
'DOM_000000106002000000','DOM_000000109002000000']:
|
||||||
|
try:
|
||||||
|
rr = fetch(S, f'https://www.buan.go.kr/index.buan?menuCd={cd}', t=10)
|
||||||
|
ss = BeautifulSoup(rr.text, 'html.parser')
|
||||||
|
has = bool(ss.select_one('div.sitemap'))
|
||||||
|
ttl = (ss.title.get_text().strip()[:30] if ss.title else '')
|
||||||
|
print(f' {cd}: div.sitemap={has} title={ttl!r}')
|
||||||
|
except Exception as e:
|
||||||
|
print(f' {cd}: ERR {type(e).__name__}')
|
||||||
|
|
||||||
|
# 완주: index 메가메뉴 + contentUid 사이트맵 후보
|
||||||
|
print('\n=== 완주군 ===')
|
||||||
|
r = fetch(S, 'https://www.wanju.go.kr/index.9is')
|
||||||
|
soup = BeautifulSoup(r.text, 'html.parser')
|
||||||
|
dump_menu_divs(soup, '완주 index')
|
||||||
|
for a in soup.find_all('a'):
|
||||||
|
t = (a.get_text() or '').strip()
|
||||||
|
if re.search(r'누리집|사이트맵|전체메뉴|지도', t):
|
||||||
|
print(' 완주 a:', repr(t[:25]), '| href=', (a.get('href') or '')[:100], '| onclick=', (a.get('onclick') or '')[:90])
|
||||||
|
|
||||||
|
# 군산: allmenubox 전체 구조(탭 포함)
|
||||||
|
print('\n=== 군산시 ===')
|
||||||
|
r = fetch(S, 'https://www.gunsan.go.kr/main')
|
||||||
|
soup = BeautifulSoup(r.text, 'html.parser')
|
||||||
|
dump_menu_divs(soup, '군산 index')
|
||||||
|
box = soup.select_one('div.allmenubox') or soup.select_one('div.allmw')
|
||||||
|
if box:
|
||||||
|
# 탭/카테고리 구조 — 직계 자식 요약
|
||||||
|
print(' allmenubox 하위 ul/div (a수):')
|
||||||
|
for el in box.find_all(['ul','div'], recursive=True):
|
||||||
|
cls = ' '.join(el.get('class', []))
|
||||||
|
n_a = len(el.find_all('a', recursive=False))
|
||||||
|
if n_a >= 5:
|
||||||
|
print(f' <{el.name} class="{cls}"> 직계a={n_a} 첫a={el.find("a").get_text().strip()[:15]!r}')
|
||||||
|
|
||||||
|
# 순창: HTML 내 sitemap/menu json/ajax 흔적
|
||||||
|
print('\n=== 순창군 ===')
|
||||||
|
r = fetch(S, 'https://www.sunchang.go.kr/')
|
||||||
|
html = r.text
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
dump_menu_divs(soup, '순창 index')
|
||||||
|
print(' html 내 sitemap/menu ajax url 흔적:')
|
||||||
|
for m in set(re.findall(r'["\']([^"\']*(?:sitemap|menu)[^"\']*\.(?:do|json|html|jsp))["\']', html, re.I)):
|
||||||
|
print(' ', m[:100])
|
||||||
|
# 순창 gnb_list가 ajax라면, 흔한 e-Gov 전체메뉴 경로 시도
|
||||||
|
for guess in ['/site/main/menu/sitemap','/kr/sitemap','/sitemap.do','/main/sitemap']:
|
||||||
|
try:
|
||||||
|
rr = fetch(S, urljoin('https://www.sunchang.go.kr', guess), t=10)
|
||||||
|
print(f' {guess} -> {rr.status_code}')
|
||||||
|
except Exception as e:
|
||||||
|
print(f' {guess} -> ERR')
|
||||||
58
_스크립트/_probe_jeonbuk5.py
Normal file
58
_스크립트/_probe_jeonbuk5.py
Normal file
@ -0,0 +1,58 @@
|
|||||||
|
"""5차: 부안/완주/군산/순창 인라인 메가메뉴 중첩 구조 정밀 덤프."""
|
||||||
|
import re, warnings
|
||||||
|
import requests
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
def fetch(url, t=20):
|
||||||
|
r = requests.get(url, timeout=t, verify=False, headers={'User-Agent': UA})
|
||||||
|
meta = re.search(rb'<meta[^>]*charset=["\']?\s*([\w-]+)', r.content[:4096], re.I)
|
||||||
|
r.encoding = meta.group(1).decode('ascii','ignore') if meta else r.apparent_encoding
|
||||||
|
return r
|
||||||
|
|
||||||
|
def tree(el, depth=0, maxd=4, maxchild=6):
|
||||||
|
if el is None or depth > maxd: return
|
||||||
|
kids = el.find_all(recursive=False)
|
||||||
|
for i, c in enumerate(kids):
|
||||||
|
if i >= maxchild and depth >= 1:
|
||||||
|
print(' '*(depth+1) + '...'); break
|
||||||
|
cls = '.'.join(c.get('class', []))
|
||||||
|
a = c.find('a', recursive=False)
|
||||||
|
atxt = a.get_text().strip()[:18] if a else ''
|
||||||
|
ah = (a.get('href') or '')[:55] if a else ''
|
||||||
|
print(' '*(depth+1) + f'{c.name}.{cls}' + (f' a={atxt!r} {ah}' if atxt else ''))
|
||||||
|
tree(c, depth+1, maxd, maxchild)
|
||||||
|
|
||||||
|
print('############ 부안 nav#onmenu ############')
|
||||||
|
soup = BeautifulSoup(fetch('https://www.buan.go.kr/index.buan?contentsSid=1').text, 'html.parser')
|
||||||
|
nav = soup.select_one('nav#onmenu')
|
||||||
|
# 첫 1~2개 top li만
|
||||||
|
if nav:
|
||||||
|
top = nav.find('ul')
|
||||||
|
print('nav>ul 첫 li 2개:')
|
||||||
|
for li in (top.find_all('li', recursive=False)[:2] if top else []):
|
||||||
|
tree(li, 0, 4, 5)
|
||||||
|
print(' ----')
|
||||||
|
|
||||||
|
print('\n############ 완주 div.top_menu_wrap ############')
|
||||||
|
soup = BeautifulSoup(fetch('https://www.wanju.go.kr/index.9is').text, 'html.parser')
|
||||||
|
w = soup.select_one('div.top_menu_wrap')
|
||||||
|
tree(w, 0, 3, 4)
|
||||||
|
|
||||||
|
print('\n############ 군산 첫 allmenubox ############')
|
||||||
|
soup = BeautifulSoup(fetch('https://www.gunsan.go.kr/main').text, 'html.parser')
|
||||||
|
pc = soup.select_one('div#all_pcmenu')
|
||||||
|
if pc:
|
||||||
|
boxes = pc.select('div.allmenubox')
|
||||||
|
print(f'allmenubox 수: {len(boxes)}')
|
||||||
|
b = boxes[0]
|
||||||
|
print('Bmenu(대분류):', repr((b.find('a', class_='Bmenu') or b.find('a')).get_text().strip()[:20]))
|
||||||
|
tree(b, 0, 4, 5)
|
||||||
|
|
||||||
|
print('\n############ 순창 ul.gnb ############')
|
||||||
|
soup = BeautifulSoup(fetch('https://www.sunchang.go.kr/').text, 'html.parser')
|
||||||
|
g = soup.select_one('ul.gnb')
|
||||||
|
if g:
|
||||||
|
print(f'ul.gnb 직계 li: {len(g.find_all("li", recursive=False))}')
|
||||||
|
for li in g.find_all('li', recursive=False)[:1]:
|
||||||
|
tree(li, 0, 4, 5)
|
||||||
46
_스크립트/_probe_jeonbuk6.py
Normal file
46
_스크립트/_probe_jeonbuk6.py
Normal file
@ -0,0 +1,46 @@
|
|||||||
|
"""6차: 완주 메뉴/사이트맵 확정 + 부안 대분류 라벨 확인."""
|
||||||
|
import re, warnings
|
||||||
|
import requests
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
def fetch(url, t=20):
|
||||||
|
r = requests.get(url, timeout=t, verify=False, headers={'User-Agent': UA})
|
||||||
|
meta = re.search(rb'<meta[^>]*charset=["\']?\s*([\w-]+)', r.content[:4096], re.I)
|
||||||
|
r.encoding = meta.group(1).decode('ascii','ignore') if meta else r.apparent_encoding
|
||||||
|
return r
|
||||||
|
|
||||||
|
# 부안 대분류 라벨: 각 depth_boxcon의 strong/p 텍스트
|
||||||
|
print('=== 부안 대분류(D) 라벨 ===')
|
||||||
|
soup = BeautifulSoup(fetch('https://www.buan.go.kr/index.buan?contentsSid=1').text, 'html.parser')
|
||||||
|
nav = soup.select_one('nav#onmenu')
|
||||||
|
top = nav.find('ul')
|
||||||
|
for li in top.find_all('li', recursive=False):
|
||||||
|
a0 = li.find('a', recursive=False)
|
||||||
|
box = li.find('div', class_='depth_boxcon')
|
||||||
|
strong = box.find('strong') if box else None
|
||||||
|
p = box.find('p') if box else None
|
||||||
|
print(' topA=', repr((a0.get_text().strip()[:20]) if a0 else ''),
|
||||||
|
'| strong=', repr(strong.get_text().strip()[:20] if strong else ''),
|
||||||
|
'| p=', repr(p.get_text(' ',strip=True)[:25] if p else ''))
|
||||||
|
|
||||||
|
# 완주: 전체 a 중 contentUid 사이트맵/메가메뉴 흔적. gnb mega menu 클래스 탐색
|
||||||
|
print('\n=== 완주 메뉴 구조 탐색 ===')
|
||||||
|
soup = BeautifulSoup(fetch('https://www.wanju.go.kr/index.9is').text, 'html.parser')
|
||||||
|
# 모든 div/ul 중 a>=30 이며 menuUid/contentUid href 다수인 것
|
||||||
|
for el in soup.find_all(['div','ul','nav']):
|
||||||
|
cls = ' '.join(el.get('class', [])); idv = el.get('id','')
|
||||||
|
n_a = len(el.find_all('a'))
|
||||||
|
if n_a >= 30:
|
||||||
|
sample = el.find('a', href=re.compile(r'contentUid|menuUid'))
|
||||||
|
print(f' <{el.name} id={idv!r} class={cls!r}> a={n_a}',
|
||||||
|
('| 샘플=' + (sample.get('href')[:60] if sample else 'none')))
|
||||||
|
# 사이트맵 페이지 후보: 전주처럼 index.9is?contentUid= 의 sitemap_Warp
|
||||||
|
# 완주 모든 contentUid 수집 후 'sitemap_Warp' 또는 'sitemap' 포함 페이지 찾기엔 비용 큼.
|
||||||
|
# 대신 footer 영역 a 전부 출력
|
||||||
|
foot = soup.find('footer') or soup.select_one('div.footer, #footer')
|
||||||
|
if foot:
|
||||||
|
print(' footer a:')
|
||||||
|
for a in foot.find_all('a')[:40]:
|
||||||
|
t=(a.get_text() or '').strip()
|
||||||
|
if t: print(' ', repr(t[:18]), (a.get('href') or '')[:60])
|
||||||
172
_스크립트/_probe_sitemap.py
Normal file
172
_스크립트/_probe_sitemap.py
Normal file
@ -0,0 +1,172 @@
|
|||||||
|
"""Probe sitemap pages directly using candidate URLs found in probe pass 1."""
|
||||||
|
import re
|
||||||
|
import warnings
|
||||||
|
from urllib.parse import urljoin, urlparse
|
||||||
|
from concurrent.futures import ThreadPoolExecutor, as_completed
|
||||||
|
|
||||||
|
import requests
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
|
||||||
|
# Round 1 results — best candidate sitemap URL per site
|
||||||
|
CANDIDATES = {
|
||||||
|
'논산시': 'https://nonsan.go.kr/kor/html/sub07/0701.html',
|
||||||
|
'당진시': 'https://www.dangjin.go.kr/kor/sitemap_11.do',
|
||||||
|
'보령시': 'https://www.brcn.go.kr/kor/sitemap_11.do',
|
||||||
|
'부여군': 'https://www.buyeo.go.kr/html/kr/sitemap.do',
|
||||||
|
'서산시': 'https://www.seosan.go.kr/www/sitemap.do',
|
||||||
|
'서천군': 'https://www.seocheon.go.kr/kor/sitemap_11.do',
|
||||||
|
'아산시': 'https://www.asan.go.kr/main/sitemap.do',
|
||||||
|
'예산군': 'https://www.yesan.go.kr/kor/sitemap.do',
|
||||||
|
'천안시': 'https://www.cheonan.go.kr/kor/sitemap.do',
|
||||||
|
'청양군': 'https://www.cheongyang.go.kr/kor/sitemap_11.do',
|
||||||
|
'태안군': 'https://www.taean.go.kr/kor/sitemap_11.do',
|
||||||
|
'홍성군': 'https://www.hongseong.go.kr/kor/sitemap.do',
|
||||||
|
}
|
||||||
|
|
||||||
|
# Alternative URLs to try if main candidate fails
|
||||||
|
ALTS = {
|
||||||
|
'부여군': ['https://www.buyeo.go.kr/html/kr/html/sub07/0701.html',
|
||||||
|
'https://www.buyeo.go.kr/html/kr/sitemap.html',
|
||||||
|
'https://www.buyeo.go.kr/html/kr/sitemap.do'],
|
||||||
|
'서산시': ['https://www.seosan.go.kr/www/sitemap.do',
|
||||||
|
'https://www.seosan.go.kr/www/contents.do?key=151'],
|
||||||
|
'아산시': ['https://www.asan.go.kr/main/sitemap.do',
|
||||||
|
'https://www.asan.go.kr/main/sub01_01.do'],
|
||||||
|
'예산군': ['https://www.yesan.go.kr/kor/sitemap.do',
|
||||||
|
'https://www.yesan.go.kr/kor/sitemap01.do'],
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def fetch(url, timeout=15):
|
||||||
|
try:
|
||||||
|
r = requests.get(url, headers=H, timeout=timeout, verify=False, allow_redirects=True)
|
||||||
|
r.encoding = r.apparent_encoding
|
||||||
|
if r.status_code == 200:
|
||||||
|
return r.url, r.text
|
||||||
|
except Exception as e:
|
||||||
|
return None, f'ERR: {e}'
|
||||||
|
return None, f'HTTP {r.status_code}'
|
||||||
|
|
||||||
|
|
||||||
|
def deep_analyze(html):
|
||||||
|
"""Deeply analyze HTML to find sitemap container regardless of class name."""
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
# Try several common selectors
|
||||||
|
selectors = [
|
||||||
|
'.sitemap', '#sitemap', '.site_map', '.sitemap_wrap',
|
||||||
|
'.allMenu', '.allmenu', '#allMenu', '.all_menu', '.allMenuWrap',
|
||||||
|
'.contents .menu', '#content .sitemap', '.menu_all',
|
||||||
|
'.gnb_all', '.totalMenu', '.totMenu',
|
||||||
|
'div[class*=sitemap]', 'div[class*=allMenu]',
|
||||||
|
]
|
||||||
|
candidates = []
|
||||||
|
for sel in selectors:
|
||||||
|
for el in soup.select(sel):
|
||||||
|
# Count nested anchors as a measure of usefulness
|
||||||
|
a_count = len(el.find_all('a'))
|
||||||
|
if a_count >= 20:
|
||||||
|
candidates.append((a_count, sel, el))
|
||||||
|
if not candidates:
|
||||||
|
# Fallback — find any container with most anchors (excluding header/footer)
|
||||||
|
all_divs = soup.find_all(['div', 'section', 'main', 'article'])
|
||||||
|
for d in all_divs:
|
||||||
|
cls = ' '.join(d.get('class', []))
|
||||||
|
if 'header' in cls.lower() or 'footer' in cls.lower() or 'gnb' in cls.lower() and 'all' not in cls.lower():
|
||||||
|
continue
|
||||||
|
a_count = len(d.find_all('a'))
|
||||||
|
if a_count >= 50:
|
||||||
|
candidates.append((a_count, f'div.{cls}', d))
|
||||||
|
candidates.sort(reverse=True)
|
||||||
|
return candidates[:3]
|
||||||
|
|
||||||
|
|
||||||
|
def describe(el):
|
||||||
|
"""Describe DOM structure of an element."""
|
||||||
|
info = {
|
||||||
|
'tag': el.name,
|
||||||
|
'class': ' '.join(el.get('class', [])),
|
||||||
|
'id': el.get('id', ''),
|
||||||
|
'a_count': len(el.find_all('a')),
|
||||||
|
'dl_count': len(el.find_all('dl')),
|
||||||
|
'ul_count': len(el.find_all('ul')),
|
||||||
|
'li_count': len(el.find_all('li')),
|
||||||
|
}
|
||||||
|
# Identify pattern
|
||||||
|
dls_direct = el.find_all('dl', recursive=False)
|
||||||
|
uls_direct = el.find_all('ul', recursive=False)
|
||||||
|
info['dl_direct'] = len(dls_direct)
|
||||||
|
info['ul_direct'] = len(uls_direct)
|
||||||
|
if dls_direct and dls_direct[0].find('dt') and dls_direct[0].find('dd'):
|
||||||
|
info['pattern'] = 'A (dl>dt|dd>ul)'
|
||||||
|
elif uls_direct:
|
||||||
|
info['pattern'] = 'B (ul nested)'
|
||||||
|
else:
|
||||||
|
# Search one level deeper
|
||||||
|
nested_dl = []
|
||||||
|
nested_ul = []
|
||||||
|
for child in el.find_all(['div', 'section'], recursive=False):
|
||||||
|
nested_dl.extend(child.find_all('dl', recursive=False))
|
||||||
|
nested_ul.extend(child.find_all('ul', recursive=False))
|
||||||
|
if nested_dl:
|
||||||
|
info['pattern'] = f'A nested 1 deep (dl={len(nested_dl)})'
|
||||||
|
elif nested_ul:
|
||||||
|
info['pattern'] = f'B nested 1 deep (ul={len(nested_ul)})'
|
||||||
|
else:
|
||||||
|
info['pattern'] = 'UNKNOWN'
|
||||||
|
return info
|
||||||
|
|
||||||
|
|
||||||
|
def probe(name, url):
|
||||||
|
print(f'\n=== {name} === {url}')
|
||||||
|
u, html = fetch(url)
|
||||||
|
if not u:
|
||||||
|
print(f' 실패: {html}')
|
||||||
|
# Try alternatives
|
||||||
|
for alt in ALTS.get(name, []):
|
||||||
|
u, html = fetch(alt)
|
||||||
|
if u:
|
||||||
|
print(f' 대체 URL: {alt}')
|
||||||
|
break
|
||||||
|
else:
|
||||||
|
return name, None
|
||||||
|
print(f' 최종 URL: {u} ({len(html)} bytes)')
|
||||||
|
cands = deep_analyze(html)
|
||||||
|
if not cands:
|
||||||
|
print(' 사이트맵 컨테이너 못 찾음')
|
||||||
|
# Dump some snippets to help diagnose
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
for tag in ['title', 'h1', 'h2']:
|
||||||
|
for t in soup.find_all(tag)[:3]:
|
||||||
|
print(f' {tag}: {t.get_text(strip=True)[:80]}')
|
||||||
|
return name, None
|
||||||
|
for a_count, sel, el in cands:
|
||||||
|
info = describe(el)
|
||||||
|
print(f' [{a_count} anchors] selector="{sel}" → {info}')
|
||||||
|
# Return best candidate info
|
||||||
|
best_count, best_sel, best_el = cands[0]
|
||||||
|
return name, {'url': u, 'selector': best_sel, **describe(best_el)}
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
results = {}
|
||||||
|
with ThreadPoolExecutor(max_workers=4) as ex:
|
||||||
|
futs = {ex.submit(probe, n, u): n for n, u in CANDIDATES.items()}
|
||||||
|
for f in as_completed(futs):
|
||||||
|
n, info = f.result()
|
||||||
|
results[n] = info
|
||||||
|
print('\n=== 최종 요약 ===')
|
||||||
|
for n in CANDIDATES:
|
||||||
|
info = results.get(n)
|
||||||
|
if info:
|
||||||
|
print(f' {n}: {info["url"]} | sel={info["selector"]} | a={info["a_count"]} dl={info["dl_count"]} ul={info["ul_count"]} | {info["pattern"]}')
|
||||||
|
else:
|
||||||
|
print(f' {n}: NOT FOUND')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
221
_스크립트/_recheck_kogl_all.py
Normal file
221
_스크립트/_recheck_kogl_all.py
Normal file
@ -0,0 +1,221 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""계룡시 방식 일괄 적용 (수정판) — 충청남도(2~15) + 충청북도(1~11).
|
||||||
|
|
||||||
|
★ 원본 phase234 검출 로직을 그대로 재사용한다(import):
|
||||||
|
- get_body(body_sel)로 '본문 영역'에 한정해 검출 → 헤더/푸터 외부링크 노이즈 제거
|
||||||
|
- extract_detail_urls(JS fn_detail 폴백 포함)로 게시판 상세를 원본과 동일하게 추적
|
||||||
|
- KOGL_IMG_PAT = img_opentype(\\d{2}).png (2자리 png), KOGL_LINK_PAT = licenseType(\\d)
|
||||||
|
|
||||||
|
각 부착 행을 재크롤링하여
|
||||||
|
1) O열(15) = 이미지(img_opentype) 유형만으로 재판정 (링크 숫자는 합치지 않음)
|
||||||
|
2) 링크가 '존재'하면서 이미지≠링크면 S열(19) 비고에 '링크주소 오기' (링크 없음은 비움)
|
||||||
|
|
||||||
|
파일별 *_backup_kogl재판정전.xlsx 백업 후 저장(기존 백업 있으면 보존).
|
||||||
|
|
||||||
|
사용: python -X utf8 _recheck_kogl_all.py [기관명 ...]
|
||||||
|
"""
|
||||||
|
import sys, os, re, shutil, warnings, openpyxl
|
||||||
|
from concurrent.futures import ThreadPoolExecutor
|
||||||
|
from urllib.parse import urlparse
|
||||||
|
|
||||||
|
sys.path.insert(0, r'D:\01.프로젝트\DB수집\_스크립트')
|
||||||
|
sys.stdout.reconfigure(encoding='utf-8')
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
import _chungnam_phase234_all as CN
|
||||||
|
import _chungbuk_phase234_all as CB
|
||||||
|
|
||||||
|
O_COL, K_COL, S_COL = 15, 11, 19
|
||||||
|
WORKERS = 6
|
||||||
|
|
||||||
|
# CN/CB 모듈 SITES에 없는 기관(개별 스크립트만 존재) 보완 — body_sel은 계룡시와 동일
|
||||||
|
EXTRA_CFG = {
|
||||||
|
'공주시': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\2.공주시\충청남도_공주시.xlsx',
|
||||||
|
'body_sel': ['#txt', '#contents', 'main']},
|
||||||
|
'금산군': {'xlsx': r'D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\3.금산군\충청남도_금산군.xlsx',
|
||||||
|
'body_sel': ['#txt', '#contents', 'main']},
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def get_cfg(name, region):
|
||||||
|
mod = CN if region == 'cn' else CB
|
||||||
|
return mod.SITES.get(name) or EXTRA_CFG[name]
|
||||||
|
|
||||||
|
# (기관명, region, weak_ssl) — xlsx/body_sel 은 각 모듈 SITES 에서 가져옴
|
||||||
|
TARGETS = [
|
||||||
|
('공주시', 'cn', False), ('금산군', 'cn', False), ('논산시', 'cn', False),
|
||||||
|
('당진시', 'cn', False), ('보령시', 'cn', False), ('부여군', 'cn', False),
|
||||||
|
('서산시', 'cn', False), ('서천군', 'cn', False), ('아산시', 'cn', False),
|
||||||
|
('예산군', 'cn', False), ('천안시', 'cn', False), ('청양군', 'cn', False),
|
||||||
|
('태안군', 'cn', False), ('홍성군', 'cn', False),
|
||||||
|
('괴산군', 'cb', False), ('단양군', 'cb', False), ('영동군', 'cb', True),
|
||||||
|
('옥천군', 'cb', False), ('음성군', 'cb', False), ('제천시', 'cb', False),
|
||||||
|
('증평군', 'cb', False), ('진천군', 'cb', False), ('청주시', 'cb', False),
|
||||||
|
('충주시', 'cb', False),
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
# 확장 이미지 패턴: img_opentype / img_opencode, 1~2자리, png/jpg/jpeg/gif
|
||||||
|
BROAD_IMG_PAT = re.compile(r'(?:new_)?img_open(?:type|code)(\d{1,2})\.(?:png|jpe?g|gif)', re.I)
|
||||||
|
|
||||||
|
|
||||||
|
def _valid(n):
|
||||||
|
return 1 <= n <= 4 # 공공누리 유형은 1~4만 유효
|
||||||
|
|
||||||
|
|
||||||
|
def detect_split(body, LINK_PAT):
|
||||||
|
"""이미지명(확장패턴) 유형과 링크(licenseType) 유형을 분리 검출. 유형 1~4만."""
|
||||||
|
img_t, link_t = set(), set()
|
||||||
|
for a in body.find_all('a', href=True):
|
||||||
|
m = LINK_PAT.search(a['href'])
|
||||||
|
if m and _valid(int(m.group(1))):
|
||||||
|
link_t.add(int(m.group(1)))
|
||||||
|
# 이미지: src 속성 + style 배경이미지 + raw HTML(누락 방지)
|
||||||
|
blob = ' '.join(filter(None, (img.get('src', '') for img in body.find_all('img'))))
|
||||||
|
blob += ' ' + ' '.join(el.get('style', '') for el in body.find_all(style=True))
|
||||||
|
blob += ' ' + str(body)
|
||||||
|
for m in BROAD_IMG_PAT.finditer(blob):
|
||||||
|
n = int(m.group(1))
|
||||||
|
if _valid(n):
|
||||||
|
img_t.add(n)
|
||||||
|
return img_t, link_t
|
||||||
|
|
||||||
|
|
||||||
|
def decide_O(img_t, link_t):
|
||||||
|
"""O열 값과 '링크주소 오기' 비고 여부 결정.
|
||||||
|
- 이미지 있으면 이미지 우선(권위). 링크 존재 & 이미지≠링크면 비고.
|
||||||
|
- 이미지 없고 링크만: 1·2·3·4 전부면 범례(설명)페이지 → 미부착. 아니면 링크 유형 인정.
|
||||||
|
- 둘 다 없으면 미부착."""
|
||||||
|
if img_t:
|
||||||
|
newO = ','.join(f'{n}유형' for n in sorted(img_t))
|
||||||
|
return newO, bool(link_t) and (link_t != img_t)
|
||||||
|
if link_t:
|
||||||
|
if {1, 2, 3, 4}.issubset(link_t):
|
||||||
|
return '미부착', False
|
||||||
|
return ','.join(f'{n}유형' for n in sorted(link_t)), False
|
||||||
|
return '미부착', False
|
||||||
|
|
||||||
|
|
||||||
|
def make_fetch(region, weak_ssl):
|
||||||
|
if region == 'cn':
|
||||||
|
return CN.fetch # fetch(url, timeout)
|
||||||
|
sess = CB.make_session(weak_ssl)
|
||||||
|
return lambda url, timeout=12: CB.fetch(sess, url, timeout)
|
||||||
|
|
||||||
|
|
||||||
|
def _domain(host):
|
||||||
|
"""go.kr/or.kr 등 2단계 TLD 고려해 등록 도메인(끝 3라벨) 반환."""
|
||||||
|
labels = (host or '').split('.')
|
||||||
|
return '.'.join(labels[-3:]) if len(labels) >= 3 else host
|
||||||
|
|
||||||
|
|
||||||
|
def same_site(base_url, target_url):
|
||||||
|
"""상세 링크가 같은 기관 도메인일 때만 True (외부 사이트 KOGL 오탐 방지)."""
|
||||||
|
return _domain(urlparse(base_url).hostname) == _domain(urlparse(target_url).hostname)
|
||||||
|
|
||||||
|
|
||||||
|
def detect_row(url, body_sel, mod, fetch_fn):
|
||||||
|
soup = fetch_fn(url)
|
||||||
|
if soup is None:
|
||||||
|
return None, None, 'fetch_fail'
|
||||||
|
body = mod.get_body(soup, body_sel)
|
||||||
|
form, _ = mod.detect_form(body)
|
||||||
|
img_t, link_t = detect_split(body, mod.KOGL_LINK_PAT)
|
||||||
|
if form == '게시판':
|
||||||
|
for du in mod.extract_detail_urls(body, url, limit=5):
|
||||||
|
if not same_site(url, du): # 외부 도메인 상세링크 제외
|
||||||
|
continue
|
||||||
|
ds = fetch_fn(du, 10)
|
||||||
|
if ds is None:
|
||||||
|
continue
|
||||||
|
db = mod.get_body(ds, body_sel)
|
||||||
|
di, dl = detect_split(db, mod.KOGL_LINK_PAT)
|
||||||
|
img_t |= di
|
||||||
|
link_t |= dl
|
||||||
|
return img_t, link_t, 'ok'
|
||||||
|
|
||||||
|
|
||||||
|
def process_site(name, region, weak_ssl):
|
||||||
|
mod = CN if region == 'cn' else CB
|
||||||
|
cfg = get_cfg(name, region)
|
||||||
|
xlsx, body_sel = cfg['xlsx'], cfg['body_sel']
|
||||||
|
print(f'\n{"="*70}\n■ {name} [{region}] ({os.path.relpath(xlsx)})')
|
||||||
|
wb = openpyxl.load_workbook(xlsx)
|
||||||
|
ws = wb.active
|
||||||
|
targets = []
|
||||||
|
for r in range(3, ws.max_row + 1):
|
||||||
|
o = ws.cell(r, O_COL).value
|
||||||
|
if not o or str(o).strip() == '미부착':
|
||||||
|
continue
|
||||||
|
url = ws.cell(r, K_COL).value
|
||||||
|
if url:
|
||||||
|
targets.append((r, str(url).strip(), o, ws.cell(r, S_COL).value))
|
||||||
|
if not targets:
|
||||||
|
print(' 부착 행 없음 → 변경 없음')
|
||||||
|
return dict(name=name, total=0, o_changed=0, to_none=0, mismatch=0, fetch_fail=0)
|
||||||
|
|
||||||
|
print(f' 부착 행 {len(targets)}건 재크롤링 (본문영역+상세추적, workers={WORKERS}, weak_ssl={weak_ssl})')
|
||||||
|
fetch_fn = make_fetch(region, weak_ssl)
|
||||||
|
|
||||||
|
def work(t):
|
||||||
|
r, url, oldO, oldS = t
|
||||||
|
img_t, link_t, status = detect_row(url, body_sel, mod, fetch_fn)
|
||||||
|
return (r, url, oldO, oldS, img_t, link_t, status)
|
||||||
|
|
||||||
|
results = []
|
||||||
|
with ThreadPoolExecutor(max_workers=WORKERS) as ex:
|
||||||
|
for res in ex.map(work, targets):
|
||||||
|
results.append(res)
|
||||||
|
|
||||||
|
o_changed = to_none = mismatch = fetch_fail = 0
|
||||||
|
for r, url, oldO, oldS, img_t, link_t, status in sorted(results):
|
||||||
|
if status != 'ok':
|
||||||
|
fetch_fail += 1
|
||||||
|
print(f' row{r:>3} | FETCH_FAIL | {url[:70]}')
|
||||||
|
continue
|
||||||
|
newO, do_flag = decide_O(img_t, link_t)
|
||||||
|
if newO != str(oldO).strip():
|
||||||
|
o_changed += 1
|
||||||
|
if newO == '미부착':
|
||||||
|
to_none += 1
|
||||||
|
ws.cell(r, O_COL).value = newO
|
||||||
|
print(f' row{r:>3} | O: {str(oldO):>16} -> {newO:<16} <== 변경 | img={sorted(img_t)} link={sorted(link_t)}')
|
||||||
|
if do_flag:
|
||||||
|
mismatch += 1
|
||||||
|
cur = (oldS or '').strip()
|
||||||
|
if '링크주소 오기' not in cur:
|
||||||
|
ws.cell(r, S_COL).value = '링크주소 오기' if not cur else cur + ' / 링크주소 오기'
|
||||||
|
print(f' row{r:>3} | 비고+= 링크주소 오기 | img={sorted(img_t)} != link={sorted(link_t)}')
|
||||||
|
|
||||||
|
backup = os.path.join(os.path.dirname(xlsx), f'{name}_backup_kogl재판정전.xlsx')
|
||||||
|
if not os.path.exists(backup):
|
||||||
|
shutil.copy2(xlsx, backup)
|
||||||
|
wb.save(xlsx)
|
||||||
|
print(f' ▶ {name} 완료: O변경 {o_changed} (미부착化 {to_none}) | 불일치비고 {mismatch} | fetch실패 {fetch_fail} | 저장')
|
||||||
|
return dict(name=name, total=len(targets), o_changed=o_changed, to_none=to_none,
|
||||||
|
mismatch=mismatch, fetch_fail=fetch_fail)
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
sel = [a for a in sys.argv[1:] if not a.startswith('-')]
|
||||||
|
sites = [t for t in TARGETS if (not sel or t[0] in sel)]
|
||||||
|
print(f'대상 {len(sites)}개: {[s[0] for s in sites]}')
|
||||||
|
summary = []
|
||||||
|
for name, region, weak in sites:
|
||||||
|
try:
|
||||||
|
summary.append(process_site(name, region, weak))
|
||||||
|
except Exception as e:
|
||||||
|
import traceback; traceback.print_exc()
|
||||||
|
print(f' !! {name} 오류: {e}')
|
||||||
|
summary.append(dict(name=name, total=-1, o_changed=0, to_none=0, mismatch=0, fetch_fail=0))
|
||||||
|
|
||||||
|
print('\n' + '=' * 70 + '\n[전체 요약]')
|
||||||
|
print(f'{"기관":<8}{"부착":>6}{"O변경":>7}{"미부착化":>8}{"불일치":>7}{"실패":>6}')
|
||||||
|
for s in summary:
|
||||||
|
print(f'{s["name"]:<8}{s["total"]:>6}{s["o_changed"]:>7}{s["to_none"]:>8}{s["mismatch"]:>7}{s["fetch_fail"]:>6}')
|
||||||
|
tot = lambda k: sum(x[k] for x in summary if x[k] >= 0)
|
||||||
|
print(f'\n합계: 부착 {tot("total")} | O변경 {tot("o_changed")} | 미부착化 {tot("to_none")} | 불일치비고 {tot("mismatch")} | fetch실패 {tot("fetch_fail")}')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
178
_스크립트/_recheck_kogl_jeonbuk.py
Normal file
178
_스크립트/_recheck_kogl_jeonbuk.py
Normal file
@ -0,0 +1,178 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""전북(14)+제주(2) 공공누리(O열) 권위 재판정 — 충청도 _recheck_kogl_all.py 방식.
|
||||||
|
|
||||||
|
_jeonbuk_phase234_all 의 검출 로직(get_body/detect_form/extract_detail_urls/KOGL_LINK_PAT)을
|
||||||
|
그대로 재사용. 각 부착 행을 재크롤링하여:
|
||||||
|
1) O열 = 이미지명(img_opentype/opencode) 유형 우선 권위 재판정
|
||||||
|
2) 이미지 존재 & 이미지≠링크 → S열 비고 '링크주소 오기'
|
||||||
|
3) 링크만 있고 1·2·3·4 전부 → 범례페이지로 미부착
|
||||||
|
규칙: feedback_kogl_image_rule. 상세추적은 같은 기관 도메인만.
|
||||||
|
|
||||||
|
사용: python -X utf8 _recheck_kogl_jeonbuk.py [기관명 ...]
|
||||||
|
"""
|
||||||
|
import sys, os, re, shutil, warnings, openpyxl
|
||||||
|
from concurrent.futures import ThreadPoolExecutor
|
||||||
|
from urllib.parse import urlparse
|
||||||
|
|
||||||
|
sys.path.insert(0, r'D:\01.프로젝트\DB수집\_스크립트')
|
||||||
|
sys.stdout.reconfigure(encoding='utf-8')
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
import _jeonbuk_phase234_all as JB
|
||||||
|
|
||||||
|
O_COL, K_COL, S_COL = 15, 11, 19
|
||||||
|
WORKERS = 12
|
||||||
|
DETAIL_TIMEOUT = 5
|
||||||
|
DETAIL_LIMIT = 3
|
||||||
|
|
||||||
|
# (기관명, weak_ssl) — xlsx/body_sel 은 JB.SITES 에서 가져옴
|
||||||
|
TARGETS = [(name, JB.SITES[name].get('weak_ssl', False)) for name in JB.SITES]
|
||||||
|
|
||||||
|
BROAD_IMG_PAT = re.compile(r'(?:new_)?img_open(?:type|code)(\d{1,2})\.(?:png|jpe?g|gif)', re.I)
|
||||||
|
|
||||||
|
|
||||||
|
def _valid(n):
|
||||||
|
return 1 <= n <= 4
|
||||||
|
|
||||||
|
|
||||||
|
def detect_split(body, LINK_PAT):
|
||||||
|
img_t, link_t = set(), set()
|
||||||
|
for a in body.find_all('a', href=True):
|
||||||
|
m = LINK_PAT.search(a['href'])
|
||||||
|
if m and _valid(int(m.group(1))):
|
||||||
|
link_t.add(int(m.group(1)))
|
||||||
|
blob = ' '.join(filter(None, (img.get('src', '') for img in body.find_all('img'))))
|
||||||
|
blob += ' ' + ' '.join(el.get('style', '') for el in body.find_all(style=True))
|
||||||
|
blob += ' ' + str(body)
|
||||||
|
for m in BROAD_IMG_PAT.finditer(blob):
|
||||||
|
n = int(m.group(1))
|
||||||
|
if _valid(n):
|
||||||
|
img_t.add(n)
|
||||||
|
return img_t, link_t
|
||||||
|
|
||||||
|
|
||||||
|
def decide_O(img_t, link_t):
|
||||||
|
if img_t:
|
||||||
|
newO = ','.join(f'{n}유형' for n in sorted(img_t))
|
||||||
|
return newO, bool(link_t) and (link_t != img_t)
|
||||||
|
if link_t:
|
||||||
|
if {1, 2, 3, 4}.issubset(link_t):
|
||||||
|
return '미부착', False
|
||||||
|
return ','.join(f'{n}유형' for n in sorted(link_t)), False
|
||||||
|
return '미부착', False
|
||||||
|
|
||||||
|
|
||||||
|
def make_fetch(weak_ssl):
|
||||||
|
sess = JB.make_session(weak_ssl)
|
||||||
|
return lambda url, timeout=DETAIL_TIMEOUT: JB.fetch(sess, url, timeout)
|
||||||
|
|
||||||
|
|
||||||
|
def _domain(host):
|
||||||
|
labels = (host or '').split('.')
|
||||||
|
return '.'.join(labels[-3:]) if len(labels) >= 3 else host
|
||||||
|
|
||||||
|
|
||||||
|
def same_site(base_url, target_url):
|
||||||
|
return _domain(urlparse(base_url).hostname) == _domain(urlparse(target_url).hostname)
|
||||||
|
|
||||||
|
|
||||||
|
def detect_row(url, body_sel, fetch_fn):
|
||||||
|
soup = fetch_fn(url)
|
||||||
|
if soup is None:
|
||||||
|
return None, None, 'fetch_fail'
|
||||||
|
body = JB.get_body(soup, body_sel)
|
||||||
|
form, _ = JB.detect_form(body)
|
||||||
|
img_t, link_t = detect_split(body, JB.KOGL_LINK_PAT)
|
||||||
|
if form == '게시판':
|
||||||
|
for du in JB.extract_detail_urls(body, url, limit=DETAIL_LIMIT):
|
||||||
|
if not same_site(url, du):
|
||||||
|
continue
|
||||||
|
ds = fetch_fn(du, DETAIL_TIMEOUT)
|
||||||
|
if ds is None:
|
||||||
|
continue
|
||||||
|
db = JB.get_body(ds, body_sel)
|
||||||
|
di, dl = detect_split(db, JB.KOGL_LINK_PAT)
|
||||||
|
img_t |= di
|
||||||
|
link_t |= dl
|
||||||
|
return img_t, link_t, 'ok'
|
||||||
|
|
||||||
|
|
||||||
|
def process_site(name, weak_ssl):
|
||||||
|
cfg = JB.SITES[name]
|
||||||
|
xlsx, body_sel = cfg['xlsx'], cfg['body_sel']
|
||||||
|
print(f'\n{"="*70}\n■ {name} ({os.path.relpath(xlsx)})')
|
||||||
|
wb = openpyxl.load_workbook(xlsx)
|
||||||
|
ws = wb.active
|
||||||
|
targets = []
|
||||||
|
for r in range(3, ws.max_row + 1):
|
||||||
|
o = ws.cell(r, O_COL).value
|
||||||
|
if not o or str(o).strip() == '미부착':
|
||||||
|
continue
|
||||||
|
url = ws.cell(r, K_COL).value
|
||||||
|
if url:
|
||||||
|
targets.append((r, str(url).strip(), o, ws.cell(r, S_COL).value))
|
||||||
|
if not targets:
|
||||||
|
print(' 부착 행 없음 → 변경 없음')
|
||||||
|
return dict(name=name, total=0, o_changed=0, to_none=0, mismatch=0, fetch_fail=0)
|
||||||
|
|
||||||
|
print(f' 부착 행 {len(targets)}건 재크롤링 (workers={WORKERS}, weak_ssl={weak_ssl})')
|
||||||
|
fetch_fn = make_fetch(weak_ssl)
|
||||||
|
|
||||||
|
def work(t):
|
||||||
|
r, url, oldO, oldS = t
|
||||||
|
img_t, link_t, status = detect_row(url, body_sel, fetch_fn)
|
||||||
|
return (r, url, oldO, oldS, img_t, link_t, status)
|
||||||
|
|
||||||
|
results = []
|
||||||
|
with ThreadPoolExecutor(max_workers=WORKERS) as ex:
|
||||||
|
for res in ex.map(work, targets):
|
||||||
|
results.append(res)
|
||||||
|
|
||||||
|
o_changed = to_none = mismatch = fetch_fail = 0
|
||||||
|
for r, url, oldO, oldS, img_t, link_t, status in sorted(results):
|
||||||
|
if status != 'ok':
|
||||||
|
fetch_fail += 1
|
||||||
|
continue
|
||||||
|
newO, do_flag = decide_O(img_t, link_t)
|
||||||
|
if newO != str(oldO).strip():
|
||||||
|
o_changed += 1
|
||||||
|
if newO == '미부착':
|
||||||
|
to_none += 1
|
||||||
|
ws.cell(r, O_COL).value = newO
|
||||||
|
if do_flag:
|
||||||
|
mismatch += 1
|
||||||
|
cur = (oldS or '').strip()
|
||||||
|
if '링크주소 오기' not in cur:
|
||||||
|
ws.cell(r, S_COL).value = '링크주소 오기' if not cur else cur + ' / 링크주소 오기'
|
||||||
|
|
||||||
|
backup = os.path.join(os.path.dirname(xlsx), f'{name}_backup_kogl재판정전.xlsx')
|
||||||
|
if not os.path.exists(backup):
|
||||||
|
shutil.copy2(xlsx, backup)
|
||||||
|
wb.save(xlsx)
|
||||||
|
print(f' ▶ {name} 완료: O변경 {o_changed} (미부착化 {to_none}) | 불일치비고 {mismatch} | fetch실패 {fetch_fail}')
|
||||||
|
return dict(name=name, total=len(targets), o_changed=o_changed, to_none=to_none,
|
||||||
|
mismatch=mismatch, fetch_fail=fetch_fail)
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
sel = [a for a in sys.argv[1:] if not a.startswith('-')]
|
||||||
|
sites = [t for t in TARGETS if (not sel or t[0] in sel)]
|
||||||
|
print(f'대상 {len(sites)}개: {[s[0] for s in sites]}')
|
||||||
|
summary = []
|
||||||
|
for name, weak in sites:
|
||||||
|
try:
|
||||||
|
summary.append(process_site(name, weak))
|
||||||
|
except Exception as e:
|
||||||
|
import traceback; traceback.print_exc()
|
||||||
|
summary.append(dict(name=name, total=-1, o_changed=0, to_none=0, mismatch=0, fetch_fail=0))
|
||||||
|
|
||||||
|
print('\n' + '=' * 70 + '\n[전체 요약]')
|
||||||
|
print(f'{"기관":<8}{"부착":>6}{"O변경":>7}{"미부착化":>8}{"불일치":>7}{"실패":>6}')
|
||||||
|
for s in summary:
|
||||||
|
print(f'{s["name"]:<8}{s["total"]:>6}{s["o_changed"]:>7}{s["to_none"]:>8}{s["mismatch"]:>7}{s["fetch_fail"]:>6}')
|
||||||
|
tot = lambda k: sum(x[k] for x in summary if x[k] >= 0)
|
||||||
|
print(f'\n합계: 부착 {tot("total")} | O변경 {tot("o_changed")} | 미부착化 {tot("to_none")} | 불일치비고 {tot("mismatch")} | fetch실패 {tot("fetch_fail")}')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
174
_스크립트/_recurse_dry.log
Normal file
174
_스크립트/_recurse_dry.log
Normal file
@ -0,0 +1,174 @@
|
|||||||
|
시트 .do 페이지행: 407 — 재귀 크롤 시작...
|
||||||
|
크롤한 고유 페이지: 598
|
||||||
|
|
||||||
|
=== 수량 변경(증가/감소) 대상: 26행 ===
|
||||||
|
행306 주민복지 M 6 → 32 https://www.gongju.go.kr/kr/sub06_01_06_07_01.do
|
||||||
|
행398 교통약자 특별교통차량이용안내 M 1 → 18 https://www.gongju.go.kr/kr/sub06_09_03.do
|
||||||
|
행400 공주시 주정차 위반 단속 M 1 → 18 https://www.gongju.go.kr/kr/sub06_09_08.do
|
||||||
|
행401 주정차단속 문자알림 M 1 → 18 https://www.gongju.go.kr/kr/sub06_09_05.do
|
||||||
|
행402 공영주차장 현황 M 1 → 18 https://www.gongju.go.kr/kr/sub06_09_06.do
|
||||||
|
행403 자전거대여 M 1 → 18 https://www.gongju.go.kr/kr/sub06_09_07.do
|
||||||
|
행404 공주시 행복택시 M 1 → 18 https://www.gongju.go.kr/kr/sub06_09_09.do
|
||||||
|
행356 국민재난안전 포털 M 1 → 13 https://www.gongju.go.kr/kr/sub06_07_01.do
|
||||||
|
행357 충청남도 재난안전포털 M 1 → 13 https://www.gongju.go.kr/kr/sub06_07_05.do
|
||||||
|
행358 공주시 재난안전포털 M 1 → 13 https://www.gongju.go.kr/kr/sub06_07_09.do
|
||||||
|
행359 도민안전점검 청구제 M 1 → 13 https://www.gongju.go.kr/kr/sub06_07_02.do
|
||||||
|
행360 비닐하우스 피해경감 농가 행동요령 M 1 → 13 https://www.gongju.go.kr/kr/sub06_07_03.do
|
||||||
|
행361 민방위 정보 M 1 → 13 https://www.gongju.go.kr/kr/sub06_07_04.do
|
||||||
|
행363 시민안전보험 M 1 → 13 https://www.gongju.go.kr/kr/sub06_07_07.do
|
||||||
|
행364 어린이 안전보험 M 1 → 13 https://www.gongju.go.kr/kr/sub06_07_08.do
|
||||||
|
행365 풍수해보험 M 1 → 13 https://www.gongju.go.kr/kr/sub06_07_11.do
|
||||||
|
행367 반려동물을 위한 재난대처법 M 1 → 13 https://www.gongju.go.kr/kr/sub06_07_15.do
|
||||||
|
행110 제안하기 M 1 → 3 https://www.gongju.go.kr/kr/sub03_03_01.do
|
||||||
|
행111 나의제안 M 1 → 3 https://www.gongju.go.kr/kr/sub03_03_02.do
|
||||||
|
행112 공개제안 M 1 → 3 https://www.gongju.go.kr/kr/sub03_03_03.do
|
||||||
|
행334 공주시 행복누림 M 1 → 3 https://www.gongju.go.kr/kr/sub06_03_03.do
|
||||||
|
행335 공주시종합사회복지관 M 1 → 3 https://www.gongju.go.kr/kr/sub06_03_02.do
|
||||||
|
행336 충청남도 사이버교육 M 1 → 3 https://www.gongju.go.kr/kr/sub06_03_04.do
|
||||||
|
행167 미래전략실 M 1 → 2 https://www.gongju.go.kr/prog/deptPerson/kr/sub05_06_28_02/45003470000/list.do
|
||||||
|
행27 자동차등록안내 M 16 → 12 https://www.gongju.go.kr/kr/sub01_06_02_01.do
|
||||||
|
행396 버스정보 M 20 → 11 https://www.gongju.go.kr/kr/sub06_09_01_01.do
|
||||||
|
|
||||||
|
=== ⚠ 하위에 게시판 있어 합치지 않음(M 보류): 140행 ===
|
||||||
|
행19 납세자보호관 (현 M=1, 재귀=3) https://www.gongju.go.kr/kr/sub01_04_03_01.do
|
||||||
|
행20 (현 M=1, 재귀=3) https://www.gongju.go.kr/kr/sub01_04_03_02.do
|
||||||
|
행28 자동차검사 예약,신청조회 (현 M=3, 재귀=3) https://www.gongju.go.kr/kr/sub01_06_03_01.do
|
||||||
|
행33 정보공개제도안내 (현 M=6, 재귀=6) https://www.gongju.go.kr/kr/sub02_15_01_01.do
|
||||||
|
행35 정보공개목록 (현 M=1, 재귀=19) https://www.gongju.go.kr/kr/sub02_15_03.do
|
||||||
|
행36 정보공개목록(구) (현 M=1, 재귀=19) https://www.gongju.go.kr/kr/sub02_15_04.do
|
||||||
|
행38 정보공개청구 (현 M=1, 재귀=19) https://www.gongju.go.kr/kr/sub02_15_06.do
|
||||||
|
행68 공유재산 (현 M=1, 재귀=3) https://www.gongju.go.kr/kr/sub02_20_07_01.do
|
||||||
|
행70 (현 M=1, 재귀=3) https://www.gongju.go.kr/kr/sub02_20_07_03.do
|
||||||
|
행71 (현 M=1, 재귀=3) https://www.gongju.go.kr/kr/sub02_20_07_04.do
|
||||||
|
행75 공공데이터 개방목록 (현 M=1, 재귀=3) https://www.gongju.go.kr/kr/sub02_24_01.do
|
||||||
|
행77 축제 분석 인포그래픽 (현 M=1, 재귀=3) https://www.gongju.go.kr/kr/sub02_24_03.do
|
||||||
|
행81 기부제 안내 (현 M=1, 재귀=5) https://www.gongju.go.kr/kr/sub03_12_01.do
|
||||||
|
행82 답례품 안내 (현 M=1, 재귀=5) https://www.gongju.go.kr/kr/sub03_12_02.do
|
||||||
|
행84 홍보영상 (현 M=1, 재귀=5) https://www.gongju.go.kr/kr/sub03_12_04.do
|
||||||
|
행85 기부자 명예의 전당 (현 M=1, 재귀=5) https://www.gongju.go.kr/kr/sub03_12_05.do
|
||||||
|
행86 공모전 (현 M=1, 재귀=5) https://www.gongju.go.kr/prog/contest/kr/sub03_02_01/list.do
|
||||||
|
행90 공공언어 개선 (현 M=1, 재귀=5) https://www.gongju.go.kr/kr/sub03_02_11.do
|
||||||
|
행91 국민생각함 (현 M=1, 재귀=13) https://www.gongju.go.kr/prog/thinkBoxData/kr/sub03_03_04/getData.do
|
||||||
|
행92 공무원불친절 (현 M=1, 재귀=13) https://www.gongju.go.kr/kr/sub03_04_01.do
|
||||||
|
행93 민원부조리·부패신고 (현 M=1, 재귀=13) https://www.gongju.go.kr/kr/sub03_04_02.do
|
||||||
|
행94 공직자부조리신고센터 (현 M=1, 재귀=13) https://www.gongju.go.kr/kr/sub03_04_04.do
|
||||||
|
행95 예산낭비신고 (현 M=1, 재귀=13) https://www.gongju.go.kr/kr/sub03_04_05.do
|
||||||
|
행96 보조금부정수급신고 (현 M=1, 재귀=13) https://www.gongju.go.kr/kr/sub03_04_06.do
|
||||||
|
행98 안전신문고 (현 M=1, 재귀=13) https://www.gongju.go.kr/kr/sub03_04_09.do
|
||||||
|
행99 식품안전소비자신고 (현 M=1, 재귀=13) https://www.gongju.go.kr/kr/sub03_04_10.do
|
||||||
|
행100 부동산불법거래신고 (현 M=1, 재귀=13) https://www.gongju.go.kr/kr/sub03_04_11.do
|
||||||
|
행101 규제개혁 (현 M=1, 재귀=13) https://www.gongju.go.kr/kr/sub03_04_12.do
|
||||||
|
행102 직장 내 성희롱·성폭력·스토킹 고충상담 (현 M=1, 재귀=13) https://www.gongju.go.kr/kr/sub03_04_13.do
|
||||||
|
행103 공익신고 (현 M=1, 재귀=13) https://www.gongju.go.kr/kr/sub03_04_14.do
|
||||||
|
행109 법령유권해석 (현 M=2, 재귀=2) https://www.gongju.go.kr/kr/sub03_05_05_01.do
|
||||||
|
행116 여성인재DB 사업안내 (현 M=1, 재귀=1) https://www.gongju.go.kr/kr/sub03_10_01.do
|
||||||
|
행171 안전총괄과 (현 M=1, 재귀=2) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_06_05_02/45003590000/list.do
|
||||||
|
행172 (현 M=1, 재귀=2) https://www.gongju.go.kr/kr/sub05_06_05_01.do
|
||||||
|
행193 도시정책과 (현 M=1, 재귀=3) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_06_20_02/45003790000/list.do
|
||||||
|
행194 (현 M=1, 재귀=3) https://www.gongju.go.kr/kr/sub05_06_20_01.do
|
||||||
|
행196 허가건축과 (현 M=1, 재귀=3) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_06_21_02/45003800000/list.do
|
||||||
|
행197 (현 M=1, 재귀=3) https://www.gongju.go.kr/kr/sub05_06_21_01.do
|
||||||
|
행206 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_01_02.do
|
||||||
|
행207 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_01_03/yugu/dongList.do
|
||||||
|
행208 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_01_04/45000520000/list.do
|
||||||
|
행210 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_01_06.do
|
||||||
|
행212 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_02_02.do
|
||||||
|
행213 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_02_03/einmyun/dongList.do
|
||||||
|
행214 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_02_04/45000530000/list.do
|
||||||
|
행216 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_02_06.do
|
||||||
|
행218 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_03_02.do
|
||||||
|
행219 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_03_03/tancheon/dongList.do
|
||||||
|
행220 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_03_04/45000540000/list.do
|
||||||
|
행222 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_03_06.do
|
||||||
|
행224 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_04_02.do
|
||||||
|
행225 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_04_03/gyeryong/dongList.do
|
||||||
|
행226 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_04_04/45000550000/list.do
|
||||||
|
행228 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_04_06.do
|
||||||
|
행230 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_05_02.do
|
||||||
|
행231 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_05_03/banpo/dongList.do
|
||||||
|
행232 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_05_04/45000560000/list.do
|
||||||
|
행234 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_05_06.do
|
||||||
|
행236 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_06_02.do
|
||||||
|
행237 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_06_03/euidang/dongList.do
|
||||||
|
행238 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_06_04/45000580000/list.do
|
||||||
|
행240 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_06_06.do
|
||||||
|
행242 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_07_02.do
|
||||||
|
행243 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_07_03/jungan/dongList.do
|
||||||
|
행244 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_07_04/45000590000/list.do
|
||||||
|
행246 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_07_06.do
|
||||||
|
행248 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_08_02.do
|
||||||
|
행249 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_08_03/woosung/dongList.do
|
||||||
|
행250 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_08_04/45000600000/list.do
|
||||||
|
행252 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_08_06.do
|
||||||
|
행254 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_09_02.do
|
||||||
|
행255 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_09_03/sagok/dongList.do
|
||||||
|
행256 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_09_04/45000610000/list.do
|
||||||
|
행258 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_09_06.do
|
||||||
|
행260 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_10_02.do
|
||||||
|
행261 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_10_03/sinpoong/dongList.do
|
||||||
|
행262 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_10_04/45000620000/list.do
|
||||||
|
행264 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_10_06.do
|
||||||
|
행266 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_11_02.do
|
||||||
|
행267 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_11_03/junghak/dongList.do
|
||||||
|
행268 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_11_04/45000630000/list.do
|
||||||
|
행270 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_11_06.do
|
||||||
|
행272 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_12_02.do
|
||||||
|
행273 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_12_03/ungjin/dongList.do
|
||||||
|
행274 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_12_04/45000660000/list.do
|
||||||
|
행276 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_12_06.do
|
||||||
|
행278 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_13_02.do
|
||||||
|
행279 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_13_03/geumhak/dongList.do
|
||||||
|
행280 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_13_04/45000670000/list.do
|
||||||
|
행282 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_13_06.do
|
||||||
|
행284 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_14_02.do
|
||||||
|
행285 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_14_03/okryong/dongList.do
|
||||||
|
행286 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_14_04/45000680000/list.do
|
||||||
|
행288 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_14_06.do
|
||||||
|
행290 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_15_02.do
|
||||||
|
행291 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_15_03/sinkwan/dongList.do
|
||||||
|
행292 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_15_04/45000690000/list.do
|
||||||
|
행294 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_15_06.do
|
||||||
|
행296 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_16_02.do
|
||||||
|
행297 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/tursmCn/kr/sub05_08_16_03/wolsong/dongList.do
|
||||||
|
행298 (현 M=1, 재귀=6) https://www.gongju.go.kr/prog/deptPerson/kr/sub05_08_16_04/45002340000/list.do
|
||||||
|
행300 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub05_08_16_06.do
|
||||||
|
행307 노인 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub06_01_01_01.do
|
||||||
|
행308 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub06_01_01_02.do
|
||||||
|
행309 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub06_01_01_03.do
|
||||||
|
행310 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub06_01_01_08.do
|
||||||
|
행311 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub06_01_01_06.do
|
||||||
|
행313 장애인 (현 M=10, 재귀=81) https://www.gongju.go.kr/kr/sub06_01_02_01.do
|
||||||
|
행316 (현 M=1, 재귀=47) https://www.gongju.go.kr/kr/sub06_01_09_02_01.do
|
||||||
|
행317 (현 M=1, 재귀=3) https://www.gongju.go.kr/kr/sub06_01_09_03_01.do
|
||||||
|
행318 (현 M=1, 재귀=47) https://www.gongju.go.kr/kr/sub06_01_09_04_03.do
|
||||||
|
행319 (현 M=1, 재귀=46) https://www.gongju.go.kr/kr/sub06_01_09_05.do
|
||||||
|
행320 (현 M=1, 재귀=100) https://www.gongju.go.kr/kr/sub06_01_09_08.do
|
||||||
|
행328 공주시일자리센터 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub06_02_01.do
|
||||||
|
행330 임금체불 등 근로피해신고 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub06_02_04.do
|
||||||
|
행331 취업지원프로그램(고용24) (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub06_02_05.do
|
||||||
|
행332 충남인력개발원 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub06_02_06.do
|
||||||
|
행344 기업 (현 M=1, 재귀=1) https://www.gongju.go.kr/biz/index.do
|
||||||
|
행347 전통시장 (현 M=1, 재귀=28) https://www.gongju.go.kr/kr/sub06_06_01.do
|
||||||
|
행349 (현 M=1, 재귀=2) https://www.gongju.go.kr/kr/sub06_06_02_02.do
|
||||||
|
행352 유가정보서비스 (현 M=2, 재귀=2) https://www.gongju.go.kr/kr/sub06_06_05_01.do
|
||||||
|
행353 사회적·마을 기업 (현 M=4, 재귀=25) https://www.gongju.go.kr/kr/sub06_06_08.do
|
||||||
|
행368 지속가능발전협의회 (현 M=1, 재귀=328) https://www.gongju.go.kr/kr/sub06_08_01.do
|
||||||
|
행371 탄소중립포인트제 (현 M=1, 재귀=328) https://www.gongju.go.kr/kr/sub06_08_04.do
|
||||||
|
행375 야생동식물보호 (현 M=1, 재귀=328) https://www.gongju.go.kr/kr/sub06_08_08.do
|
||||||
|
행376 동물등록제 (현 M=1, 재귀=328) https://www.gongju.go.kr/kr/sub06_08_09.do
|
||||||
|
행377 상수도 (현 M=2, 재귀=5) https://www.gongju.go.kr/kr/sub06_08_12_01.do
|
||||||
|
행378 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub06_08_12_02.do
|
||||||
|
행379 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub06_08_12_03.do
|
||||||
|
행380 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sub06_08_12_04.do
|
||||||
|
행382 하수도 (현 M=1, 재귀=5) https://www.gongju.go.kr/kr/sub06_08_13_01.do
|
||||||
|
행383 (현 M=1, 재귀=5) https://www.gongju.go.kr/kr/sub06_08_13_02_01.do
|
||||||
|
행385 (현 M=1, 재귀=5) https://www.gongju.go.kr/kr/sub06_08_13_04.do
|
||||||
|
행386 (현 M=1, 재귀=5) https://www.gongju.go.kr/kr/sub06_08_13_05.do
|
||||||
|
행388 수질검사 안내 (현 M=1, 재귀=4) https://www.gongju.go.kr/kr/sub06_08_14_01.do
|
||||||
|
행389 (현 M=1, 재귀=4) https://www.gongju.go.kr/kr/sub06_08_14_02.do
|
||||||
|
행391 (현 M=1, 재귀=4) https://www.gongju.go.kr/kr/sub06_08_14_04.do
|
||||||
|
행416 공주시 개인정보처리방침 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sitemap_03_01.do
|
||||||
|
행419 홈페이지 개인정보처리방침 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sitemap_03_03.do
|
||||||
|
행420 표준지방세정보시스템 개인정보 처리방침 (현 M=1, 재귀=6) https://www.gongju.go.kr/kr/sitemap_03_04.do
|
||||||
|
|
||||||
|
(DRY — 적용하려면 --write)
|
||||||
200
_스크립트/_redo_N_all.py
Normal file
200
_스크립트/_redo_N_all.py
Normal file
@ -0,0 +1,200 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""수집완료 기관 N(저작물 유형) 일괄 재검토 (2026-06-01, 당진 규칙 일반화).
|
||||||
|
|
||||||
|
매뉴얼 3-1a/3-3/3-4b 반영:
|
||||||
|
- 이미지 = 별도 정보 주는 것만(사진·평면도·지도·QR·악보·소식지·상징물·도표).
|
||||||
|
- 어문 = 글 설명 인포그래픽(한눈에式)·로고·파트너로고·아이콘·배너·헤드라인텍스트·버튼·팝업·웹접근성마크.
|
||||||
|
- 오디오 = <audio>·mp3/wav/m4a 링크(다운로드 쿼리형 포함).
|
||||||
|
- 지도/PDF 임베드 = 이미지.
|
||||||
|
- 빈 게시판(M=0) = 없음. 사이트(외부) = N 미변경(빈칸 유지).
|
||||||
|
|
||||||
|
N만 재계산하고 L/M/O/P/Q 등 다른 컬럼은 보존. 지역 phase234 모듈의 fetch/get_body/
|
||||||
|
extract_detail_urls/정규식/SITES 재사용. 검수완료(계룡·공주·금산·논산·보령)+당진 제외.
|
||||||
|
|
||||||
|
사용:
|
||||||
|
python -X utf8 _redo_N_all.py dry [기관...] # 변경 로그만(읽기전용)
|
||||||
|
python -X utf8 _redo_N_all.py run [기관...] # 백업 후 N 기입
|
||||||
|
(기관 미지정 시 3개 지역 전체 ✅ 기관)
|
||||||
|
"""
|
||||||
|
import sys, os, io, re, shutil, importlib.util, warnings
|
||||||
|
from concurrent.futures import ThreadPoolExecutor, as_completed
|
||||||
|
import openpyxl
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
HERE = os.path.dirname(os.path.abspath(__file__))
|
||||||
|
MODULES = ['_chungnam_phase234_all.py', '_chungbuk_phase234_all.py', '_jeonbuk_phase234_all.py']
|
||||||
|
EXCLUDE_INST = {'계룡시', '공주시', '금산군', '논산시', '보령시', '당진시'} # 검수완료 + 당진
|
||||||
|
|
||||||
|
# ── 이미지 분류 (벤더 공통) ───────────────────────────────
|
||||||
|
DECO = re.compile(r'/common/|move\.png|no[-_]?img|blank|spacer|/ico|/btn|bullet|arrow|/bg|icon|see_btn|/sample|mimetype|/file_|filedown|btn_dir', re.I)
|
||||||
|
EXCLUDE = re.compile(
|
||||||
|
r'한눈에|흐름도|절차도|처리절차|이용절차'
|
||||||
|
r'|로고(?!송)|logo(?!song)|아이콘|배너|banner'
|
||||||
|
r'|신문고|relation_item|tracer|headline'
|
||||||
|
r'|카피라이트|copyright|copy_logo|popup|/pup/|wa_mk|웹접근성|품질인증', re.I)
|
||||||
|
AUDIO = re.compile(r'\.(?:mp3|wav|m4a|ogg|flac)\b', re.I)
|
||||||
|
MAPPDF = re.compile(r'pdf|viewer\.html|/map|kakao|daum.*map', re.I)
|
||||||
|
|
||||||
|
|
||||||
|
def load_module(fname):
|
||||||
|
spec = importlib.util.spec_from_file_location(fname[:-3], os.path.join(HERE, fname))
|
||||||
|
m = importlib.util.module_from_spec(spec); spec.loader.exec_module(m)
|
||||||
|
return m
|
||||||
|
|
||||||
|
|
||||||
|
def real_imgs(M, body):
|
||||||
|
out = []
|
||||||
|
for img in body.find_all('img'):
|
||||||
|
src = img.get('src') or ''
|
||||||
|
if not src or M.KOGL_IMG_PAT.search(src) or DECO.search(src):
|
||||||
|
continue
|
||||||
|
if EXCLUDE.search(src + ' ' + (img.get('alt') or '')):
|
||||||
|
continue
|
||||||
|
out.append(img)
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def media2(M, body):
|
||||||
|
has_text = len(body.get_text(strip=True)) > 30
|
||||||
|
img = len(real_imgs(M, body)) > 0
|
||||||
|
vid = False
|
||||||
|
for ifr in body.find_all('iframe'):
|
||||||
|
s = ifr.get('src') or ''
|
||||||
|
if M.YOUTUBE_PAT.search(s):
|
||||||
|
vid = True
|
||||||
|
elif MAPPDF.search(s):
|
||||||
|
img = True
|
||||||
|
if not vid and (body.find('video') or body.find('a', href=M.YOUTUBE_PAT) or M.VIDEO_EXT.search(str(body))):
|
||||||
|
vid = True
|
||||||
|
aud = bool(body.find('audio')) or bool(AUDIO.search(str(body)))
|
||||||
|
return img, vid, aud, has_text
|
||||||
|
|
||||||
|
|
||||||
|
def make_fetch(M, cfg):
|
||||||
|
"""모듈별 fetch 시그니처 차이 흡수: 충남=fetch(url), 충북/전북=fetch(session,url)."""
|
||||||
|
if hasattr(M, 'make_session'):
|
||||||
|
sess = M.make_session(weak_ssl=cfg.get('weak_ssl', False))
|
||||||
|
return lambda u: M.fetch(sess, u)
|
||||||
|
return lambda u: M.fetch(u)
|
||||||
|
|
||||||
|
|
||||||
|
def _fetch_body(M, url, body_sel, fetchfn, tries=3):
|
||||||
|
"""throttle 대비 재시도. 본문 텍스트>30 또는 미디어 잡히면 즉시 반환."""
|
||||||
|
import time
|
||||||
|
last = None
|
||||||
|
for k in range(tries):
|
||||||
|
soup = fetchfn(url)
|
||||||
|
if soup is not None:
|
||||||
|
body = M.get_body(soup, body_sel)
|
||||||
|
last = media2(M, body) + (body,)
|
||||||
|
if last[3] or last[0] or last[1] or last[2]: # txt/img/vid/aud 중 하나라도
|
||||||
|
return last
|
||||||
|
time.sleep(0.6 * (k + 1))
|
||||||
|
return last # 끝까지 비면 마지막(또는 None)
|
||||||
|
|
||||||
|
|
||||||
|
def n_of(M, url, L, body_sel, fetchfn):
|
||||||
|
res = _fetch_body(M, url, body_sel, fetchfn)
|
||||||
|
if res is None:
|
||||||
|
return None # 접근 실패 → 기존 N 유지
|
||||||
|
img, vid, aud, txt, body = res
|
||||||
|
if not (txt or img or vid or aud):
|
||||||
|
return None # 재시도해도 빈 본문 → 신뢰불가, 기존 N 유지(거짓 '없음' 차단)
|
||||||
|
if L == '게시판':
|
||||||
|
detail_urls = M.extract_detail_urls(body, url, limit=5)
|
||||||
|
data_rows = [tr for tr in body.select('table tbody tr, .board_list li, ul.bbs_list li') if tr.find('a')]
|
||||||
|
empty_msg = bool(re.search(r'게시물이?\s*없|등록된\s*(?:게시물|자료)\s*가?\s*없|자료가\s*없', body.get_text(' ', strip=True)))
|
||||||
|
if not detail_urls and not data_rows and (empty_msg or not txt):
|
||||||
|
return '없음' # 실제 글 0개 = 진짜 빈 게시판 (M값 무시, 내용기반)
|
||||||
|
for du in detail_urls:
|
||||||
|
ds = fetchfn(du)
|
||||||
|
if not ds:
|
||||||
|
continue
|
||||||
|
db = M.get_body(ds, body_sel)
|
||||||
|
i2, v2, a2, t2 = media2(M, db)
|
||||||
|
img = img or i2; vid = vid or v2; aud = aud or a2; txt = txt or t2
|
||||||
|
parts = []
|
||||||
|
if txt:
|
||||||
|
parts.append('어문')
|
||||||
|
if img:
|
||||||
|
parts.append('이미지')
|
||||||
|
if vid:
|
||||||
|
parts.append('영상')
|
||||||
|
if aud:
|
||||||
|
parts.append('오디오')
|
||||||
|
return ','.join(parts) if parts else '없음'
|
||||||
|
|
||||||
|
|
||||||
|
def do_inst(M, name, cfg, mode, log):
|
||||||
|
xlsx = cfg['xlsx']
|
||||||
|
if not os.path.exists(xlsx):
|
||||||
|
log.write(f'[{name}] 엑셀 없음\n'); return (name, 0, 0)
|
||||||
|
body_sel = cfg['body_sel']
|
||||||
|
fetchfn = make_fetch(M, cfg)
|
||||||
|
wb = openpyxl.load_workbook(xlsx); ws = wb.active
|
||||||
|
jobs = []
|
||||||
|
for r in range(3, ws.max_row + 1):
|
||||||
|
L = ws.cell(r, 12).value
|
||||||
|
K = ws.cell(r, 11).value
|
||||||
|
Mq = ws.cell(r, 13).value
|
||||||
|
if not (isinstance(K, str) and K.startswith('http')):
|
||||||
|
continue
|
||||||
|
if L == '사이트':
|
||||||
|
continue
|
||||||
|
if L not in ('페이지', '게시판'):
|
||||||
|
continue
|
||||||
|
jobs.append((r, K, L)) # M=0이어도 내용기반으로 n_of가 판정(거짓 M=0 대응)
|
||||||
|
results = {}
|
||||||
|
with ThreadPoolExecutor(max_workers=4) as ex:
|
||||||
|
futs = {}
|
||||||
|
for r, K, L in jobs:
|
||||||
|
if K == '__EMPTY__':
|
||||||
|
results[r] = '없음'
|
||||||
|
else:
|
||||||
|
futs[ex.submit(n_of, M, K, L, body_sel, fetchfn)] = r
|
||||||
|
for f in as_completed(futs):
|
||||||
|
r = futs[f]
|
||||||
|
try:
|
||||||
|
results[r] = f.result()
|
||||||
|
except Exception:
|
||||||
|
results[r] = None
|
||||||
|
changed = 0
|
||||||
|
for r, newN in sorted(results.items()):
|
||||||
|
if newN is None:
|
||||||
|
continue
|
||||||
|
oldN = ws.cell(r, 14).value
|
||||||
|
if newN != oldN:
|
||||||
|
changed += 1
|
||||||
|
log.write(f' [{name}] r{r} {oldN}→{newN}\n')
|
||||||
|
if mode == 'run':
|
||||||
|
ws.cell(r, 14).value = newN
|
||||||
|
total = len([j for j in jobs])
|
||||||
|
if mode == 'run' and changed:
|
||||||
|
bak = xlsx.replace('.xlsx', '_backup_N재검토전.xlsx')
|
||||||
|
shutil.copy(xlsx, bak)
|
||||||
|
wb.save(xlsx)
|
||||||
|
return (name, total, changed)
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
mode = sys.argv[1] if len(sys.argv) > 1 else 'dry'
|
||||||
|
only = set(a for a in sys.argv[2:] if not a.startswith('--'))
|
||||||
|
log = io.open(os.path.join(HERE, '_redo_N_all_log.txt'), 'w', encoding='utf-8')
|
||||||
|
grand = []
|
||||||
|
for fname in MODULES:
|
||||||
|
M = load_module(fname)
|
||||||
|
for name, cfg in M.SITES.items():
|
||||||
|
if name in EXCLUDE_INST:
|
||||||
|
continue
|
||||||
|
if only and name not in only:
|
||||||
|
continue
|
||||||
|
res = do_inst(M, name, cfg, mode, log)
|
||||||
|
grand.append(res)
|
||||||
|
print(f'{name:7} 처리 {res[1]:4}행 | N변경 {res[2]:4} ({mode})', flush=True)
|
||||||
|
log.close()
|
||||||
|
print('=' * 50)
|
||||||
|
print(f'총 {len(grand)}개 기관 | 변경합계 {sum(r[2] for r in grand)}행 ({mode}) | 로그 _redo_N_all_log.txt')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
60
_스크립트/_redo_N_all_log.txt
Normal file
60
_스크립트/_redo_N_all_log.txt
Normal file
@ -0,0 +1,60 @@
|
|||||||
|
[무주군] r3 어문,이미지→어문
|
||||||
|
[무주군] r4 어문,이미지→어문
|
||||||
|
[무주군] r6 어문,이미지→어문
|
||||||
|
[무주군] r12 어문,이미지→어문
|
||||||
|
[무주군] r13 어문,이미지→어문
|
||||||
|
[무주군] r14 어문,이미지→어문
|
||||||
|
[무주군] r15 어문,이미지→어문
|
||||||
|
[무주군] r17 어문,이미지→어문
|
||||||
|
[무주군] r26 어문,이미지→어문
|
||||||
|
[무주군] r27 어문,이미지→어문
|
||||||
|
[무주군] r30 어문,이미지→어문
|
||||||
|
[무주군] r89 어문,이미지→어문
|
||||||
|
[무주군] r90 어문,이미지→어문
|
||||||
|
[무주군] r91 어문,이미지→어문
|
||||||
|
[무주군] r92 어문,이미지→어문
|
||||||
|
[무주군] r99 어문,이미지→어문
|
||||||
|
[무주군] r103 어문,이미지→어문
|
||||||
|
[무주군] r116 어문,이미지→어문
|
||||||
|
[무주군] r120 어문,이미지→어문
|
||||||
|
[무주군] r121 어문,이미지→어문
|
||||||
|
[무주군] r126 어문,이미지→어문
|
||||||
|
[무주군] r134 어문,이미지→어문
|
||||||
|
[무주군] r155 어문,이미지→어문
|
||||||
|
[무주군] r156 어문,이미지→어문
|
||||||
|
[무주군] r157 어문,이미지→어문
|
||||||
|
[무주군] r158 어문,이미지→어문
|
||||||
|
[무주군] r159 어문,이미지→어문
|
||||||
|
[무주군] r163 어문,이미지→어문
|
||||||
|
[무주군] r164 어문,이미지→어문
|
||||||
|
[무주군] r165 어문,이미지→어문
|
||||||
|
[무주군] r185 어문,이미지→어문
|
||||||
|
[무주군] r186 어문,이미지→어문
|
||||||
|
[무주군] r187 어문,이미지→어문
|
||||||
|
[무주군] r188 어문,이미지→어문
|
||||||
|
[무주군] r189 어문,이미지→어문
|
||||||
|
[무주군] r190 어문,이미지→어문
|
||||||
|
[무주군] r192 어문,이미지→어문
|
||||||
|
[무주군] r193 어문,이미지→어문
|
||||||
|
[무주군] r194 어문,이미지→어문
|
||||||
|
[무주군] r195 어문,이미지→어문
|
||||||
|
[무주군] r196 어문,이미지→어문
|
||||||
|
[무주군] r197 어문,이미지→어문
|
||||||
|
[무주군] r198 어문,이미지→어문
|
||||||
|
[무주군] r199 어문,이미지→어문
|
||||||
|
[무주군] r200 어문,이미지→어문
|
||||||
|
[무주군] r207 어문,이미지→어문
|
||||||
|
[무주군] r239 어문,이미지,영상→어문,영상
|
||||||
|
[무주군] r247 어문,이미지→어문
|
||||||
|
[무주군] r262 어문,이미지→어문
|
||||||
|
[무주군] r263 어문,이미지→어문
|
||||||
|
[무주군] r314 어문,이미지→어문
|
||||||
|
[무주군] r323 어문,이미지→어문
|
||||||
|
[무주군] r326 어문,이미지→어문
|
||||||
|
[무주군] r349 어문,이미지→어문
|
||||||
|
[무주군] r350 어문,이미지→어문
|
||||||
|
[무주군] r360 어문,이미지→어문
|
||||||
|
[무주군] r362 어문,이미지→어문
|
||||||
|
[무주군] r368 어문,이미지→어문
|
||||||
|
[무주군] r370 어문,이미지→어문
|
||||||
|
[무주군] r379 어문,이미지→어문
|
||||||
101
_스크립트/_rename_prefix.py
Normal file
101
_스크립트/_rename_prefix.py
Normal file
@ -0,0 +1,101 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""작업파일명 앞에 광역(도) 접두 — {시}.xlsx → {도}_{시}.xlsx (2026-06-01).
|
||||||
|
|
||||||
|
1) 광역_사이트맵/{도}/{N.시}/{시}.xlsx 메인 파일 rename (백업·부수파일 제외).
|
||||||
|
2) 모든 .py(_스크립트/ + 광역_사이트맵/**)에서 리터럴 '{시}.xlsx' → '{도}_{시}.xlsx' 치환.
|
||||||
|
(충남·충북 phase234 SITES 하드코딩 경로, 기관폴더 스크립트 등)
|
||||||
|
※ 동적 빌더(f'{name}.xlsx' 등 5곳)는 별도 수동 수정.
|
||||||
|
|
||||||
|
사용: python -X utf8 _rename_prefix.py dry|run
|
||||||
|
"""
|
||||||
|
import os, sys, io, re
|
||||||
|
|
||||||
|
ROOT = r'D:\01.프로젝트\DB수집'
|
||||||
|
MAP = os.path.join(ROOT, '작업파일', '광역_사이트맵')
|
||||||
|
PROVS = ['충청남도', '충청북도', '전북특별자치도', '제주특별자치도']
|
||||||
|
|
||||||
|
|
||||||
|
def build_map():
|
||||||
|
"""반환 [(도, 폴더명(N.시), 시, 기존xlsx경로, 새xlsx경로)]"""
|
||||||
|
out = []
|
||||||
|
for prov in PROVS:
|
||||||
|
pdir = os.path.join(MAP, prov)
|
||||||
|
if not os.path.isdir(pdir):
|
||||||
|
continue
|
||||||
|
for fold in sorted(os.listdir(pdir)):
|
||||||
|
fdir = os.path.join(pdir, fold)
|
||||||
|
if not os.path.isdir(fdir):
|
||||||
|
continue
|
||||||
|
si = fold.split('.', 1)[-1] # 'N.시' → 시
|
||||||
|
old = os.path.join(fdir, f'{si}.xlsx')
|
||||||
|
if os.path.exists(old):
|
||||||
|
new = os.path.join(fdir, f'{prov}_{si}.xlsx')
|
||||||
|
out.append((prov, fold, si, old, new))
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def py_files():
|
||||||
|
fs = []
|
||||||
|
sdir = os.path.join(ROOT, '_스크립트')
|
||||||
|
for f in os.listdir(sdir):
|
||||||
|
if f.endswith('.py'):
|
||||||
|
fs.append(os.path.join(sdir, f))
|
||||||
|
for dp, _, names in os.walk(MAP):
|
||||||
|
for n in names:
|
||||||
|
if n.endswith('.py'):
|
||||||
|
fs.append(os.path.join(dp, n))
|
||||||
|
return fs
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
mode = sys.argv[1] if len(sys.argv) > 1 else 'dry'
|
||||||
|
m = build_map()
|
||||||
|
# 시→도 (리터럴 치환용). 시명 길이 내림차순(부분일치 방지)
|
||||||
|
s2p = {si: prov for prov, fold, si, o, n in m}
|
||||||
|
print(f'== rename 대상 {len(m)}개 ==')
|
||||||
|
for prov, fold, si, o, n in m:
|
||||||
|
print(f' {prov}/{fold}/{si}.xlsx → {prov}_{si}.xlsx')
|
||||||
|
if mode == 'run':
|
||||||
|
if os.path.exists(n):
|
||||||
|
print(' !! 이미 존재, 스킵'); continue
|
||||||
|
os.rename(o, n)
|
||||||
|
# 리터럴 치환
|
||||||
|
print('\n== .py 리터럴 치환 ==')
|
||||||
|
keys = sorted(s2p.keys(), key=len, reverse=True)
|
||||||
|
total_files = 0; total_hits = 0
|
||||||
|
for fp in py_files():
|
||||||
|
try:
|
||||||
|
txt = io.open(fp, encoding='utf-8').read()
|
||||||
|
except Exception:
|
||||||
|
continue
|
||||||
|
orig = txt; hits = 0
|
||||||
|
for si in keys:
|
||||||
|
prov = s2p[si]
|
||||||
|
pat = si + '.xlsx'
|
||||||
|
rep = f'{prov}_{si}.xlsx'
|
||||||
|
# 이미 접두된 경우 중복 방지: {prov}_{si}.xlsx 는 건드리지 않음
|
||||||
|
# 음수 룩비하인드로 바로 앞이 '_'(도접두 직후)나 한글이면 스킵
|
||||||
|
def _sub(mo):
|
||||||
|
start = mo.start()
|
||||||
|
before = txt[max(0, start - 1):start]
|
||||||
|
# 직전이 '_' 또는 한글(다른 시명 꼬리)이면 치환 안 함
|
||||||
|
if before and (before == '_' or '가' <= before <= '힣'):
|
||||||
|
return mo.group(0)
|
||||||
|
return rep
|
||||||
|
new = re.sub(re.escape(pat), _sub, txt)
|
||||||
|
if new != txt:
|
||||||
|
hits += new.count(rep) - orig.count(rep)
|
||||||
|
txt = new
|
||||||
|
if txt != orig:
|
||||||
|
total_files += 1
|
||||||
|
n_changed = sum(txt.count(f'{s2p[si]}_{si}.xlsx') for si in keys) - sum(orig.count(f'{s2p[si]}_{si}.xlsx') for si in keys)
|
||||||
|
print(f' {os.path.relpath(fp, ROOT)} (+{n_changed})')
|
||||||
|
total_hits += n_changed
|
||||||
|
if mode == 'run':
|
||||||
|
io.open(fp, 'w', encoding='utf-8').write(txt)
|
||||||
|
print(f'\n치환 파일 {total_files}개, 경로 {total_hits}건 ({mode})')
|
||||||
|
print('※ 동적 빌더 수동수정 필요: _dedup_all.py xpath, _tab_batch.py xlsx_path, _jeonbuk_phase234_all.py site(), phase1 3개')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
59
_스크립트/_scan_textonly.py
Normal file
59
_스크립트/_scan_textonly.py
Normal file
@ -0,0 +1,59 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""조회 전용(저장 안 함): 부착 행 중 이미지(img_opentype) 없이
|
||||||
|
링크(licenseType)만 있는 '텍스트 부착' 행의 URL을 시·군별로 나열."""
|
||||||
|
import sys, os, openpyxl
|
||||||
|
from concurrent.futures import ThreadPoolExecutor
|
||||||
|
|
||||||
|
sys.path.insert(0, r'D:\01.프로젝트\DB수집\_스크립트')
|
||||||
|
sys.stdout.reconfigure(encoding='utf-8')
|
||||||
|
|
||||||
|
import _recheck_kogl_all as RC
|
||||||
|
|
||||||
|
O_COL, K_COL = 15, 11
|
||||||
|
|
||||||
|
|
||||||
|
def scan(name, region, weak):
|
||||||
|
mod = RC.CN if region == 'cn' else RC.CB
|
||||||
|
cfg = RC.get_cfg(name, region)
|
||||||
|
xlsx, body_sel = cfg['xlsx'], cfg['body_sel']
|
||||||
|
wb = openpyxl.load_workbook(xlsx); ws = wb.active
|
||||||
|
targets = []
|
||||||
|
for r in range(3, ws.max_row + 1):
|
||||||
|
o = ws.cell(r, O_COL).value
|
||||||
|
if o and str(o).strip() != '미부착' and ws.cell(r, K_COL).value:
|
||||||
|
targets.append((r, str(ws.cell(r, K_COL).value).strip(), o))
|
||||||
|
if not targets:
|
||||||
|
return name, []
|
||||||
|
fetch_fn = RC.make_fetch(region, weak)
|
||||||
|
|
||||||
|
def work(t):
|
||||||
|
r, url, oldO = t
|
||||||
|
img_t, link_t, st = RC.detect_row(url, body_sel, mod, fetch_fn)
|
||||||
|
return (r, url, oldO, img_t, link_t, st)
|
||||||
|
|
||||||
|
out = []
|
||||||
|
with ThreadPoolExecutor(max_workers=6) as ex:
|
||||||
|
for r, url, oldO, img_t, link_t, st in ex.map(work, targets):
|
||||||
|
if st == 'ok' and not img_t and link_t:
|
||||||
|
out.append((r, url, str(oldO).strip(), sorted(link_t)))
|
||||||
|
return name, sorted(out)
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
sel = [a for a in sys.argv[1:] if not a.startswith('-')]
|
||||||
|
sites = [t for t in RC.TARGETS if (not sel or t[0] in sel)]
|
||||||
|
grand = 0
|
||||||
|
for name, region, weak in sites:
|
||||||
|
nm, rows = scan(name, region, weak)
|
||||||
|
if rows:
|
||||||
|
print(f'\n■ {nm} — 텍스트(링크)만 부착 {len(rows)}건')
|
||||||
|
for r, url, oldO, lk in rows:
|
||||||
|
print(f' row{r:>3} | 기존O={oldO:<14} | 링크={lk} | {url}')
|
||||||
|
grand += len(rows)
|
||||||
|
else:
|
||||||
|
print(f'■ {nm} — 없음')
|
||||||
|
print(f'\n총 텍스트-only 부착 행: {grand}건')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
1009
_스크립트/_split_plan.json
Normal file
1009
_스크립트/_split_plan.json
Normal file
File diff suppressed because it is too large
Load Diff
174
_스크립트/_tab_batch.py
Normal file
174
_스크립트/_tab_batch.py
Normal file
@ -0,0 +1,174 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""수집완료(✅) 기관 일괄 본문탭(1-5b) 확장 + 신규행 Phase2~4 수집 오케스트레이터.
|
||||||
|
|
||||||
|
매뉴얼 1-5b 의 _tab_expand.py / _tab_phase234.py 를 기관 표대로 순차 호출한다.
|
||||||
|
- 검수완료(계룡시)·미처리(보은군)·이미 탭확장(공주시)는 제외.
|
||||||
|
- base/domain 은 각 엑셀 K열 URL에서 자동 도출(가장 흔한 host 기준).
|
||||||
|
- 영동군만 가중 SSL(--weak-ssl).
|
||||||
|
|
||||||
|
모드:
|
||||||
|
python -X utf8 _tab_batch.py probe # 설정·행수·base/domain 검증만
|
||||||
|
python -X utf8 _tab_batch.py dry # 전 기관 확장계획(읽기전용)만 산출
|
||||||
|
python -X utf8 _tab_batch.py run [기관...] # 확장(--write)+Phase234 실제 수행
|
||||||
|
"""
|
||||||
|
import os
|
||||||
|
import re
|
||||||
|
import sys
|
||||||
|
import subprocess
|
||||||
|
from collections import Counter
|
||||||
|
from urllib.parse import urlsplit
|
||||||
|
|
||||||
|
import openpyxl
|
||||||
|
|
||||||
|
HERE = os.path.dirname(os.path.abspath(__file__))
|
||||||
|
ROOT = os.path.dirname(HERE)
|
||||||
|
MAP = os.path.join(ROOT, '작업파일', '광역_사이트맵')
|
||||||
|
|
||||||
|
CN_ALL = os.path.join(HERE, '_chungnam_phase234_all.py')
|
||||||
|
CB_ALL = os.path.join(HERE, '_chungbuk_phase234_all.py')
|
||||||
|
JB_ALL = os.path.join(HERE, '_jeonbuk_phase234_all.py')
|
||||||
|
GEUMSAN = os.path.join(MAP, '충청남도', '3.금산군', '_phase234.py')
|
||||||
|
JB_SEL = '#main-contents,#content,#contents,.contents,#txt,main,#container,#sub'
|
||||||
|
|
||||||
|
# (광역, idx, 기관, 폴더명, phase234모듈, body_sel(콤마/None), weak_ssl)
|
||||||
|
SITES = [
|
||||||
|
# 충청남도 (계룡=검수완료 제외, 공주=이미 탭확장 제외)
|
||||||
|
('충청남도', 3, '금산군', GEUMSAN, None, False),
|
||||||
|
('충청남도', 4, '논산시', CN_ALL, '#txt,#contents,main', False),
|
||||||
|
('충청남도', 5, '당진시', CN_ALL, '#txt,#contents,main', False),
|
||||||
|
('충청남도', 6, '보령시', CN_ALL, '#txt,#contents,main', False),
|
||||||
|
('충청남도', 7, '부여군', CN_ALL, '#txt,#contents,main', False),
|
||||||
|
('충청남도', 8, '서산시', CN_ALL, '#contents,#txt,main', False),
|
||||||
|
('충청남도', 9, '서천군', CN_ALL, '#txt,#contents,main', False),
|
||||||
|
('충청남도', 10, '아산시', CN_ALL, '#contents,main,#txt', False),
|
||||||
|
('충청남도', 11, '예산군', CN_ALL, '#txt,#contents,main', False),
|
||||||
|
('충청남도', 12, '천안시', CN_ALL, '#txt,#contents,main', False),
|
||||||
|
('충청남도', 13, '청양군', CN_ALL, '#txt,#contents,main', False),
|
||||||
|
('충청남도', 14, '태안군', CN_ALL, '#txt,#contents,main', False),
|
||||||
|
('충청남도', 15, '홍성군', CN_ALL, '#txt,#contents,main', False),
|
||||||
|
# 충청북도 (보은=미처리 제외)
|
||||||
|
('충청북도', 1, '괴산군', CB_ALL, '#contents,#txt,main', False),
|
||||||
|
('충청북도', 2, '단양군', CB_ALL, '#contents,#txt,main', False),
|
||||||
|
('충청북도', 4, '영동군', CB_ALL, '#txt,#contents,main', True),
|
||||||
|
('충청북도', 5, '옥천군', CB_ALL, '#contents,#txt,main', False),
|
||||||
|
('충청북도', 6, '음성군', CB_ALL, '#contents,#txt,main', False),
|
||||||
|
('충청북도', 7, '제천시', CB_ALL, '#contents,#txt,main', False),
|
||||||
|
('충청북도', 8, '증평군', CB_ALL, '#txt,#contents,main', False),
|
||||||
|
('충청북도', 9, '진천군', CB_ALL, '#contents,#txt,main', False),
|
||||||
|
('충청북도', 10, '청주시', CB_ALL, '#contents,#txt,main', False),
|
||||||
|
('충청북도', 11, '충주시', CB_ALL, '#contents,#txt,main', False),
|
||||||
|
# 전북특별자치도
|
||||||
|
('전북특별자치도', 1, '고창군', JB_ALL, JB_SEL, False),
|
||||||
|
('전북특별자치도', 2, '군산시', JB_ALL, JB_SEL, False),
|
||||||
|
('전북특별자치도', 3, '김제시', JB_ALL, JB_SEL, False),
|
||||||
|
('전북특별자치도', 4, '남원시', JB_ALL, JB_SEL, False),
|
||||||
|
('전북특별자치도', 5, '무주군', JB_ALL, JB_SEL, False),
|
||||||
|
('전북특별자치도', 6, '부안군', JB_ALL, JB_SEL, False),
|
||||||
|
('전북특별자치도', 7, '순창군', JB_ALL, JB_SEL, False),
|
||||||
|
('전북특별자치도', 8, '완주군', JB_ALL, JB_SEL, False),
|
||||||
|
('전북특별자치도', 9, '익산시', JB_ALL, JB_SEL, False),
|
||||||
|
('전북특별자치도', 10, '임실군', JB_ALL, JB_SEL, False),
|
||||||
|
('전북특별자치도', 11, '장수군', JB_ALL, JB_SEL, False),
|
||||||
|
('전북특별자치도', 12, '전주시', JB_ALL, JB_SEL, False),
|
||||||
|
('전북특별자치도', 13, '정읍시', JB_ALL, JB_SEL, False),
|
||||||
|
('전북특별자치도', 14, '진안군', JB_ALL, JB_SEL, False),
|
||||||
|
# 제주특별자치도
|
||||||
|
('제주특별자치도', 1, '서귀포시', JB_ALL, JB_SEL, False),
|
||||||
|
('제주특별자치도', 2, '제주시', JB_ALL, JB_SEL, False),
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def xlsx_path(prov, idx, name):
|
||||||
|
return os.path.join(MAP, prov, f'{idx}.{name}', f'{prov}_{name}.xlsx')
|
||||||
|
|
||||||
|
|
||||||
|
def derive_base_domain(xlsx):
|
||||||
|
"""K열 URL에서 가장 흔한 host → base(scheme://host), domain(끝 3라벨)."""
|
||||||
|
wb = openpyxl.load_workbook(xlsx, read_only=True)
|
||||||
|
ws = wb.active
|
||||||
|
hosts = Counter()
|
||||||
|
scheme_by_host = {}
|
||||||
|
nrow = 0
|
||||||
|
for r in ws.iter_rows(min_row=3, min_col=11, max_col=11, values_only=True):
|
||||||
|
u = r[0]
|
||||||
|
if isinstance(u, str) and u.startswith('http'):
|
||||||
|
nrow += 1
|
||||||
|
sp = urlsplit(u)
|
||||||
|
if sp.hostname:
|
||||||
|
hosts[sp.hostname] += 1
|
||||||
|
scheme_by_host.setdefault(sp.hostname, sp.scheme)
|
||||||
|
# 데이터 행수(K 무관)
|
||||||
|
maxrow = ws.max_row
|
||||||
|
wb.close()
|
||||||
|
if not hosts:
|
||||||
|
return None, None, maxrow
|
||||||
|
host = hosts.most_common(1)[0][0]
|
||||||
|
base = f'{scheme_by_host[host]}://{host}'
|
||||||
|
labels = host.split('.')
|
||||||
|
domain = '.'.join(labels[-3:]) if len(labels) >= 3 else host
|
||||||
|
return base, domain, maxrow
|
||||||
|
|
||||||
|
|
||||||
|
def run(cmd):
|
||||||
|
print(' $', ' '.join(os.path.basename(c) if c.endswith('.py') else c for c in cmd[2:]))
|
||||||
|
p = subprocess.run(cmd, capture_output=True, text=True, encoding='utf-8', errors='replace')
|
||||||
|
out = (p.stdout or '') + (p.stderr or '')
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
mode = sys.argv[1] if len(sys.argv) > 1 else 'probe'
|
||||||
|
only = set(sys.argv[2:]) if len(sys.argv) > 2 else None
|
||||||
|
py = [sys.executable, '-X', 'utf8']
|
||||||
|
|
||||||
|
sites = [s for s in SITES if not only or s[2] in only]
|
||||||
|
print(f'대상 기관: {len(sites)}개 (모드={mode})\n')
|
||||||
|
|
||||||
|
for prov, idx, name, mod, sel, weak in sites:
|
||||||
|
xp = xlsx_path(prov, idx, name)
|
||||||
|
if not os.path.exists(xp):
|
||||||
|
print(f'❌ {prov} {name}: 엑셀 없음 → {xp}\n')
|
||||||
|
continue
|
||||||
|
base, domain, maxrow = derive_base_domain(xp)
|
||||||
|
tag = ' [weak-ssl]' if weak else ''
|
||||||
|
print(f'━━━ {prov} {name}{tag} (행~{maxrow}, base={base}, domain={domain})')
|
||||||
|
|
||||||
|
if mode == 'probe':
|
||||||
|
print(f' 모듈={os.path.basename(mod)} body_sel={sel}\n')
|
||||||
|
continue
|
||||||
|
|
||||||
|
# 1) 항상 dry 로 먼저 확장계획 산출(읽기전용)
|
||||||
|
ex = py + [os.path.join(HERE, '_tab_expand.py'), xp, base, domain]
|
||||||
|
if weak:
|
||||||
|
ex += ['--weak-ssl']
|
||||||
|
out = run(ex)
|
||||||
|
m = re.search(r'신규 추가 행:\s*(\d+)개', out)
|
||||||
|
newcnt = int(m.group(1)) if m else 0
|
||||||
|
grp = re.search(r'탭 확장 계획:\s*(\d+)개', out)
|
||||||
|
print(f' → 확장그룹 {grp.group(1) if grp else "?"}개 / 신규행 {newcnt}개')
|
||||||
|
if mode == 'dry':
|
||||||
|
print()
|
||||||
|
continue
|
||||||
|
|
||||||
|
# 2) run 모드 & 신규행 있을 때만 --write 후 Phase234
|
||||||
|
if newcnt > 0:
|
||||||
|
exw = ex + ['--write']
|
||||||
|
run(exw)
|
||||||
|
ph = py + [os.path.join(HERE, '_tab_phase234.py'), xp, mod, domain]
|
||||||
|
if sel:
|
||||||
|
ph += [sel]
|
||||||
|
if weak:
|
||||||
|
ph += ['--weak-ssl']
|
||||||
|
out2 = run(ph)
|
||||||
|
for line in out2.splitlines():
|
||||||
|
if any(k in line for k in ('대상 신규행', 'L 분포', 'O 분포', '접근실패', '저장', '대체 저장')):
|
||||||
|
print(' ' + line)
|
||||||
|
else:
|
||||||
|
print(' (신규행 없음 → Phase234 생략)')
|
||||||
|
print()
|
||||||
|
|
||||||
|
print('\n=== 배치 종료 ===')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
81
_스크립트/_tab_classfind.py
Normal file
81
_스크립트/_tab_classfind.py
Normal file
@ -0,0 +1,81 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""클래스명 무관 본문탭 탐지: 표본 페이지에서 '자기 URL을 포함하면서 2+실제링크 &
|
||||||
|
1+ 신규(엑셀에 없는) URL'을 가진 UL을 찾아 그 class 를 집계. 1-5b 조건2~4의 클래스 비의존 버전.
|
||||||
|
미지의 탭 클래스명을 발견하기 위함.
|
||||||
|
|
||||||
|
사용: python -X utf8 _tab_classfind.py <엑셀> <base> <도메인> [표본=50] [--weak-ssl]
|
||||||
|
"""
|
||||||
|
import sys, warnings, ssl
|
||||||
|
from collections import Counter
|
||||||
|
from urllib.parse import urlsplit, urljoin
|
||||||
|
from concurrent.futures import ThreadPoolExecutor, as_completed
|
||||||
|
import openpyxl, requests
|
||||||
|
from requests.adapters import HTTPAdapter
|
||||||
|
from urllib3.util.ssl_ import create_urllib3_context
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
H={'User-Agent':'Mozilla/5.0 Chrome/120 Safari/537.36'}
|
||||||
|
class W(HTTPAdapter):
|
||||||
|
def init_poolmanager(self,*a,**k):
|
||||||
|
c=create_urllib3_context();c.set_ciphers('DEFAULT@SECLEVEL=0');c.options|=0x4
|
||||||
|
c.check_hostname=False;c.verify_mode=ssl.CERT_NONE;k['ssl_context']=c
|
||||||
|
return super().init_poolmanager(*a,**k)
|
||||||
|
S=requests.Session();S.headers.update(H)
|
||||||
|
if '--weak-ssl' in sys.argv: S.mount('https://',W())
|
||||||
|
|
||||||
|
def norm(u):
|
||||||
|
s=urlsplit(u);return urlsplit(urljoin('http://x/',s.path)).path.rstrip('/').lower()
|
||||||
|
def absu(h,base):
|
||||||
|
h=(h or '').strip()
|
||||||
|
if not h or h.startswith(('javascript:','#')): return ''
|
||||||
|
if h.startswith(('http://','https://')): return h
|
||||||
|
return urljoin(base.rstrip('/')+'/',h.lstrip('/'))
|
||||||
|
|
||||||
|
def main():
|
||||||
|
xlsx,base,domain=sys.argv[1],sys.argv[2],sys.argv[3]
|
||||||
|
nums=[a for a in sys.argv[4:] if a.isdigit()]
|
||||||
|
nsamp=int(nums[0]) if nums else 50
|
||||||
|
wb=openpyxl.load_workbook(xlsx,read_only=True);ws=wb.active
|
||||||
|
urls=[];existing=set()
|
||||||
|
for row in ws.iter_rows(min_row=3,min_col=11,max_col=11,values_only=True):
|
||||||
|
u=row[0]
|
||||||
|
if isinstance(u,str) and u.startswith('http'):
|
||||||
|
existing.add(norm(u))
|
||||||
|
if domain in u: urls.append(u)
|
||||||
|
wb.close()
|
||||||
|
step=max(1,len(urls)//nsamp);sample=urls[::step][:nsamp]
|
||||||
|
print(f'표본 {len(sample)}/{len(urls)} ({domain})')
|
||||||
|
cls_cnt=Counter();examples={}
|
||||||
|
def work(u):
|
||||||
|
try:
|
||||||
|
r=S.get(u,timeout=15,verify=False);return u,r.content
|
||||||
|
except Exception:return u,None
|
||||||
|
with ThreadPoolExecutor(max_workers=8) as ex:
|
||||||
|
for f in as_completed([ex.submit(work,u) for u in sample]):
|
||||||
|
u,html=f.result()
|
||||||
|
if not html:continue
|
||||||
|
cur=norm(u);soup=BeautifulSoup(html,'html.parser')
|
||||||
|
for ul in soup.find_all('ul'):
|
||||||
|
links=[]
|
||||||
|
ok=True
|
||||||
|
for a in ul.find_all('a'):
|
||||||
|
t=a.get_text(strip=True)
|
||||||
|
au=absu(a.get('href'),base)
|
||||||
|
if not t:continue
|
||||||
|
if not au: ok=False;break
|
||||||
|
links.append(au)
|
||||||
|
if not ok or len(links)<2:continue
|
||||||
|
norms=[norm(x) for x in links]
|
||||||
|
if cur not in norms:continue # 조건3: 자기 탭그룹
|
||||||
|
if sum(1 for n in norms if n not in existing)<1:continue # 조건4: 신규1+
|
||||||
|
cls=' '.join(ul.get('class') or []) or '(no-class)'
|
||||||
|
cls_cnt[cls]+=1
|
||||||
|
examples.setdefault(cls,(u,[ (a.get_text(strip=True), absu(a.get('href'),base)) for a in ul.find_all('a') if a.get_text(strip=True)][:6]))
|
||||||
|
print('\n=== 본문탭(자기URL포함+신규有) UL class 빈도 ===')
|
||||||
|
if not cls_cnt: print(' (없음 — 본문탭 진짜 없음 가능성)')
|
||||||
|
for cls,c in cls_cnt.most_common(15):
|
||||||
|
print(f' {c:3d}회 [{cls}]')
|
||||||
|
eu,el=examples[cls];print(f' 예: {eu}')
|
||||||
|
for t,h in el: print(f' - {t} -> {h}')
|
||||||
|
|
||||||
|
if __name__=='__main__':main()
|
||||||
91
_스크립트/_tab_discover.py
Normal file
91
_스크립트/_tab_discover.py
Normal file
@ -0,0 +1,91 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""미지 탭 클래스 발견: 표본 페이지 본문에서 '같은 본문에 2+ 실제링크를 가진 UL'의
|
||||||
|
class 를 빈도순 집계한다. 사이트 전역 nav(거의 모든 페이지에 동일 링크집합) 제외 목적.
|
||||||
|
|
||||||
|
사용: python -X utf8 _tab_discover.py <엑셀> <도메인> [표본수=40]
|
||||||
|
"""
|
||||||
|
import sys, warnings, ssl
|
||||||
|
from collections import Counter
|
||||||
|
from urllib.parse import urlsplit
|
||||||
|
from concurrent.futures import ThreadPoolExecutor, as_completed
|
||||||
|
|
||||||
|
import openpyxl, requests
|
||||||
|
from requests.adapters import HTTPAdapter
|
||||||
|
from urllib3.util.ssl_ import create_urllib3_context
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
H = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) Chrome/120 Safari/537.36'}
|
||||||
|
|
||||||
|
|
||||||
|
class WeakSSLAdapter(HTTPAdapter):
|
||||||
|
def init_poolmanager(self, *a, **k):
|
||||||
|
ctx = create_urllib3_context(); ctx.set_ciphers('DEFAULT@SECLEVEL=0')
|
||||||
|
ctx.options |= 0x4; ctx.check_hostname = False; ctx.verify_mode = ssl.CERT_NONE
|
||||||
|
k['ssl_context'] = ctx; return super().init_poolmanager(*a, **k)
|
||||||
|
|
||||||
|
|
||||||
|
S = requests.Session(); S.headers.update(H)
|
||||||
|
if '--weak-ssl' in sys.argv:
|
||||||
|
S.mount('https://', WeakSSLAdapter())
|
||||||
|
|
||||||
|
|
||||||
|
def norm(u):
|
||||||
|
s = urlsplit(u); return (s.path.rstrip('/')).lower()
|
||||||
|
|
||||||
|
|
||||||
|
def fetch(r, u):
|
||||||
|
try:
|
||||||
|
x = S.get(u, timeout=15, verify=False)
|
||||||
|
return r, u, x.content
|
||||||
|
except Exception:
|
||||||
|
return r, u, None
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
xlsx, domain = sys.argv[1], sys.argv[2]
|
||||||
|
nsamp = int([a for a in sys.argv[3:] if a.isdigit()][0]) if any(a.isdigit() for a in sys.argv[3:]) else 40
|
||||||
|
wb = openpyxl.load_workbook(xlsx, read_only=True); ws = wb.active
|
||||||
|
urls = []
|
||||||
|
for row in ws.iter_rows(min_row=3, min_col=11, max_col=11, values_only=True):
|
||||||
|
u = row[0]
|
||||||
|
if isinstance(u, str) and domain in u:
|
||||||
|
urls.append(u)
|
||||||
|
wb.close()
|
||||||
|
step = max(1, len(urls) // nsamp)
|
||||||
|
sample = urls[::step][:nsamp]
|
||||||
|
print(f'표본 {len(sample)}/{len(urls)} ({domain})')
|
||||||
|
|
||||||
|
# class별 등장 페이지 수 / 링크집합 다양성
|
||||||
|
cls_pages = Counter()
|
||||||
|
cls_linksets = {}
|
||||||
|
with ThreadPoolExecutor(max_workers=8) as ex:
|
||||||
|
futs = [ex.submit(fetch, i, u) for i, u in enumerate(sample)]
|
||||||
|
for f in as_completed(futs):
|
||||||
|
r, u, html = f.result()
|
||||||
|
if not html:
|
||||||
|
continue
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
for ul in soup.find_all(['ul', 'ol']):
|
||||||
|
links = []
|
||||||
|
for a in ul.find_all('a'):
|
||||||
|
t = a.get_text(strip=True)
|
||||||
|
h = (a.get('href') or '').strip()
|
||||||
|
if t and h and not h.startswith(('#', 'javascript:')):
|
||||||
|
links.append(norm(h))
|
||||||
|
if len(links) < 2:
|
||||||
|
continue
|
||||||
|
cls = ' '.join(ul.get('class') or []).strip() or '(no-class)'
|
||||||
|
cls_pages[cls] += 1
|
||||||
|
cls_linksets.setdefault(cls, set()).add(tuple(links))
|
||||||
|
|
||||||
|
print('\n=== UL/OL class별 (2+링크) — 등장페이지수 / 서로다른 링크집합수 ===')
|
||||||
|
print('(전역 nav = 등장多+집합1 / 본문탭 = 집합 다양) \n')
|
||||||
|
for cls, cnt in cls_pages.most_common(40):
|
||||||
|
nset = len(cls_linksets[cls])
|
||||||
|
flag = '★탭후보' if nset >= 3 and cnt >= 3 else ''
|
||||||
|
print(f' {cnt:3d}p / {nset:3d}집합 [{cls}] {flag}')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
107
_스크립트/_tab_dry.log
Normal file
107
_스크립트/_tab_dry.log
Normal file
@ -0,0 +1,107 @@
|
|||||||
|
대상 기관: 39개 (모드=dry)
|
||||||
|
|
||||||
|
━━━ 충청남도 금산군 (행~526, base=https://www.geumsan.go.kr, domain=geumsan.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\3.금산군\금산군.xlsx https://www.geumsan.go.kr geumsan.go.kr
|
||||||
|
→ 확장그룹 3개 / 신규행 9개
|
||||||
|
|
||||||
|
━━━ 충청남도 논산시 (행~738, base=https://nonsan.go.kr, domain=nonsan.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\4.논산시\논산시.xlsx https://nonsan.go.kr nonsan.go.kr
|
||||||
|
→ 확장그룹 7개 / 신규행 22개
|
||||||
|
|
||||||
|
━━━ 충청남도 당진시 (행~315, base=https://www.dangjin.go.kr, domain=dangjin.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\5.당진시\당진시.xlsx https://www.dangjin.go.kr dangjin.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
|
||||||
|
━━━ 충청남도 보령시 (행~573, base=https://www.brcn.go.kr, domain=brcn.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\6.보령시\보령시.xlsx https://www.brcn.go.kr brcn.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
|
||||||
|
━━━ 충청남도 부여군 (행~315, base=https://www.buyeo.go.kr, domain=buyeo.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\7.부여군\부여군.xlsx https://www.buyeo.go.kr buyeo.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
|
||||||
|
━━━ 충청남도 서산시 (행~349, base=https://www.seosan.go.kr, domain=seosan.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\8.서산시\서산시.xlsx https://www.seosan.go.kr seosan.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
|
||||||
|
━━━ 충청남도 서천군 (행~330, base=https://www.seocheon.go.kr, domain=seocheon.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\9.서천군\서천군.xlsx https://www.seocheon.go.kr seocheon.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
|
||||||
|
━━━ 충청남도 아산시 (행~315, base=https://www.asan.go.kr, domain=asan.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\10.아산시\아산시.xlsx https://www.asan.go.kr asan.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
|
||||||
|
━━━ 충청남도 예산군 (행~387, base=https://www.yesan.go.kr, domain=yesan.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\11.예산군\예산군.xlsx https://www.yesan.go.kr yesan.go.kr
|
||||||
|
→ 확장그룹 91개 / 신규행 233개
|
||||||
|
|
||||||
|
━━━ 충청남도 천안시 (행~371, base=https://www.cheonan.go.kr, domain=cheonan.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\12.천안시\천안시.xlsx https://www.cheonan.go.kr cheonan.go.kr
|
||||||
|
→ 확장그룹 107개 / 신규행 297개
|
||||||
|
|
||||||
|
━━━ 충청남도 청양군 (행~315, base=http://www.cheongyang.go.kr, domain=cheongyang.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\13.청양군\청양군.xlsx http://www.cheongyang.go.kr cheongyang.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
|
||||||
|
━━━ 충청남도 태안군 (행~315, base=https://www.taean.go.kr, domain=taean.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\14.태안군\태안군.xlsx https://www.taean.go.kr taean.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
|
||||||
|
━━━ 충청남도 홍성군 (행~315, base=https://www.hongseong.go.kr, domain=hongseong.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\15.홍성군\홍성군.xlsx https://www.hongseong.go.kr hongseong.go.kr
|
||||||
|
→ 확장그룹 67개 / 신규행 215개
|
||||||
|
|
||||||
|
━━━ 충청북도 괴산군 (행~315, base=https://www.goesan.go.kr, domain=goesan.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\1.괴산군\괴산군.xlsx https://www.goesan.go.kr goesan.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
|
||||||
|
━━━ 충청북도 단양군 (행~486, base=https://www.danyang.go.kr, domain=danyang.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\2.단양군\단양군.xlsx https://www.danyang.go.kr danyang.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
|
||||||
|
━━━ 충청북도 영동군 [weak-ssl] (행~674, base=https://www.yd21.go.kr, domain=yd21.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\4.영동군\영동군.xlsx https://www.yd21.go.kr yd21.go.kr --weak-ssl
|
||||||
|
→ 확장그룹 20개 / 신규행 143개
|
||||||
|
|
||||||
|
━━━ 충청북도 옥천군 (행~564, base=https://www.oc.go.kr, domain=oc.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\5.옥천군\옥천군.xlsx https://www.oc.go.kr oc.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
|
||||||
|
━━━ 충청북도 음성군 (행~652, base=https://www.eumseong.go.kr, domain=eumseong.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\6.음성군\음성군.xlsx https://www.eumseong.go.kr eumseong.go.kr
|
||||||
|
→ 확장그룹 9개 / 신규행 18개
|
||||||
|
|
||||||
|
━━━ 충청북도 제천시 (행~644, base=https://www.jecheon.go.kr, domain=jecheon.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\7.제천시\제천시.xlsx https://www.jecheon.go.kr jecheon.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
|
||||||
|
━━━ 충청북도 증평군 (행~315, base=https://www.jp.go.kr, domain=jp.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\8.증평군\증평군.xlsx https://www.jp.go.kr jp.go.kr
|
||||||
|
→ 확장그룹 42개 / 신규행 175개
|
||||||
|
|
||||||
|
━━━ 충청북도 진천군 (행~334, base=https://www.jincheon.go.kr, domain=jincheon.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\9.진천군\진천군.xlsx https://www.jincheon.go.kr jincheon.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
|
||||||
|
━━━ 충청북도 청주시 (행~342, base=https://www.cheongju.go.kr, domain=cheongju.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\10.청주시\청주시.xlsx https://www.cheongju.go.kr cheongju.go.kr
|
||||||
|
→ 확장그룹 5개 / 신규행 97개
|
||||||
|
|
||||||
|
━━━ 충청북도 충주시 (행~403, base=https://www.chungju.go.kr, domain=chungju.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\11.충주시\충주시.xlsx https://www.chungju.go.kr chungju.go.kr
|
||||||
|
→ 확장그룹 3개 / 신규행 33개
|
||||||
|
|
||||||
|
━━━ 전북특별자치도 고창군 (행~381, base=https://www.gochang.go.kr, domain=gochang.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\1.고창군\고창군.xlsx https://www.gochang.go.kr gochang.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
|
||||||
|
━━━ 전북특별자치도 군산시 (행~639, base=https://www.gunsan.go.kr, domain=gunsan.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\2.군산시\군산시.xlsx https://www.gunsan.go.kr gunsan.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
|
||||||
|
━━━ 전북특별자치도 김제시 (행~454, base=https://www.gimje.go.kr, domain=gimje.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\3.김제시\김제시.xlsx https://www.gimje.go.kr gimje.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
|
||||||
|
━━━ 전북특별자치도 남원시 (행~333, base=https://www.namwon.go.kr, domain=namwon.go.kr)
|
||||||
433
_스크립트/_tab_expand.py
Normal file
433
_스크립트/_tab_expand.py
Normal file
@ -0,0 +1,433 @@
|
|||||||
|
"""본문 탭(서브내비) → 카테고리 하위 확장.
|
||||||
|
|
||||||
|
규칙(공주시 기준, 일반화):
|
||||||
|
본문에서 '탭 UL'을 찾아, 아래 조건을 모두 만족하는 탭 그룹만 하위 카테고리로 확장한다.
|
||||||
|
1) UL(또는 직계 div)의 class 에 탭 패턴(tab-ul 등)이 있다.
|
||||||
|
2) 탭이 2개 이상이고, 모든 탭 href 가 '실제 페이지 링크'다.
|
||||||
|
(#anchor 인페이지 탭, javascript:, 빈 href, 쿼리스트링만 다른 필터 탭은 제외)
|
||||||
|
3) 현재 페이지 URL 이 탭 URL 집합에 포함된다(= 자기 자신의 탭 그룹).
|
||||||
|
4) 탭 중 '엑셀에 아직 없는 URL'이 1개 이상 있다(이미 사이트맵에 다 있으면 상위 nav → 제외).
|
||||||
|
|
||||||
|
확장 방식(=사용자 지시: 탭들은 한 단계 아래 컬럼으로):
|
||||||
|
- 원래 행의 leaf 컬럼(E~J 중 가장 깊은 값)을 부모(카테고리)로 두고,
|
||||||
|
탭들을 leaf+1 컬럼에 탭 순서대로 채운다.
|
||||||
|
- 현재 페이지와 같은 URL의 탭 = 원래 행을 재사용(L~T 기존 데이터 보존, G만 탭명 기입).
|
||||||
|
- 나머지 탭 = 신규 행(L~T 비움, Phase 2~4 별도 수행 대상).
|
||||||
|
|
||||||
|
사용법:
|
||||||
|
python -X utf8 _tab_expand.py <엑셀경로> <BASE_URL> <도메인키워드> [--write]
|
||||||
|
예) python -X utf8 _tab_expand.py 충청남도/2.공주시/충청남도_공주시.xlsx https://www.gongju.go.kr gongju.go.kr
|
||||||
|
(--write 없으면 계획만 출력 / 있으면 백업 후 실제 기입)
|
||||||
|
"""
|
||||||
|
import re
|
||||||
|
import sys
|
||||||
|
import shutil
|
||||||
|
import warnings
|
||||||
|
from copy import copy
|
||||||
|
from concurrent.futures import ThreadPoolExecutor, as_completed
|
||||||
|
from urllib.parse import urljoin, urlsplit, urlunsplit
|
||||||
|
|
||||||
|
import ssl
|
||||||
|
import openpyxl
|
||||||
|
import requests
|
||||||
|
from requests.adapters import HTTPAdapter
|
||||||
|
from urllib3.util.ssl_ import create_urllib3_context
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
from openpyxl.styles import Alignment, Font
|
||||||
|
from openpyxl.utils import get_column_letter
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
|
||||||
|
|
||||||
|
class WeakSSLAdapter(HTTPAdapter):
|
||||||
|
def init_poolmanager(self, *args, **kwargs):
|
||||||
|
ctx = create_urllib3_context()
|
||||||
|
ctx.set_ciphers('DEFAULT@SECLEVEL=0')
|
||||||
|
ctx.options |= 0x4
|
||||||
|
ctx.check_hostname = False
|
||||||
|
ctx.verify_mode = ssl.CERT_NONE
|
||||||
|
kwargs['ssl_context'] = ctx
|
||||||
|
return super().init_poolmanager(*args, **kwargs)
|
||||||
|
|
||||||
|
|
||||||
|
SESSION = requests.Session()
|
||||||
|
SESSION.headers.update(H)
|
||||||
|
|
||||||
|
TAB_CLASS_PATS = ['tab-ul', 'tab_ul', 'tabmenu', 'tab-menu', 'tab_menu',
|
||||||
|
'tablist', 'tab-list', 'tab_list', 'subtab', 'sub-tab', 'sub_tab']
|
||||||
|
# 토큰 단위 탭 클래스 정규식: 'basic_tab', 'tab_wrap', 'tab-ul', 단독 'tab' 등 매칭.
|
||||||
|
# (UL 자신뿐 아니라 직계 부모 div 클래스도 검사 → div.basic_tab > ul 구조 대응)
|
||||||
|
TAB_CLASS_RE = re.compile(
|
||||||
|
r'(?:^|[-_ ])tab(?:[-_ ]|$)|tabmenu|tablist|tab[-_]?(?:ul|wrap|list|menu)|basic[-_]tab')
|
||||||
|
|
||||||
|
MAXCOL = 27 # AA
|
||||||
|
CAT_COLS = [5, 6, 7, 8, 9, 10] # E F G H I J
|
||||||
|
HEADER_MERGES = {'B1:R1', 'S1:W1', 'Y1:AA1'}
|
||||||
|
|
||||||
|
|
||||||
|
def norm_url(u):
|
||||||
|
"""비교용 키: 프래그먼트 제거 + 끝 슬래시 정리 + 쿼리 정렬 보존.
|
||||||
|
(쿼리만 다른 게시판 분류 탭 ?code=A vs ?code=B 를 서로 다른 URL로 구분 →
|
||||||
|
is_cur·existing 중복판정 오류 방지)"""
|
||||||
|
if not u:
|
||||||
|
return ''
|
||||||
|
s = urlsplit(u)
|
||||||
|
path = s.path.rstrip('/')
|
||||||
|
q = '&'.join(sorted(s.query.split('&'))) if s.query else ''
|
||||||
|
return urlunsplit((s.scheme, s.netloc, path, q, '')).lower()
|
||||||
|
|
||||||
|
|
||||||
|
def path_key(u):
|
||||||
|
"""쿼리 제거한 경로 키 (필터 탭 판별용)."""
|
||||||
|
if not u:
|
||||||
|
return ''
|
||||||
|
s = urlsplit(u)
|
||||||
|
return urlunsplit((s.scheme, s.netloc, s.path.rstrip('/'), '', '')).lower()
|
||||||
|
|
||||||
|
|
||||||
|
def abs_url(href, base):
|
||||||
|
if not href:
|
||||||
|
return ''
|
||||||
|
href = href.strip()
|
||||||
|
if href.startswith(('javascript:', '#')):
|
||||||
|
return ''
|
||||||
|
if href.startswith(('http://', 'https://')):
|
||||||
|
return href
|
||||||
|
return urljoin(base.rstrip('/') + '/', href.lstrip('/'))
|
||||||
|
|
||||||
|
|
||||||
|
def has_tab_class(el):
|
||||||
|
if el is None or not getattr(el, 'get', None):
|
||||||
|
return False, ''
|
||||||
|
cls = ' '.join(el.get('class') or []).lower()
|
||||||
|
ok = any(p in cls for p in TAB_CLASS_PATS) or bool(TAB_CLASS_RE.search(cls))
|
||||||
|
return ok, cls
|
||||||
|
|
||||||
|
|
||||||
|
def _has_on_li(ul):
|
||||||
|
"""ul 직계 li(또는 그 a)에 활성 탭 마커(on/active/current/selected)가 있나."""
|
||||||
|
marks = {'on', 'active', 'current', 'selected', 'sel'}
|
||||||
|
for li in ul.find_all('li', recursive=False):
|
||||||
|
if marks & set(c.lower() for c in (li.get('class') or [])):
|
||||||
|
return True
|
||||||
|
a = li.find('a')
|
||||||
|
if a and (marks & set(c.lower() for c in (a.get('class') or []))):
|
||||||
|
return True
|
||||||
|
return False
|
||||||
|
|
||||||
|
|
||||||
|
def find_tab_groups(soup, base):
|
||||||
|
"""조건 1·2 만족하는 탭 그룹 후보 반환: [(cls, [(label, abs_url),...], has_on), ...]"""
|
||||||
|
out = []
|
||||||
|
seen = set()
|
||||||
|
for ul in soup.find_all('ul'):
|
||||||
|
ok, cls = has_tab_class(ul)
|
||||||
|
if not ok:
|
||||||
|
# UL 무클래스라도 직계 부모 div 에 탭 클래스가 있으면 인정 (div.basic_tab > ul)
|
||||||
|
pok, pcls = has_tab_class(ul.parent)
|
||||||
|
if not pok:
|
||||||
|
continue # 쿼리필터·무클래스 ul 배제
|
||||||
|
cls = pcls
|
||||||
|
links = []
|
||||||
|
real = True
|
||||||
|
for a in ul.find_all('a'):
|
||||||
|
t = a.get_text(strip=True).replace('\xa0', '').strip()
|
||||||
|
raw = (a.get('href') or '').strip()
|
||||||
|
au = abs_url(raw, base)
|
||||||
|
if not t:
|
||||||
|
continue
|
||||||
|
if not au: # #anchor / javascript / 빈 href → 인페이지 탭
|
||||||
|
real = False
|
||||||
|
break
|
||||||
|
links.append((t, au))
|
||||||
|
if not real or len(links) < 2:
|
||||||
|
continue
|
||||||
|
key = tuple(nu for _, (nu) in [(t, u) for t, u in links])
|
||||||
|
if key in seen:
|
||||||
|
continue
|
||||||
|
seen.add(key)
|
||||||
|
out.append((cls, links, _has_on_li(ul)))
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def fetch(r, url):
|
||||||
|
try:
|
||||||
|
resp = SESSION.get(url, timeout=15, verify=False)
|
||||||
|
# 바이트로 받아 BeautifulSoup가 meta charset으로 인코딩 판별 (apparent_encoding 오판 방지)
|
||||||
|
return r, url, resp.content, None
|
||||||
|
except Exception as e:
|
||||||
|
return r, url, None, str(e)[:60]
|
||||||
|
|
||||||
|
|
||||||
|
def load_flat(ws):
|
||||||
|
"""병합값을 채워넣어 행별 완전 데이터로 평탄화. 반환: rows(list of dict), 스타일/하이퍼링크 캡처."""
|
||||||
|
# 1) 병합값 채우기 (헤더 제외, 데이터 컬럼)
|
||||||
|
for mr in list(ws.merged_cells.ranges):
|
||||||
|
s = str(mr)
|
||||||
|
if s in HEADER_MERGES:
|
||||||
|
continue
|
||||||
|
top = ws.cell(mr.min_row, mr.min_col).value
|
||||||
|
ws.unmerge_cells(s)
|
||||||
|
for rr in range(mr.min_row, mr.max_row + 1):
|
||||||
|
for cc in range(mr.min_col, mr.max_col + 1):
|
||||||
|
if ws.cell(rr, cc).value in (None, ''):
|
||||||
|
ws.cell(rr, cc).value = top
|
||||||
|
rows = []
|
||||||
|
for r in range(3, ws.max_row + 1):
|
||||||
|
# 빈 행 스킵
|
||||||
|
if all(ws.cell(r, c).value in (None, '') for c in range(2, MAXCOL + 1)):
|
||||||
|
continue
|
||||||
|
vals = {c: ws.cell(r, c).value for c in range(1, MAXCOL + 1)}
|
||||||
|
styles = {}
|
||||||
|
for c in range(1, MAXCOL + 1):
|
||||||
|
sc = ws.cell(r, c)
|
||||||
|
styles[c] = (copy(sc.font), copy(sc.fill), copy(sc.border),
|
||||||
|
copy(sc.alignment), sc.number_format, copy(sc.protection))
|
||||||
|
hl = ws.cell(r, 11).hyperlink
|
||||||
|
rows.append({'src': r, 'vals': vals, 'styles': styles,
|
||||||
|
'hyperlink': hl.target if hl else None})
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def leaf_col(vals):
|
||||||
|
"""E~J 중 값이 있는 가장 깊은 컬럼 인덱스."""
|
||||||
|
deep = 5
|
||||||
|
for c in CAT_COLS:
|
||||||
|
if vals.get(c) not in (None, ''):
|
||||||
|
deep = c
|
||||||
|
return deep
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
xlsx, base, domain = sys.argv[1], sys.argv[2], sys.argv[3]
|
||||||
|
write = '--write' in sys.argv
|
||||||
|
# menuCd 벤더(고창·임실·정읍·진안 등 index.{name}?menuCd=…): 모든 페이지가 같은 경로,
|
||||||
|
# 쿼리(menuCd)만 다름 = 다른 페이지. → 쿼리-경로 가드(조건2) 스킵 + cur∈탭(조건3) 대신
|
||||||
|
# div.basic_tab 등에 활성탭(li.on) 마커가 있는 '자기 sub-nav'만 인정(랜딩 URL이 탭과 달라도).
|
||||||
|
menucd = '--menucd' in sys.argv
|
||||||
|
if '--weak-ssl' in sys.argv:
|
||||||
|
SESSION.mount('https://', WeakSSLAdapter())
|
||||||
|
|
||||||
|
wb = openpyxl.load_workbook(xlsx)
|
||||||
|
ws = wb.active
|
||||||
|
rows = load_flat(ws)
|
||||||
|
existing = set()
|
||||||
|
for row in rows:
|
||||||
|
u = row['vals'].get(11)
|
||||||
|
if isinstance(u, str) and u.startswith('http'):
|
||||||
|
existing.add(norm_url(u))
|
||||||
|
|
||||||
|
# 같은 도메인 행 fetch
|
||||||
|
targets = [(row['src'], row['vals'].get(11)) for row in rows
|
||||||
|
if isinstance(row['vals'].get(11), str) and domain in row['vals'].get(11)]
|
||||||
|
html_by_src = {}
|
||||||
|
with ThreadPoolExecutor(max_workers=8) as ex:
|
||||||
|
futs = [ex.submit(fetch, r, u) for r, u in targets]
|
||||||
|
for fut in as_completed(futs):
|
||||||
|
r, url, html, err = fut.result()
|
||||||
|
if html:
|
||||||
|
html_by_src[r] = (url, html)
|
||||||
|
|
||||||
|
# 이미 사이트맵에 자식행이 있는 '랜딩행' 집합 (다음 평탄행이 같은 상위카테고리 + 더 깊은 leaf).
|
||||||
|
# menuCd 벤더 랜딩은 첫 자식의 탭을 렌더하므로, 확장하면 그 자식행과 중복 → 제외.
|
||||||
|
landing_src = set()
|
||||||
|
for i in range(len(rows) - 1):
|
||||||
|
v, nv = rows[i]['vals'], rows[i + 1]['vals']
|
||||||
|
lc = leaf_col(v)
|
||||||
|
if lc < 10 and nv.get(lc + 1) not in (None, '') \
|
||||||
|
and all((v.get(c) or '') == (nv.get(c) or '') for c in CAT_COLS if c <= lc):
|
||||||
|
landing_src.add(rows[i]['src'])
|
||||||
|
|
||||||
|
# 행별 확장 계획
|
||||||
|
plan = {} # src_row -> ordered [(label, url, is_existing)]
|
||||||
|
planned_new = set() # 이미 어느 부모행이 추가한 신규 탭 URL(전역 중복 방지)
|
||||||
|
for row in rows:
|
||||||
|
src = row['src']
|
||||||
|
if src not in html_by_src:
|
||||||
|
continue
|
||||||
|
if menucd and src in landing_src:
|
||||||
|
continue # 자식 보유 랜딩 → 확장 금지(첫 자식 탭 중복 방지)
|
||||||
|
url, html = html_by_src[src]
|
||||||
|
cur = norm_url(url)
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
groups = find_tab_groups(soup, base)
|
||||||
|
chosen = None
|
||||||
|
for cls, links, has_on in groups:
|
||||||
|
if menucd:
|
||||||
|
# menuCd 벤더: 현재 페이지와 '같은 경로'(menuCd만 다른 형제)인 탭만 인정.
|
||||||
|
# → /gochang/toc/GC…(향토문화대전 백과 목록) 등 다른경로 링크 배제.
|
||||||
|
# 활성탭(li.on) 마커가 있는 자기 sub-nav만(랜딩 URL이 탭에 없어도 OK).
|
||||||
|
if not has_on:
|
||||||
|
continue
|
||||||
|
links = [(t, u) for t, u in links if path_key(u) == path_key(cur)]
|
||||||
|
if len(links) < 2:
|
||||||
|
continue
|
||||||
|
tab_norms = [norm_url(u) for _, u in links]
|
||||||
|
else:
|
||||||
|
tab_norms = [norm_url(u) for _, u in links]
|
||||||
|
if len({path_key(u) for _, u in links}) == 1:
|
||||||
|
continue # 조건2: 같은 경로(쿼리만 다른 필터 탭) → 제외
|
||||||
|
if cur not in tab_norms:
|
||||||
|
continue # 조건3: 자기 탭그룹만
|
||||||
|
new_cnt = sum(1 for n in tab_norms if n not in existing)
|
||||||
|
if new_cnt < 1:
|
||||||
|
continue # 조건4: 신규 0 → 상위nav, 제외
|
||||||
|
chosen = links
|
||||||
|
break
|
||||||
|
if chosen:
|
||||||
|
ch = []
|
||||||
|
for label, u in chosen:
|
||||||
|
is_cur = norm_url(u) == cur
|
||||||
|
nu = norm_url(u)
|
||||||
|
if not is_cur and nu in existing:
|
||||||
|
continue # 이미 사이트맵에 별도 행으로 존재 → 중복행 방지(skip)
|
||||||
|
if not is_cur and nu in planned_new:
|
||||||
|
continue # 다른 부모행이 이미 추가한 신규 URL → 전역 중복 방지
|
||||||
|
if not is_cur:
|
||||||
|
planned_new.add(nu)
|
||||||
|
is_ext = domain not in u # 같은 도메인 아님 = 외부 사이트 탭
|
||||||
|
ch.append((label, u, is_cur, is_ext))
|
||||||
|
# 신규(비-cur) 탭이 1개 이상일 때만 확장. 아니면 원본행 그대로 둠(삭제 방지).
|
||||||
|
if any(not c[2] for c in ch):
|
||||||
|
plan[src] = ch
|
||||||
|
|
||||||
|
# 계획 출력
|
||||||
|
print(f'=== 탭 확장 계획: {len(plan)}개 그룹 ===\n')
|
||||||
|
total_new = 0
|
||||||
|
for src in sorted(plan):
|
||||||
|
row = next(r for r in rows if r['src'] == src)
|
||||||
|
v = row['vals']
|
||||||
|
lc = leaf_col(v)
|
||||||
|
childc = lc + 1
|
||||||
|
ev = v.get(5) or ''
|
||||||
|
fv = v.get(6) or ''
|
||||||
|
print(f'[행{src}] E={ev} F={fv} '
|
||||||
|
f'(leaf={get_column_letter(lc)} → 자식={get_column_letter(childc)})')
|
||||||
|
for label, u, is_cur, is_ext in plan[src]:
|
||||||
|
if is_cur:
|
||||||
|
tag = '재사용'
|
||||||
|
elif is_ext:
|
||||||
|
tag = '신규+(외부=사이트)'
|
||||||
|
total_new += 1
|
||||||
|
else:
|
||||||
|
tag = '신규+'
|
||||||
|
total_new += 1
|
||||||
|
print(f' [{tag}] {label} -> {u}')
|
||||||
|
print()
|
||||||
|
print(f'신규 추가 행: {total_new}개 / 기존 {len(rows)}행 → 최종 {len(rows)+total_new}행')
|
||||||
|
|
||||||
|
if not write:
|
||||||
|
print('\n(계획만 출력. 실제 기입하려면 --write)')
|
||||||
|
return
|
||||||
|
|
||||||
|
# ===== 실제 기입 =====
|
||||||
|
backup = xlsx.replace('.xlsx', '_backup_tab전.xlsx')
|
||||||
|
shutil.copy(xlsx, backup)
|
||||||
|
print(f'\n백업: {backup}')
|
||||||
|
|
||||||
|
# 새 평탄 행 목록 구성
|
||||||
|
out_rows = [] # 각 원소 = {'vals':..., 'style_src':row_dict, 'url':..}
|
||||||
|
for row in rows:
|
||||||
|
src = row['src']
|
||||||
|
if src in plan:
|
||||||
|
v = row['vals']
|
||||||
|
lc = leaf_col(v)
|
||||||
|
childc = lc + 1
|
||||||
|
for label, u, is_cur, is_ext in plan[src]:
|
||||||
|
nv = dict(v)
|
||||||
|
# leaf 보다 깊은 컬럼 초기화 후 자식 컬럼에 탭명
|
||||||
|
for c in CAT_COLS:
|
||||||
|
if c > lc:
|
||||||
|
nv[c] = None
|
||||||
|
nv[childc] = label
|
||||||
|
if is_cur:
|
||||||
|
nv[11] = v.get(11) # 기존 URL/데이터 유지
|
||||||
|
out_rows.append({'vals': nv, 'style': row, 'url': nv.get(11)})
|
||||||
|
else:
|
||||||
|
for c in range(12, MAXCOL + 1):
|
||||||
|
nv[c] = None # L~T 비움 (신규)
|
||||||
|
nv[11] = u
|
||||||
|
nv[19] = None
|
||||||
|
if is_ext: # 외부 사이트 탭 → 즉시 L=사이트/M=1(Phase234 대상 아님)
|
||||||
|
nv[12] = '사이트'
|
||||||
|
nv[13] = 1
|
||||||
|
out_rows.append({'vals': nv, 'style': row, 'url': u})
|
||||||
|
else:
|
||||||
|
out_rows.append({'vals': dict(row['vals']), 'style': row, 'url': row['vals'].get(11)})
|
||||||
|
|
||||||
|
# 데이터 영역 클리어
|
||||||
|
for r in range(3, ws.max_row + 1):
|
||||||
|
for c in range(1, MAXCOL + 1):
|
||||||
|
ws.cell(r, c).value = None
|
||||||
|
ws.cell(r, c).hyperlink = None
|
||||||
|
|
||||||
|
START = 3
|
||||||
|
for i, orow in enumerate(out_rows):
|
||||||
|
r = START + i
|
||||||
|
sty = orow['style']['styles']
|
||||||
|
for c in range(1, MAXCOL + 1):
|
||||||
|
cell = ws.cell(r, c)
|
||||||
|
f, fl, bd, al, nf, pr = sty[c]
|
||||||
|
cell.font = copy(f); cell.fill = copy(fl); cell.border = copy(bd)
|
||||||
|
cell.alignment = copy(al); cell.number_format = nf; cell.protection = copy(pr)
|
||||||
|
v = orow['vals']
|
||||||
|
for c in range(1, MAXCOL + 1):
|
||||||
|
ws.cell(r, c).value = v.get(c)
|
||||||
|
ws.cell(r, 2).value = i + 1 # B 순번 재부여
|
||||||
|
END = START + len(out_rows) - 1
|
||||||
|
|
||||||
|
# D/E/F 재병합 (G는 leaf라 병합 안 함)
|
||||||
|
center = Alignment(horizontal='center', vertical='center', wrap_text=True)
|
||||||
|
left = Alignment(horizontal='left', vertical='center', wrap_text=False)
|
||||||
|
|
||||||
|
def merge_runs(col_letter, col_idx, group_cols=()):
|
||||||
|
cur_val = ws.cell(START, col_idx).value
|
||||||
|
cur_grp = tuple(ws.cell(START, g).value for g in group_cols)
|
||||||
|
run_start = START
|
||||||
|
runs = []
|
||||||
|
for r in range(START + 1, END + 1):
|
||||||
|
val = ws.cell(r, col_idx).value
|
||||||
|
grp = tuple(ws.cell(r, gg).value for gg in group_cols)
|
||||||
|
if val == cur_val and grp == cur_grp:
|
||||||
|
continue
|
||||||
|
if cur_val not in (None, '') and r - 1 > run_start:
|
||||||
|
runs.append((run_start, r - 1))
|
||||||
|
cur_val, cur_grp, run_start = val, grp, r
|
||||||
|
if cur_val not in (None, '') and END > run_start:
|
||||||
|
runs.append((run_start, END))
|
||||||
|
for s, e in runs:
|
||||||
|
ws.merge_cells(f'{col_letter}{s}:{col_letter}{e}')
|
||||||
|
ws.cell(s, col_idx).alignment = center
|
||||||
|
|
||||||
|
merge_runs('F', 6, group_cols=(4, 5))
|
||||||
|
merge_runs('E', 5, group_cols=(4,))
|
||||||
|
merge_runs('D', 4)
|
||||||
|
for r in range(START, END + 1):
|
||||||
|
for c in (4, 5, 6, 7):
|
||||||
|
if ws.cell(r, c).value is not None:
|
||||||
|
ws.cell(r, c).alignment = center
|
||||||
|
|
||||||
|
for r in range(1, END + 1):
|
||||||
|
ws.row_dimensions[r].height = 15
|
||||||
|
|
||||||
|
# K 하이퍼링크 재설정
|
||||||
|
for r in range(START, END + 1):
|
||||||
|
cell = ws.cell(r, 11)
|
||||||
|
u = cell.value
|
||||||
|
if isinstance(u, str) and u.startswith(('http://', 'https://')):
|
||||||
|
cell.hyperlink = u
|
||||||
|
old = cell.font
|
||||||
|
cell.font = Font(name=old.name or '맑은 고딕', size=old.size or 11,
|
||||||
|
bold=old.bold, italic=old.italic,
|
||||||
|
color='0000FF', underline='single')
|
||||||
|
cell.alignment = left
|
||||||
|
|
||||||
|
wb.save(xlsx)
|
||||||
|
print(f'저장 완료: {xlsx} (총 {len(out_rows)}행)')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
350
_스크립트/_tab_expand_backup_basictab전.py
Normal file
350
_스크립트/_tab_expand_backup_basictab전.py
Normal file
@ -0,0 +1,350 @@
|
|||||||
|
"""본문 탭(서브내비) → 카테고리 하위 확장.
|
||||||
|
|
||||||
|
규칙(공주시 기준, 일반화):
|
||||||
|
본문에서 '탭 UL'을 찾아, 아래 조건을 모두 만족하는 탭 그룹만 하위 카테고리로 확장한다.
|
||||||
|
1) UL(또는 직계 div)의 class 에 탭 패턴(tab-ul 등)이 있다.
|
||||||
|
2) 탭이 2개 이상이고, 모든 탭 href 가 '실제 페이지 링크'다.
|
||||||
|
(#anchor 인페이지 탭, javascript:, 빈 href, 쿼리스트링만 다른 필터 탭은 제외)
|
||||||
|
3) 현재 페이지 URL 이 탭 URL 집합에 포함된다(= 자기 자신의 탭 그룹).
|
||||||
|
4) 탭 중 '엑셀에 아직 없는 URL'이 1개 이상 있다(이미 사이트맵에 다 있으면 상위 nav → 제외).
|
||||||
|
|
||||||
|
확장 방식(=사용자 지시: 탭들은 한 단계 아래 컬럼으로):
|
||||||
|
- 원래 행의 leaf 컬럼(E~J 중 가장 깊은 값)을 부모(카테고리)로 두고,
|
||||||
|
탭들을 leaf+1 컬럼에 탭 순서대로 채운다.
|
||||||
|
- 현재 페이지와 같은 URL의 탭 = 원래 행을 재사용(L~T 기존 데이터 보존, G만 탭명 기입).
|
||||||
|
- 나머지 탭 = 신규 행(L~T 비움, Phase 2~4 별도 수행 대상).
|
||||||
|
|
||||||
|
사용법:
|
||||||
|
python -X utf8 _tab_expand.py <엑셀경로> <BASE_URL> <도메인키워드> [--write]
|
||||||
|
예) python -X utf8 _tab_expand.py 충청남도/2.공주시/충청남도_공주시.xlsx https://www.gongju.go.kr gongju.go.kr
|
||||||
|
(--write 없으면 계획만 출력 / 있으면 백업 후 실제 기입)
|
||||||
|
"""
|
||||||
|
import re
|
||||||
|
import sys
|
||||||
|
import shutil
|
||||||
|
import warnings
|
||||||
|
from copy import copy
|
||||||
|
from concurrent.futures import ThreadPoolExecutor, as_completed
|
||||||
|
from urllib.parse import urljoin, urlsplit, urlunsplit
|
||||||
|
|
||||||
|
import ssl
|
||||||
|
import openpyxl
|
||||||
|
import requests
|
||||||
|
from requests.adapters import HTTPAdapter
|
||||||
|
from urllib3.util.ssl_ import create_urllib3_context
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
from openpyxl.styles import Alignment, Font
|
||||||
|
from openpyxl.utils import get_column_letter
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
|
||||||
|
|
||||||
|
class WeakSSLAdapter(HTTPAdapter):
|
||||||
|
def init_poolmanager(self, *args, **kwargs):
|
||||||
|
ctx = create_urllib3_context()
|
||||||
|
ctx.set_ciphers('DEFAULT@SECLEVEL=0')
|
||||||
|
ctx.options |= 0x4
|
||||||
|
ctx.check_hostname = False
|
||||||
|
ctx.verify_mode = ssl.CERT_NONE
|
||||||
|
kwargs['ssl_context'] = ctx
|
||||||
|
return super().init_poolmanager(*args, **kwargs)
|
||||||
|
|
||||||
|
|
||||||
|
SESSION = requests.Session()
|
||||||
|
SESSION.headers.update(H)
|
||||||
|
|
||||||
|
TAB_CLASS_PATS = ['tab-ul', 'tab_ul', 'tabmenu', 'tab-menu', 'tab_menu',
|
||||||
|
'tablist', 'tab-list', 'tab_list', 'subtab', 'sub-tab', 'sub_tab']
|
||||||
|
|
||||||
|
MAXCOL = 27 # AA
|
||||||
|
CAT_COLS = [5, 6, 7, 8, 9, 10] # E F G H I J
|
||||||
|
HEADER_MERGES = {'B1:R1', 'S1:W1', 'Y1:AA1'}
|
||||||
|
|
||||||
|
|
||||||
|
def norm_url(u):
|
||||||
|
"""쿼리/프래그먼트 제거 + 끝 슬래시 정리한 비교용 키."""
|
||||||
|
if not u:
|
||||||
|
return ''
|
||||||
|
s = urlsplit(u)
|
||||||
|
path = s.path.rstrip('/')
|
||||||
|
return urlunsplit((s.scheme, s.netloc, path, '', '')).lower()
|
||||||
|
|
||||||
|
|
||||||
|
def abs_url(href, base):
|
||||||
|
if not href:
|
||||||
|
return ''
|
||||||
|
href = href.strip()
|
||||||
|
if href.startswith(('javascript:', '#')):
|
||||||
|
return ''
|
||||||
|
if href.startswith(('http://', 'https://')):
|
||||||
|
return href
|
||||||
|
return urljoin(base.rstrip('/') + '/', href.lstrip('/'))
|
||||||
|
|
||||||
|
|
||||||
|
def has_tab_class(el):
|
||||||
|
cls = ' '.join(el.get('class') or []).lower()
|
||||||
|
return any(p in cls for p in TAB_CLASS_PATS), cls
|
||||||
|
|
||||||
|
|
||||||
|
def find_tab_groups(soup, base):
|
||||||
|
"""조건 1·2 만족하는 탭 그룹 후보 반환: [(cls, [(label, abs_url),...]), ...]"""
|
||||||
|
out = []
|
||||||
|
seen = set()
|
||||||
|
for ul in soup.find_all('ul'):
|
||||||
|
ok, cls = has_tab_class(ul)
|
||||||
|
if not ok:
|
||||||
|
continue # UL 자신에 탭 클래스 필요 (쿼리필터·무클래스 ul 배제)
|
||||||
|
links = []
|
||||||
|
real = True
|
||||||
|
for a in ul.find_all('a'):
|
||||||
|
t = a.get_text(strip=True).replace('\xa0', '').strip()
|
||||||
|
raw = (a.get('href') or '').strip()
|
||||||
|
au = abs_url(raw, base)
|
||||||
|
if not t:
|
||||||
|
continue
|
||||||
|
if not au: # #anchor / javascript / 빈 href → 인페이지 탭
|
||||||
|
real = False
|
||||||
|
break
|
||||||
|
links.append((t, au))
|
||||||
|
if not real or len(links) < 2:
|
||||||
|
continue
|
||||||
|
key = tuple(nu for _, (nu) in [(t, u) for t, u in links])
|
||||||
|
if key in seen:
|
||||||
|
continue
|
||||||
|
seen.add(key)
|
||||||
|
out.append((cls, links))
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def fetch(r, url):
|
||||||
|
try:
|
||||||
|
resp = SESSION.get(url, timeout=15, verify=False)
|
||||||
|
# 바이트로 받아 BeautifulSoup가 meta charset으로 인코딩 판별 (apparent_encoding 오판 방지)
|
||||||
|
return r, url, resp.content, None
|
||||||
|
except Exception as e:
|
||||||
|
return r, url, None, str(e)[:60]
|
||||||
|
|
||||||
|
|
||||||
|
def load_flat(ws):
|
||||||
|
"""병합값을 채워넣어 행별 완전 데이터로 평탄화. 반환: rows(list of dict), 스타일/하이퍼링크 캡처."""
|
||||||
|
# 1) 병합값 채우기 (헤더 제외, 데이터 컬럼)
|
||||||
|
for mr in list(ws.merged_cells.ranges):
|
||||||
|
s = str(mr)
|
||||||
|
if s in HEADER_MERGES:
|
||||||
|
continue
|
||||||
|
top = ws.cell(mr.min_row, mr.min_col).value
|
||||||
|
ws.unmerge_cells(s)
|
||||||
|
for rr in range(mr.min_row, mr.max_row + 1):
|
||||||
|
for cc in range(mr.min_col, mr.max_col + 1):
|
||||||
|
if ws.cell(rr, cc).value in (None, ''):
|
||||||
|
ws.cell(rr, cc).value = top
|
||||||
|
rows = []
|
||||||
|
for r in range(3, ws.max_row + 1):
|
||||||
|
# 빈 행 스킵
|
||||||
|
if all(ws.cell(r, c).value in (None, '') for c in range(2, MAXCOL + 1)):
|
||||||
|
continue
|
||||||
|
vals = {c: ws.cell(r, c).value for c in range(1, MAXCOL + 1)}
|
||||||
|
styles = {}
|
||||||
|
for c in range(1, MAXCOL + 1):
|
||||||
|
sc = ws.cell(r, c)
|
||||||
|
styles[c] = (copy(sc.font), copy(sc.fill), copy(sc.border),
|
||||||
|
copy(sc.alignment), sc.number_format, copy(sc.protection))
|
||||||
|
hl = ws.cell(r, 11).hyperlink
|
||||||
|
rows.append({'src': r, 'vals': vals, 'styles': styles,
|
||||||
|
'hyperlink': hl.target if hl else None})
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def leaf_col(vals):
|
||||||
|
"""E~J 중 값이 있는 가장 깊은 컬럼 인덱스."""
|
||||||
|
deep = 5
|
||||||
|
for c in CAT_COLS:
|
||||||
|
if vals.get(c) not in (None, ''):
|
||||||
|
deep = c
|
||||||
|
return deep
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
xlsx, base, domain = sys.argv[1], sys.argv[2], sys.argv[3]
|
||||||
|
write = '--write' in sys.argv
|
||||||
|
if '--weak-ssl' in sys.argv:
|
||||||
|
SESSION.mount('https://', WeakSSLAdapter())
|
||||||
|
|
||||||
|
wb = openpyxl.load_workbook(xlsx)
|
||||||
|
ws = wb.active
|
||||||
|
rows = load_flat(ws)
|
||||||
|
existing = set()
|
||||||
|
for row in rows:
|
||||||
|
u = row['vals'].get(11)
|
||||||
|
if isinstance(u, str) and u.startswith('http'):
|
||||||
|
existing.add(norm_url(u))
|
||||||
|
|
||||||
|
# 같은 도메인 행 fetch
|
||||||
|
targets = [(row['src'], row['vals'].get(11)) for row in rows
|
||||||
|
if isinstance(row['vals'].get(11), str) and domain in row['vals'].get(11)]
|
||||||
|
html_by_src = {}
|
||||||
|
with ThreadPoolExecutor(max_workers=8) as ex:
|
||||||
|
futs = [ex.submit(fetch, r, u) for r, u in targets]
|
||||||
|
for fut in as_completed(futs):
|
||||||
|
r, url, html, err = fut.result()
|
||||||
|
if html:
|
||||||
|
html_by_src[r] = (url, html)
|
||||||
|
|
||||||
|
# 행별 확장 계획
|
||||||
|
plan = {} # src_row -> ordered [(label, url, is_existing)]
|
||||||
|
for row in rows:
|
||||||
|
src = row['src']
|
||||||
|
if src not in html_by_src:
|
||||||
|
continue
|
||||||
|
url, html = html_by_src[src]
|
||||||
|
cur = norm_url(url)
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
groups = find_tab_groups(soup, base)
|
||||||
|
chosen = None
|
||||||
|
for cls, links in groups:
|
||||||
|
tab_norms = [norm_url(u) for _, u in links]
|
||||||
|
if cur not in tab_norms:
|
||||||
|
continue # 조건3: 자기 탭그룹만
|
||||||
|
new_cnt = sum(1 for n in tab_norms if n not in existing)
|
||||||
|
if new_cnt < 1:
|
||||||
|
continue # 조건4: 신규 0 → 상위nav, 제외
|
||||||
|
chosen = links
|
||||||
|
break
|
||||||
|
if chosen:
|
||||||
|
ch = []
|
||||||
|
for label, u in chosen:
|
||||||
|
ch.append((label, u, norm_url(u) == cur))
|
||||||
|
plan[src] = ch
|
||||||
|
|
||||||
|
# 계획 출력
|
||||||
|
print(f'=== 탭 확장 계획: {len(plan)}개 그룹 ===\n')
|
||||||
|
total_new = 0
|
||||||
|
for src in sorted(plan):
|
||||||
|
row = next(r for r in rows if r['src'] == src)
|
||||||
|
v = row['vals']
|
||||||
|
lc = leaf_col(v)
|
||||||
|
childc = lc + 1
|
||||||
|
ev = v.get(5) or ''
|
||||||
|
fv = v.get(6) or ''
|
||||||
|
print(f'[행{src}] E={ev} F={fv} '
|
||||||
|
f'(leaf={get_column_letter(lc)} → 자식={get_column_letter(childc)})')
|
||||||
|
for label, u, exist in plan[src]:
|
||||||
|
tag = '재사용' if exist else '신규+'
|
||||||
|
if not exist:
|
||||||
|
total_new += 1
|
||||||
|
print(f' [{tag}] {label} -> {u}')
|
||||||
|
print()
|
||||||
|
print(f'신규 추가 행: {total_new}개 / 기존 {len(rows)}행 → 최종 {len(rows)+total_new}행')
|
||||||
|
|
||||||
|
if not write:
|
||||||
|
print('\n(계획만 출력. 실제 기입하려면 --write)')
|
||||||
|
return
|
||||||
|
|
||||||
|
# ===== 실제 기입 =====
|
||||||
|
backup = xlsx.replace('.xlsx', '_backup_tab전.xlsx')
|
||||||
|
shutil.copy(xlsx, backup)
|
||||||
|
print(f'\n백업: {backup}')
|
||||||
|
|
||||||
|
# 새 평탄 행 목록 구성
|
||||||
|
out_rows = [] # 각 원소 = {'vals':..., 'style_src':row_dict, 'url':..}
|
||||||
|
for row in rows:
|
||||||
|
src = row['src']
|
||||||
|
if src in plan:
|
||||||
|
v = row['vals']
|
||||||
|
lc = leaf_col(v)
|
||||||
|
childc = lc + 1
|
||||||
|
for label, u, exist in plan[src]:
|
||||||
|
nv = dict(v)
|
||||||
|
# leaf 보다 깊은 컬럼 초기화 후 자식 컬럼에 탭명
|
||||||
|
for c in CAT_COLS:
|
||||||
|
if c > lc:
|
||||||
|
nv[c] = None
|
||||||
|
nv[childc] = label
|
||||||
|
if exist:
|
||||||
|
nv[11] = v.get(11) # 기존 URL/데이터 유지
|
||||||
|
out_rows.append({'vals': nv, 'style': row, 'url': nv.get(11)})
|
||||||
|
else:
|
||||||
|
for c in range(12, MAXCOL + 1):
|
||||||
|
nv[c] = None # L~T 비움 (신규)
|
||||||
|
nv[11] = u
|
||||||
|
nv[19] = None
|
||||||
|
out_rows.append({'vals': nv, 'style': row, 'url': u})
|
||||||
|
else:
|
||||||
|
out_rows.append({'vals': dict(row['vals']), 'style': row, 'url': row['vals'].get(11)})
|
||||||
|
|
||||||
|
# 데이터 영역 클리어
|
||||||
|
for r in range(3, ws.max_row + 1):
|
||||||
|
for c in range(1, MAXCOL + 1):
|
||||||
|
ws.cell(r, c).value = None
|
||||||
|
ws.cell(r, c).hyperlink = None
|
||||||
|
|
||||||
|
START = 3
|
||||||
|
for i, orow in enumerate(out_rows):
|
||||||
|
r = START + i
|
||||||
|
sty = orow['style']['styles']
|
||||||
|
for c in range(1, MAXCOL + 1):
|
||||||
|
cell = ws.cell(r, c)
|
||||||
|
f, fl, bd, al, nf, pr = sty[c]
|
||||||
|
cell.font = copy(f); cell.fill = copy(fl); cell.border = copy(bd)
|
||||||
|
cell.alignment = copy(al); cell.number_format = nf; cell.protection = copy(pr)
|
||||||
|
v = orow['vals']
|
||||||
|
for c in range(1, MAXCOL + 1):
|
||||||
|
ws.cell(r, c).value = v.get(c)
|
||||||
|
ws.cell(r, 2).value = i + 1 # B 순번 재부여
|
||||||
|
END = START + len(out_rows) - 1
|
||||||
|
|
||||||
|
# D/E/F 재병합 (G는 leaf라 병합 안 함)
|
||||||
|
center = Alignment(horizontal='center', vertical='center', wrap_text=True)
|
||||||
|
left = Alignment(horizontal='left', vertical='center', wrap_text=False)
|
||||||
|
|
||||||
|
def merge_runs(col_letter, col_idx, group_cols=()):
|
||||||
|
cur_val = ws.cell(START, col_idx).value
|
||||||
|
cur_grp = tuple(ws.cell(START, g).value for g in group_cols)
|
||||||
|
run_start = START
|
||||||
|
runs = []
|
||||||
|
for r in range(START + 1, END + 1):
|
||||||
|
val = ws.cell(r, col_idx).value
|
||||||
|
grp = tuple(ws.cell(r, gg).value for gg in group_cols)
|
||||||
|
if val == cur_val and grp == cur_grp:
|
||||||
|
continue
|
||||||
|
if cur_val not in (None, '') and r - 1 > run_start:
|
||||||
|
runs.append((run_start, r - 1))
|
||||||
|
cur_val, cur_grp, run_start = val, grp, r
|
||||||
|
if cur_val not in (None, '') and END > run_start:
|
||||||
|
runs.append((run_start, END))
|
||||||
|
for s, e in runs:
|
||||||
|
ws.merge_cells(f'{col_letter}{s}:{col_letter}{e}')
|
||||||
|
ws.cell(s, col_idx).alignment = center
|
||||||
|
|
||||||
|
merge_runs('F', 6, group_cols=(4, 5))
|
||||||
|
merge_runs('E', 5, group_cols=(4,))
|
||||||
|
merge_runs('D', 4)
|
||||||
|
for r in range(START, END + 1):
|
||||||
|
for c in (4, 5, 6, 7):
|
||||||
|
if ws.cell(r, c).value is not None:
|
||||||
|
ws.cell(r, c).alignment = center
|
||||||
|
|
||||||
|
for r in range(1, END + 1):
|
||||||
|
ws.row_dimensions[r].height = 15
|
||||||
|
|
||||||
|
# K 하이퍼링크 재설정
|
||||||
|
for r in range(START, END + 1):
|
||||||
|
cell = ws.cell(r, 11)
|
||||||
|
u = cell.value
|
||||||
|
if isinstance(u, str) and u.startswith(('http://', 'https://')):
|
||||||
|
cell.hyperlink = u
|
||||||
|
old = cell.font
|
||||||
|
cell.font = Font(name=old.name or '맑은 고딕', size=old.size or 11,
|
||||||
|
bold=old.bold, italic=old.italic,
|
||||||
|
color='0000FF', underline='single')
|
||||||
|
cell.alignment = left
|
||||||
|
|
||||||
|
wb.save(xlsx)
|
||||||
|
print(f'저장 완료: {xlsx} (총 {len(out_rows)}행)')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
439
_스크립트/_tab_expand_seosan.py
Normal file
439
_스크립트/_tab_expand_seosan.py
Normal file
@ -0,0 +1,439 @@
|
|||||||
|
"""본문 탭(서브내비) → 카테고리 하위 확장.
|
||||||
|
|
||||||
|
규칙(공주시 기준, 일반화):
|
||||||
|
본문에서 '탭 UL'을 찾아, 아래 조건을 모두 만족하는 탭 그룹만 하위 카테고리로 확장한다.
|
||||||
|
1) UL(또는 직계 div)의 class 에 탭 패턴(tab-ul 등)이 있다.
|
||||||
|
2) 탭이 2개 이상이고, 모든 탭 href 가 '실제 페이지 링크'다.
|
||||||
|
(#anchor 인페이지 탭, javascript:, 빈 href, 쿼리스트링만 다른 필터 탭은 제외)
|
||||||
|
3) 현재 페이지 URL 이 탭 URL 집합에 포함된다(= 자기 자신의 탭 그룹).
|
||||||
|
4) 탭 중 '엑셀에 아직 없는 URL'이 1개 이상 있다(이미 사이트맵에 다 있으면 상위 nav → 제외).
|
||||||
|
|
||||||
|
확장 방식(=사용자 지시: 탭들은 한 단계 아래 컬럼으로):
|
||||||
|
- 원래 행의 leaf 컬럼(E~J 중 가장 깊은 값)을 부모(카테고리)로 두고,
|
||||||
|
탭들을 leaf+1 컬럼에 탭 순서대로 채운다.
|
||||||
|
- 현재 페이지와 같은 URL의 탭 = 원래 행을 재사용(L~T 기존 데이터 보존, G만 탭명 기입).
|
||||||
|
- 나머지 탭 = 신규 행(L~T 비움, Phase 2~4 별도 수행 대상).
|
||||||
|
|
||||||
|
사용법:
|
||||||
|
python -X utf8 _tab_expand.py <엑셀경로> <BASE_URL> <도메인키워드> [--write]
|
||||||
|
예) python -X utf8 _tab_expand.py 충청남도/2.공주시/충청남도_공주시.xlsx https://www.gongju.go.kr gongju.go.kr
|
||||||
|
(--write 없으면 계획만 출력 / 있으면 백업 후 실제 기입)
|
||||||
|
"""
|
||||||
|
import re
|
||||||
|
import sys
|
||||||
|
import shutil
|
||||||
|
import warnings
|
||||||
|
from copy import copy
|
||||||
|
from concurrent.futures import ThreadPoolExecutor, as_completed
|
||||||
|
from urllib.parse import urljoin, urlsplit, urlunsplit
|
||||||
|
|
||||||
|
import ssl
|
||||||
|
import openpyxl
|
||||||
|
import requests
|
||||||
|
from requests.adapters import HTTPAdapter
|
||||||
|
from urllib3.util.ssl_ import create_urllib3_context
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
from openpyxl.styles import Alignment, Font
|
||||||
|
from openpyxl.utils import get_column_letter
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
|
||||||
|
|
||||||
|
class WeakSSLAdapter(HTTPAdapter):
|
||||||
|
def init_poolmanager(self, *args, **kwargs):
|
||||||
|
ctx = create_urllib3_context()
|
||||||
|
ctx.set_ciphers('DEFAULT@SECLEVEL=0')
|
||||||
|
ctx.options |= 0x4
|
||||||
|
ctx.check_hostname = False
|
||||||
|
ctx.verify_mode = ssl.CERT_NONE
|
||||||
|
kwargs['ssl_context'] = ctx
|
||||||
|
return super().init_poolmanager(*args, **kwargs)
|
||||||
|
|
||||||
|
|
||||||
|
SESSION = requests.Session()
|
||||||
|
SESSION.headers.update(H)
|
||||||
|
|
||||||
|
TAB_CLASS_PATS = ['tab-ul', 'tab_ul', 'tabmenu', 'tab-menu', 'tab_menu',
|
||||||
|
'tablist', 'tab-list', 'tab_list', 'subtab', 'sub-tab', 'sub_tab']
|
||||||
|
# 토큰 단위 탭 클래스 정규식: 'basic_tab', 'tab_wrap', 'tab-ul', 단독 'tab' 등 매칭.
|
||||||
|
# (UL 자신뿐 아니라 직계 부모 div 클래스도 검사 → div.basic_tab > ul 구조 대응)
|
||||||
|
TAB_CLASS_RE = re.compile(
|
||||||
|
r'(?:^|[-_ ])tab(?:[-_ ]|$)|tabmenu|tablist|tab[-_]?(?:ul|wrap|list|menu)|basic[-_]tab')
|
||||||
|
|
||||||
|
MAXCOL = 27 # AA
|
||||||
|
CAT_COLS = [5, 6, 7, 8, 9, 10] # E F G H I J
|
||||||
|
HEADER_MERGES = {'B1:R1', 'S1:W1', 'Y1:AA1'}
|
||||||
|
|
||||||
|
|
||||||
|
def norm_url(u):
|
||||||
|
"""비교용 키: 프래그먼트 제거 + 끝 슬래시 정리 + 쿼리 정렬 보존.
|
||||||
|
(쿼리만 다른 게시판 분류 탭 ?code=A vs ?code=B 를 서로 다른 URL로 구분 →
|
||||||
|
is_cur·existing 중복판정 오류 방지)"""
|
||||||
|
if not u:
|
||||||
|
return ''
|
||||||
|
s = urlsplit(u)
|
||||||
|
path = s.path.rstrip('/')
|
||||||
|
q = '&'.join(sorted(s.query.split('&'))) if s.query else ''
|
||||||
|
return urlunsplit((s.scheme, s.netloc, path, q, '')).lower()
|
||||||
|
|
||||||
|
|
||||||
|
def path_key(u):
|
||||||
|
"""쿼리 제거한 경로 키 (필터 탭 판별용)."""
|
||||||
|
if not u:
|
||||||
|
return ''
|
||||||
|
s = urlsplit(u)
|
||||||
|
return urlunsplit((s.scheme, s.netloc, s.path.rstrip('/'), '', '')).lower()
|
||||||
|
|
||||||
|
|
||||||
|
def abs_url(href, base):
|
||||||
|
if not href:
|
||||||
|
return ''
|
||||||
|
href = href.strip()
|
||||||
|
if href.startswith(('javascript:', '#')):
|
||||||
|
return ''
|
||||||
|
if href.startswith(('http://', 'https://')):
|
||||||
|
return href
|
||||||
|
return urljoin(base.rstrip('/') + '/', href.lstrip('/'))
|
||||||
|
|
||||||
|
|
||||||
|
def has_tab_class(el):
|
||||||
|
if el is None or not getattr(el, 'get', None):
|
||||||
|
return False, ''
|
||||||
|
cls = ' '.join(el.get('class') or []).lower()
|
||||||
|
ok = any(p in cls for p in TAB_CLASS_PATS) or bool(TAB_CLASS_RE.search(cls))
|
||||||
|
return ok, cls
|
||||||
|
|
||||||
|
|
||||||
|
def _has_on_li(ul):
|
||||||
|
"""ul 직계 li(또는 그 a)에 활성 탭 마커(on/active/current/selected)가 있나."""
|
||||||
|
marks = {'on', 'active', 'current', 'selected', 'sel'}
|
||||||
|
for li in ul.find_all('li', recursive=False):
|
||||||
|
if marks & set(c.lower() for c in (li.get('class') or [])):
|
||||||
|
return True
|
||||||
|
a = li.find('a')
|
||||||
|
if a and (marks & set(c.lower() for c in (a.get('class') or []))):
|
||||||
|
return True
|
||||||
|
return False
|
||||||
|
|
||||||
|
|
||||||
|
def find_tab_groups(soup, base):
|
||||||
|
"""조건 1·2 만족하는 탭 그룹 후보 반환: [(cls, [(label, abs_url),...], has_on), ...]"""
|
||||||
|
out = []
|
||||||
|
seen = set()
|
||||||
|
for ul in soup.find_all('ul'):
|
||||||
|
ok, cls = has_tab_class(ul)
|
||||||
|
if not ok:
|
||||||
|
# UL 무클래스라도 직계 부모 div 에 탭 클래스가 있으면 인정 (div.basic_tab > ul)
|
||||||
|
pok, pcls = has_tab_class(ul.parent)
|
||||||
|
if not pok:
|
||||||
|
continue # 쿼리필터·무클래스 ul 배제
|
||||||
|
cls = pcls
|
||||||
|
links = []
|
||||||
|
real = True
|
||||||
|
for a in ul.find_all('a'):
|
||||||
|
t = a.get_text(strip=True).replace('\xa0', '').strip()
|
||||||
|
raw = (a.get('href') or '').strip()
|
||||||
|
au = abs_url(raw, base)
|
||||||
|
if not t:
|
||||||
|
continue
|
||||||
|
if not au: # #anchor / javascript / 빈 href → 인페이지 탭
|
||||||
|
real = False
|
||||||
|
break
|
||||||
|
links.append((t, au))
|
||||||
|
if not real or len(links) < 2:
|
||||||
|
continue
|
||||||
|
key = tuple(nu for _, (nu) in [(t, u) for t, u in links])
|
||||||
|
if key in seen:
|
||||||
|
continue
|
||||||
|
seen.add(key)
|
||||||
|
out.append((cls, links, _has_on_li(ul)))
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def fetch(r, url):
|
||||||
|
try:
|
||||||
|
resp = SESSION.get(url, timeout=15, verify=False)
|
||||||
|
# 바이트로 받아 BeautifulSoup가 meta charset으로 인코딩 판별 (apparent_encoding 오판 방지)
|
||||||
|
return r, url, resp.content, None
|
||||||
|
except Exception as e:
|
||||||
|
return r, url, None, str(e)[:60]
|
||||||
|
|
||||||
|
|
||||||
|
def load_flat(ws):
|
||||||
|
"""병합값을 채워넣어 행별 완전 데이터로 평탄화. 반환: rows(list of dict), 스타일/하이퍼링크 캡처."""
|
||||||
|
# 1) 병합값 채우기 (헤더 제외, 데이터 컬럼)
|
||||||
|
for mr in list(ws.merged_cells.ranges):
|
||||||
|
s = str(mr)
|
||||||
|
if s in HEADER_MERGES:
|
||||||
|
continue
|
||||||
|
top = ws.cell(mr.min_row, mr.min_col).value
|
||||||
|
ws.unmerge_cells(s)
|
||||||
|
for rr in range(mr.min_row, mr.max_row + 1):
|
||||||
|
for cc in range(mr.min_col, mr.max_col + 1):
|
||||||
|
if ws.cell(rr, cc).value in (None, ''):
|
||||||
|
ws.cell(rr, cc).value = top
|
||||||
|
rows = []
|
||||||
|
for r in range(3, ws.max_row + 1):
|
||||||
|
# 빈 행 스킵
|
||||||
|
if all(ws.cell(r, c).value in (None, '') for c in range(2, MAXCOL + 1)):
|
||||||
|
continue
|
||||||
|
vals = {c: ws.cell(r, c).value for c in range(1, MAXCOL + 1)}
|
||||||
|
styles = {}
|
||||||
|
for c in range(1, MAXCOL + 1):
|
||||||
|
sc = ws.cell(r, c)
|
||||||
|
styles[c] = (copy(sc.font), copy(sc.fill), copy(sc.border),
|
||||||
|
copy(sc.alignment), sc.number_format, copy(sc.protection))
|
||||||
|
hl = ws.cell(r, 11).hyperlink
|
||||||
|
rows.append({'src': r, 'vals': vals, 'styles': styles,
|
||||||
|
'hyperlink': hl.target if hl else None})
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def leaf_col(vals):
|
||||||
|
"""E~J 중 값이 있는 가장 깊은 컬럼 인덱스."""
|
||||||
|
deep = 5
|
||||||
|
for c in CAT_COLS:
|
||||||
|
if vals.get(c) not in (None, ''):
|
||||||
|
deep = c
|
||||||
|
return deep
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
xlsx, base, domain = sys.argv[1], sys.argv[2], sys.argv[3]
|
||||||
|
write = '--write' in sys.argv
|
||||||
|
# menuCd 벤더(고창·임실·정읍·진안 등 index.{name}?menuCd=…): 모든 페이지가 같은 경로,
|
||||||
|
# 쿼리(menuCd)만 다름 = 다른 페이지. → 쿼리-경로 가드(조건2) 스킵 + cur∈탭(조건3) 대신
|
||||||
|
# div.basic_tab 등에 활성탭(li.on) 마커가 있는 '자기 sub-nav'만 인정(랜딩 URL이 탭과 달라도).
|
||||||
|
menucd = '--menucd' in sys.argv
|
||||||
|
# 서산 등 쿼리기반(contents.do?key=, selectBbsNttList.do?bbsNo=) 벤더:
|
||||||
|
# 모든 탭이 같은 경로·쿼리만 다름(=다른 페이지)이고 한 탭바에 페이지+게시판 혼재.
|
||||||
|
# → 쿼리가드(같은경로=필터 제외) 끄고, div.tab_menu>ul.tab_button 그룹만 인정.
|
||||||
|
noqg = '--noqg' in sys.argv
|
||||||
|
if '--weak-ssl' in sys.argv:
|
||||||
|
SESSION.mount('https://', WeakSSLAdapter())
|
||||||
|
|
||||||
|
wb = openpyxl.load_workbook(xlsx)
|
||||||
|
ws = wb.active
|
||||||
|
rows = load_flat(ws)
|
||||||
|
existing = set()
|
||||||
|
for row in rows:
|
||||||
|
u = row['vals'].get(11)
|
||||||
|
if isinstance(u, str) and u.startswith('http'):
|
||||||
|
existing.add(norm_url(u))
|
||||||
|
|
||||||
|
# 같은 도메인 행 fetch
|
||||||
|
targets = [(row['src'], row['vals'].get(11)) for row in rows
|
||||||
|
if isinstance(row['vals'].get(11), str) and domain in row['vals'].get(11)]
|
||||||
|
html_by_src = {}
|
||||||
|
with ThreadPoolExecutor(max_workers=8) as ex:
|
||||||
|
futs = [ex.submit(fetch, r, u) for r, u in targets]
|
||||||
|
for fut in as_completed(futs):
|
||||||
|
r, url, html, err = fut.result()
|
||||||
|
if html:
|
||||||
|
html_by_src[r] = (url, html)
|
||||||
|
|
||||||
|
# 이미 사이트맵에 자식행이 있는 '랜딩행' 집합 (다음 평탄행이 같은 상위카테고리 + 더 깊은 leaf).
|
||||||
|
# menuCd 벤더 랜딩은 첫 자식의 탭을 렌더하므로, 확장하면 그 자식행과 중복 → 제외.
|
||||||
|
landing_src = set()
|
||||||
|
for i in range(len(rows) - 1):
|
||||||
|
v, nv = rows[i]['vals'], rows[i + 1]['vals']
|
||||||
|
lc = leaf_col(v)
|
||||||
|
if lc < 10 and nv.get(lc + 1) not in (None, '') \
|
||||||
|
and all((v.get(c) or '') == (nv.get(c) or '') for c in CAT_COLS if c <= lc):
|
||||||
|
landing_src.add(rows[i]['src'])
|
||||||
|
|
||||||
|
# 행별 확장 계획
|
||||||
|
plan = {} # src_row -> ordered [(label, url, is_existing)]
|
||||||
|
planned_new = set() # 이미 어느 부모행이 추가한 신규 탭 URL(전역 중복 방지)
|
||||||
|
for row in rows:
|
||||||
|
src = row['src']
|
||||||
|
if src not in html_by_src:
|
||||||
|
continue
|
||||||
|
if menucd and src in landing_src:
|
||||||
|
continue # 자식 보유 랜딩 → 확장 금지(첫 자식 탭 중복 방지)
|
||||||
|
url, html = html_by_src[src]
|
||||||
|
cur = norm_url(url)
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
groups = find_tab_groups(soup, base)
|
||||||
|
chosen = None
|
||||||
|
for cls, links, has_on in groups:
|
||||||
|
if menucd:
|
||||||
|
# menuCd 벤더: 현재 페이지와 '같은 경로'(menuCd만 다른 형제)인 탭만 인정.
|
||||||
|
# → /gochang/toc/GC…(향토문화대전 백과 목록) 등 다른경로 링크 배제.
|
||||||
|
# 활성탭(li.on) 마커가 있는 자기 sub-nav만(랜딩 URL이 탭에 없어도 OK).
|
||||||
|
if not has_on:
|
||||||
|
continue
|
||||||
|
links = [(t, u) for t, u in links if path_key(u) == path_key(cur)]
|
||||||
|
if len(links) < 2:
|
||||||
|
continue
|
||||||
|
tab_norms = [norm_url(u) for _, u in links]
|
||||||
|
else:
|
||||||
|
if noqg and not ('tab_button' in cls or 'tab_menu' in cls):
|
||||||
|
continue # 서산: 본문 탭바(tab_menu/tab_button)만 — GNB·푸터 nav 배제
|
||||||
|
tab_norms = [norm_url(u) for _, u in links]
|
||||||
|
if not noqg and len({path_key(u) for _, u in links}) == 1:
|
||||||
|
continue # 조건2: 같은 경로(쿼리만 다른 필터 탭) → 제외 (noqg면 쿼리=다른페이지라 허용)
|
||||||
|
if cur not in tab_norms:
|
||||||
|
continue # 조건3: 자기 탭그룹만
|
||||||
|
new_cnt = sum(1 for n in tab_norms if n not in existing)
|
||||||
|
if new_cnt < 1:
|
||||||
|
continue # 조건4: 신규 0 → 상위nav, 제외
|
||||||
|
chosen = links
|
||||||
|
break
|
||||||
|
if chosen:
|
||||||
|
ch = []
|
||||||
|
for label, u in chosen:
|
||||||
|
is_cur = norm_url(u) == cur
|
||||||
|
nu = norm_url(u)
|
||||||
|
if not is_cur and nu in existing:
|
||||||
|
continue # 이미 사이트맵에 별도 행으로 존재 → 중복행 방지(skip)
|
||||||
|
if not is_cur and nu in planned_new:
|
||||||
|
continue # 다른 부모행이 이미 추가한 신규 URL → 전역 중복 방지
|
||||||
|
if not is_cur:
|
||||||
|
planned_new.add(nu)
|
||||||
|
is_ext = domain not in u # 같은 도메인 아님 = 외부 사이트 탭
|
||||||
|
ch.append((label, u, is_cur, is_ext))
|
||||||
|
# 신규(비-cur) 탭이 1개 이상일 때만 확장. 아니면 원본행 그대로 둠(삭제 방지).
|
||||||
|
if any(not c[2] for c in ch):
|
||||||
|
plan[src] = ch
|
||||||
|
|
||||||
|
# 계획 출력
|
||||||
|
print(f'=== 탭 확장 계획: {len(plan)}개 그룹 ===\n')
|
||||||
|
total_new = 0
|
||||||
|
for src in sorted(plan):
|
||||||
|
row = next(r for r in rows if r['src'] == src)
|
||||||
|
v = row['vals']
|
||||||
|
lc = leaf_col(v)
|
||||||
|
childc = lc + 1
|
||||||
|
ev = v.get(5) or ''
|
||||||
|
fv = v.get(6) or ''
|
||||||
|
print(f'[행{src}] E={ev} F={fv} '
|
||||||
|
f'(leaf={get_column_letter(lc)} → 자식={get_column_letter(childc)})')
|
||||||
|
for label, u, is_cur, is_ext in plan[src]:
|
||||||
|
if is_cur:
|
||||||
|
tag = '재사용'
|
||||||
|
elif is_ext:
|
||||||
|
tag = '신규+(외부=사이트)'
|
||||||
|
total_new += 1
|
||||||
|
else:
|
||||||
|
tag = '신규+'
|
||||||
|
total_new += 1
|
||||||
|
print(f' [{tag}] {label} -> {u}')
|
||||||
|
print()
|
||||||
|
print(f'신규 추가 행: {total_new}개 / 기존 {len(rows)}행 → 최종 {len(rows)+total_new}행')
|
||||||
|
|
||||||
|
if not write:
|
||||||
|
print('\n(계획만 출력. 실제 기입하려면 --write)')
|
||||||
|
return
|
||||||
|
|
||||||
|
# ===== 실제 기입 =====
|
||||||
|
backup = xlsx.replace('.xlsx', '_backup_tab전.xlsx')
|
||||||
|
shutil.copy(xlsx, backup)
|
||||||
|
print(f'\n백업: {backup}')
|
||||||
|
|
||||||
|
# 새 평탄 행 목록 구성
|
||||||
|
out_rows = [] # 각 원소 = {'vals':..., 'style_src':row_dict, 'url':..}
|
||||||
|
for row in rows:
|
||||||
|
src = row['src']
|
||||||
|
if src in plan:
|
||||||
|
v = row['vals']
|
||||||
|
lc = leaf_col(v)
|
||||||
|
childc = lc + 1
|
||||||
|
for label, u, is_cur, is_ext in plan[src]:
|
||||||
|
nv = dict(v)
|
||||||
|
# leaf 보다 깊은 컬럼 초기화 후 자식 컬럼에 탭명
|
||||||
|
for c in CAT_COLS:
|
||||||
|
if c > lc:
|
||||||
|
nv[c] = None
|
||||||
|
nv[childc] = label
|
||||||
|
if is_cur:
|
||||||
|
nv[11] = v.get(11) # 기존 URL/데이터 유지
|
||||||
|
out_rows.append({'vals': nv, 'style': row, 'url': nv.get(11)})
|
||||||
|
else:
|
||||||
|
for c in range(12, MAXCOL + 1):
|
||||||
|
nv[c] = None # L~T 비움 (신규)
|
||||||
|
nv[11] = u
|
||||||
|
nv[19] = None
|
||||||
|
if is_ext: # 외부 사이트 탭 → 즉시 L=사이트/M=1(Phase234 대상 아님)
|
||||||
|
nv[12] = '사이트'
|
||||||
|
nv[13] = 1
|
||||||
|
out_rows.append({'vals': nv, 'style': row, 'url': u})
|
||||||
|
else:
|
||||||
|
out_rows.append({'vals': dict(row['vals']), 'style': row, 'url': row['vals'].get(11)})
|
||||||
|
|
||||||
|
# 데이터 영역 클리어
|
||||||
|
for r in range(3, ws.max_row + 1):
|
||||||
|
for c in range(1, MAXCOL + 1):
|
||||||
|
ws.cell(r, c).value = None
|
||||||
|
ws.cell(r, c).hyperlink = None
|
||||||
|
|
||||||
|
START = 3
|
||||||
|
for i, orow in enumerate(out_rows):
|
||||||
|
r = START + i
|
||||||
|
sty = orow['style']['styles']
|
||||||
|
for c in range(1, MAXCOL + 1):
|
||||||
|
cell = ws.cell(r, c)
|
||||||
|
f, fl, bd, al, nf, pr = sty[c]
|
||||||
|
cell.font = copy(f); cell.fill = copy(fl); cell.border = copy(bd)
|
||||||
|
cell.alignment = copy(al); cell.number_format = nf; cell.protection = copy(pr)
|
||||||
|
v = orow['vals']
|
||||||
|
for c in range(1, MAXCOL + 1):
|
||||||
|
ws.cell(r, c).value = v.get(c)
|
||||||
|
ws.cell(r, 2).value = i + 1 # B 순번 재부여
|
||||||
|
END = START + len(out_rows) - 1
|
||||||
|
|
||||||
|
# D/E/F 재병합 (G는 leaf라 병합 안 함)
|
||||||
|
center = Alignment(horizontal='center', vertical='center', wrap_text=True)
|
||||||
|
left = Alignment(horizontal='left', vertical='center', wrap_text=False)
|
||||||
|
|
||||||
|
def merge_runs(col_letter, col_idx, group_cols=()):
|
||||||
|
cur_val = ws.cell(START, col_idx).value
|
||||||
|
cur_grp = tuple(ws.cell(START, g).value for g in group_cols)
|
||||||
|
run_start = START
|
||||||
|
runs = []
|
||||||
|
for r in range(START + 1, END + 1):
|
||||||
|
val = ws.cell(r, col_idx).value
|
||||||
|
grp = tuple(ws.cell(r, gg).value for gg in group_cols)
|
||||||
|
if val == cur_val and grp == cur_grp:
|
||||||
|
continue
|
||||||
|
if cur_val not in (None, '') and r - 1 > run_start:
|
||||||
|
runs.append((run_start, r - 1))
|
||||||
|
cur_val, cur_grp, run_start = val, grp, r
|
||||||
|
if cur_val not in (None, '') and END > run_start:
|
||||||
|
runs.append((run_start, END))
|
||||||
|
for s, e in runs:
|
||||||
|
ws.merge_cells(f'{col_letter}{s}:{col_letter}{e}')
|
||||||
|
ws.cell(s, col_idx).alignment = center
|
||||||
|
|
||||||
|
merge_runs('F', 6, group_cols=(4, 5))
|
||||||
|
merge_runs('E', 5, group_cols=(4,))
|
||||||
|
merge_runs('D', 4)
|
||||||
|
for r in range(START, END + 1):
|
||||||
|
for c in (4, 5, 6, 7):
|
||||||
|
if ws.cell(r, c).value is not None:
|
||||||
|
ws.cell(r, c).alignment = center
|
||||||
|
|
||||||
|
for r in range(1, END + 1):
|
||||||
|
ws.row_dimensions[r].height = 15
|
||||||
|
|
||||||
|
# K 하이퍼링크 재설정
|
||||||
|
for r in range(START, END + 1):
|
||||||
|
cell = ws.cell(r, 11)
|
||||||
|
u = cell.value
|
||||||
|
if isinstance(u, str) and u.startswith(('http://', 'https://')):
|
||||||
|
cell.hyperlink = u
|
||||||
|
old = cell.font
|
||||||
|
cell.font = Font(name=old.name or '맑은 고딕', size=old.size or 11,
|
||||||
|
bold=old.bold, italic=old.italic,
|
||||||
|
color='0000FF', underline='single')
|
||||||
|
cell.alignment = left
|
||||||
|
|
||||||
|
wb.save(xlsx)
|
||||||
|
print(f'저장 완료: {xlsx} (총 {len(out_rows)}행)')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
220
_스크립트/_tab_phase234.py
Normal file
220
_스크립트/_tab_phase234.py
Normal file
@ -0,0 +1,220 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""탭 확장으로 새로 추가된 행만 골라 Phase 2~4(L/M/N/O/P) 수집.
|
||||||
|
|
||||||
|
- 대상: L(게시판형태)이 비어있고 K가 같은 도메인 http URL 인 행 = 신규 탭 행.
|
||||||
|
- L/M/N: 해당 기관 _phase234.py 의 get_body/detect_form/detect_media/extract_detail_urls 재사용.
|
||||||
|
- O/P : 확정 KOGL 규칙(_recheck_kogl_all 의 BROAD_IMG_PAT + detect_split + decide_O).
|
||||||
|
이미지명 우선, 게시판은 같은도메인 상세 5건 추적, 이미지≠링크면 S열 '링크주소 오기'.
|
||||||
|
- 인코딩: 바이트로 받아 BeautifulSoup 자동판별(읍면동 등 오판 방지).
|
||||||
|
- 기존 행(L 이미 채워짐)은 절대 건드리지 않음.
|
||||||
|
|
||||||
|
사용: python -X utf8 _tab_phase234.py <엑셀경로> <Phase234 모듈경로> <도메인키워드> [body_sel(콤마)]
|
||||||
|
- 개별형(공주시 _phase234.py: get_body(soup)) → body_sel 생략
|
||||||
|
- 일괄형(_chungnam_phase234_all.py: get_body(soup, sel)) → body_sel 지정(미지정 시 #txt,#contents,main)
|
||||||
|
예) python -X utf8 _tab_phase234.py 충청남도/2.공주시/공주시_탭확장.xlsx 충청남도/2.공주시/_phase234.py gongju.go.kr
|
||||||
|
python -X utf8 _tab_phase234.py 충청남도/4.논산시/충청남도_논산시.xlsx _chungnam_phase234_all.py nonsan.go.kr "#txt,#contents,main"
|
||||||
|
"""
|
||||||
|
import re
|
||||||
|
import sys
|
||||||
|
import time
|
||||||
|
import inspect
|
||||||
|
import importlib.util
|
||||||
|
import warnings
|
||||||
|
from urllib.parse import urlparse
|
||||||
|
from concurrent.futures import ThreadPoolExecutor, as_completed
|
||||||
|
|
||||||
|
import ssl
|
||||||
|
import openpyxl
|
||||||
|
import requests
|
||||||
|
from requests.adapters import HTTPAdapter
|
||||||
|
from urllib3.util.ssl_ import create_urllib3_context
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
|
||||||
|
|
||||||
|
class WeakSSLAdapter(HTTPAdapter):
|
||||||
|
def init_poolmanager(self, *args, **kwargs):
|
||||||
|
ctx = create_urllib3_context()
|
||||||
|
ctx.set_ciphers('DEFAULT@SECLEVEL=0')
|
||||||
|
ctx.options |= 0x4
|
||||||
|
ctx.check_hostname = False
|
||||||
|
ctx.verify_mode = ssl.CERT_NONE
|
||||||
|
kwargs['ssl_context'] = ctx
|
||||||
|
return super().init_poolmanager(*args, **kwargs)
|
||||||
|
|
||||||
|
|
||||||
|
SESSION = requests.Session()
|
||||||
|
SESSION.headers.update(H)
|
||||||
|
|
||||||
|
BROAD_IMG_PAT = re.compile(r'(?:new_)?img_open(?:type|code)(\d{1,2})\.(?:png|jpe?g|gif)', re.I)
|
||||||
|
|
||||||
|
|
||||||
|
def load_module(path):
|
||||||
|
spec = importlib.util.spec_from_file_location('city_p234', path)
|
||||||
|
mod = importlib.util.module_from_spec(spec)
|
||||||
|
spec.loader.exec_module(mod)
|
||||||
|
return mod
|
||||||
|
|
||||||
|
|
||||||
|
def fetch_soup(url, timeout=14):
|
||||||
|
try:
|
||||||
|
r = SESSION.get(url, timeout=timeout, verify=False)
|
||||||
|
if r.status_code == 200:
|
||||||
|
return BeautifulSoup(r.content, 'html.parser') # 바이트 → 자동 인코딩
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def valid(n):
|
||||||
|
return 1 <= n <= 4
|
||||||
|
|
||||||
|
|
||||||
|
def detect_split(body, LINK_PAT):
|
||||||
|
img_t, link_t = set(), set()
|
||||||
|
for a in body.find_all('a', href=True):
|
||||||
|
m = LINK_PAT.search(a['href'])
|
||||||
|
if m and valid(int(m.group(1))):
|
||||||
|
link_t.add(int(m.group(1)))
|
||||||
|
blob = ' '.join(filter(None, (img.get('src', '') for img in body.find_all('img'))))
|
||||||
|
blob += ' ' + ' '.join(el.get('style', '') for el in body.find_all(style=True))
|
||||||
|
blob += ' ' + str(body)
|
||||||
|
for m in BROAD_IMG_PAT.finditer(blob):
|
||||||
|
n = int(m.group(1))
|
||||||
|
if valid(n):
|
||||||
|
img_t.add(n)
|
||||||
|
return img_t, link_t
|
||||||
|
|
||||||
|
|
||||||
|
def decide_O(img_t, link_t):
|
||||||
|
if img_t:
|
||||||
|
return ','.join(f'{n}유형' for n in sorted(img_t)), (bool(link_t) and link_t != img_t)
|
||||||
|
if link_t:
|
||||||
|
if {1, 2, 3, 4}.issubset(link_t):
|
||||||
|
return '미부착', False
|
||||||
|
return ','.join(f'{n}유형' for n in sorted(link_t)), False
|
||||||
|
return '미부착', False
|
||||||
|
|
||||||
|
|
||||||
|
def domain3(host):
|
||||||
|
labels = (host or '').split('.')
|
||||||
|
return '.'.join(labels[-3:]) if len(labels) >= 3 else host
|
||||||
|
|
||||||
|
|
||||||
|
def same_site(a, b):
|
||||||
|
return domain3(urlparse(a).hostname) == domain3(urlparse(b).hostname)
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
xlsx, p234_path, domain = sys.argv[1], sys.argv[2], sys.argv[3]
|
||||||
|
if '--weak-ssl' in sys.argv:
|
||||||
|
SESSION.mount('https://', WeakSSLAdapter())
|
||||||
|
body_sel = None
|
||||||
|
if len(sys.argv) > 4 and sys.argv[4].strip() and not sys.argv[4].startswith('--'):
|
||||||
|
body_sel = [s.strip() for s in sys.argv[4].split(',') if s.strip()]
|
||||||
|
mod = load_module(p234_path)
|
||||||
|
LINK_PAT = mod.KOGL_LINK_PAT
|
||||||
|
|
||||||
|
# get_body 시그니처 자동 대응: 일괄형은 (soup, selectors), 개별형은 (soup)
|
||||||
|
needs_sel = len(inspect.signature(mod.get_body).parameters) >= 2
|
||||||
|
sel = body_sel or ['#txt', '#contents', 'main']
|
||||||
|
get_body = (lambda s: mod.get_body(s, sel)) if needs_sel else mod.get_body
|
||||||
|
print(f'get_body 인자 {2 if needs_sel else 1}개 | body_sel={sel if needs_sel else "(미사용)"}')
|
||||||
|
|
||||||
|
wb = openpyxl.load_workbook(xlsx)
|
||||||
|
ws = wb.active
|
||||||
|
|
||||||
|
targets = []
|
||||||
|
for r in range(3, ws.max_row + 1):
|
||||||
|
L = ws.cell(r, 12).value
|
||||||
|
url = ws.cell(r, 11).value
|
||||||
|
if L not in (None, '') :
|
||||||
|
continue # 기존 행 보존
|
||||||
|
if not (isinstance(url, str) and domain in url):
|
||||||
|
continue
|
||||||
|
targets.append((r, url))
|
||||||
|
print(f'대상 신규행: {len(targets)}개')
|
||||||
|
|
||||||
|
def work(t):
|
||||||
|
r, url = t
|
||||||
|
soup = fetch_soup(url)
|
||||||
|
if soup is None:
|
||||||
|
return r, {'note': '접근 실패'}
|
||||||
|
body = get_body(soup)
|
||||||
|
form, count = mod.detect_form(body)
|
||||||
|
has_img, has_vid, has_txt = mod.detect_media(body)
|
||||||
|
img_t, link_t = detect_split(body, LINK_PAT)
|
||||||
|
P_loc = '게시판' if img_t or link_t else ''
|
||||||
|
if form == '게시판':
|
||||||
|
for du in mod.extract_detail_urls(body, url, limit=5):
|
||||||
|
if not same_site(url, du):
|
||||||
|
continue
|
||||||
|
ds = fetch_soup(du, 10)
|
||||||
|
if ds is None:
|
||||||
|
continue
|
||||||
|
db = get_body(ds)
|
||||||
|
di, dv, dt = mod.detect_media(db)
|
||||||
|
has_img = has_img or di; has_vid = has_vid or dv; has_txt = has_txt or dt
|
||||||
|
dimg, dlink = detect_split(db, LINK_PAT)
|
||||||
|
if (dimg or dlink) and not P_loc:
|
||||||
|
P_loc = '게시물'
|
||||||
|
img_t |= dimg; link_t |= dlink
|
||||||
|
N = mod.n_string(has_txt, has_img, has_vid)
|
||||||
|
O, mismatch = decide_O(img_t, link_t)
|
||||||
|
P = '' if O == '미부착' else (P_loc or '게시판')
|
||||||
|
return r, {'L': form, 'M': count if form == '게시판' else 1,
|
||||||
|
'N': N, 'O': O, 'P': P, 'mismatch': mismatch}
|
||||||
|
|
||||||
|
t0 = time.time()
|
||||||
|
results = {}
|
||||||
|
done = 0
|
||||||
|
with ThreadPoolExecutor(max_workers=8) as ex:
|
||||||
|
futs = [ex.submit(work, t) for t in targets]
|
||||||
|
for fut in as_completed(futs):
|
||||||
|
r, res = fut.result()
|
||||||
|
results[r] = res
|
||||||
|
done += 1
|
||||||
|
if done % 40 == 0:
|
||||||
|
print(f' 진행 {done}/{len(targets)} ({time.time()-t0:.0f}s)')
|
||||||
|
print(f'크롤링 완료 ({time.time()-t0:.0f}s)')
|
||||||
|
|
||||||
|
fail = 0
|
||||||
|
for r, res in results.items():
|
||||||
|
if res.get('note'):
|
||||||
|
fail += 1
|
||||||
|
if not ws.cell(r, 19).value:
|
||||||
|
ws.cell(r, 19).value = res['note']
|
||||||
|
continue
|
||||||
|
ws.cell(r, 12).value = res['L']
|
||||||
|
ws.cell(r, 13).value = res['M']
|
||||||
|
ws.cell(r, 14).value = res['N']
|
||||||
|
ws.cell(r, 15).value = res['O']
|
||||||
|
if res['P']:
|
||||||
|
ws.cell(r, 16).value = res['P']
|
||||||
|
if res['mismatch']:
|
||||||
|
cur = (ws.cell(r, 19).value or '').strip()
|
||||||
|
ws.cell(r, 19).value = '링크주소 오기' if not cur else cur + ' / 링크주소 오기'
|
||||||
|
|
||||||
|
try:
|
||||||
|
wb.save(xlsx)
|
||||||
|
saved = xlsx
|
||||||
|
except PermissionError:
|
||||||
|
saved = xlsx.replace('.xlsx', '_LP.xlsx')
|
||||||
|
wb.save(saved)
|
||||||
|
print(f'!! 원본 잠김(Excel 열림). 대체 저장: {saved}')
|
||||||
|
|
||||||
|
from collections import Counter
|
||||||
|
Lc = Counter(res.get('L') for res in results.values())
|
||||||
|
Oc = Counter(res.get('O') for res in results.values())
|
||||||
|
print(f'저장: {xlsx}')
|
||||||
|
print('L 분포:', dict(Lc))
|
||||||
|
print('O 분포:', dict(Oc))
|
||||||
|
print(f'접근실패: {fail}')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
265
_스크립트/_tab_run.log
Normal file
265
_스크립트/_tab_run.log
Normal file
@ -0,0 +1,265 @@
|
|||||||
|
대상 기관: 39개 (모드=run)
|
||||||
|
|
||||||
|
━━━ 충청남도 금산군 (행~526, base=https://www.geumsan.go.kr, domain=geumsan.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\3.금산군\금산군.xlsx https://www.geumsan.go.kr geumsan.go.kr
|
||||||
|
→ 확장그룹 3개 / 신규행 9개
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\3.금산군\금산군.xlsx https://www.geumsan.go.kr geumsan.go.kr --write
|
||||||
|
$ utf8 _tab_phase234.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\3.금산군\금산군.xlsx _phase234.py geumsan.go.kr
|
||||||
|
대상 신규행: 13개
|
||||||
|
저장: D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\3.금산군\금산군.xlsx
|
||||||
|
L 분포: {None: 4, '페이지': 4, '게시판': 5}
|
||||||
|
O 분포: {None: 4, '미부착': 7, '4유형': 2}
|
||||||
|
접근실패: 4
|
||||||
|
|
||||||
|
━━━ 충청남도 논산시 (행~738, base=https://nonsan.go.kr, domain=nonsan.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\4.논산시\논산시.xlsx https://nonsan.go.kr nonsan.go.kr
|
||||||
|
→ 확장그룹 7개 / 신규행 22개
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\4.논산시\논산시.xlsx https://nonsan.go.kr nonsan.go.kr --write
|
||||||
|
$ utf8 _tab_phase234.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\4.논산시\논산시.xlsx _chungnam_phase234_all.py nonsan.go.kr #txt,#contents,main
|
||||||
|
대상 신규행: 22개
|
||||||
|
저장: D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\4.논산시\논산시.xlsx
|
||||||
|
L 분포: {'페이지': 19, '게시판': 3}
|
||||||
|
O 분포: {'미부착': 22}
|
||||||
|
접근실패: 0
|
||||||
|
|
||||||
|
━━━ 충청남도 당진시 (행~315, base=https://www.dangjin.go.kr, domain=dangjin.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\5.당진시\당진시.xlsx https://www.dangjin.go.kr dangjin.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
(신규행 없음 → Phase234 생략)
|
||||||
|
|
||||||
|
━━━ 충청남도 보령시 (행~573, base=https://www.brcn.go.kr, domain=brcn.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\6.보령시\보령시.xlsx https://www.brcn.go.kr brcn.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
(신규행 없음 → Phase234 생략)
|
||||||
|
|
||||||
|
━━━ 충청남도 부여군 (행~315, base=https://www.buyeo.go.kr, domain=buyeo.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\7.부여군\부여군.xlsx https://www.buyeo.go.kr buyeo.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
(신규행 없음 → Phase234 생략)
|
||||||
|
|
||||||
|
━━━ 충청남도 서산시 (행~349, base=https://www.seosan.go.kr, domain=seosan.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\8.서산시\서산시.xlsx https://www.seosan.go.kr seosan.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
(신규행 없음 → Phase234 생략)
|
||||||
|
|
||||||
|
━━━ 충청남도 서천군 (행~330, base=https://www.seocheon.go.kr, domain=seocheon.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\9.서천군\서천군.xlsx https://www.seocheon.go.kr seocheon.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
(신규행 없음 → Phase234 생략)
|
||||||
|
|
||||||
|
━━━ 충청남도 아산시 (행~315, base=https://www.asan.go.kr, domain=asan.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\10.아산시\아산시.xlsx https://www.asan.go.kr asan.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
(신규행 없음 → Phase234 생략)
|
||||||
|
|
||||||
|
━━━ 충청남도 예산군 (행~387, base=https://www.yesan.go.kr, domain=yesan.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\11.예산군\예산군.xlsx https://www.yesan.go.kr yesan.go.kr
|
||||||
|
→ 확장그룹 91개 / 신규행 233개
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\11.예산군\예산군.xlsx https://www.yesan.go.kr yesan.go.kr --write
|
||||||
|
$ utf8 _tab_phase234.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\11.예산군\예산군.xlsx _chungnam_phase234_all.py yesan.go.kr #txt,#contents,main
|
||||||
|
대상 신규행: 237개
|
||||||
|
저장: D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\11.예산군\예산군.xlsx
|
||||||
|
L 분포: {'페이지': 167, '게시판': 70}
|
||||||
|
O 분포: {'미부착': 237}
|
||||||
|
접근실패: 0
|
||||||
|
|
||||||
|
━━━ 충청남도 천안시 (행~371, base=https://www.cheonan.go.kr, domain=cheonan.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\12.천안시\천안시.xlsx https://www.cheonan.go.kr cheonan.go.kr
|
||||||
|
→ 확장그룹 107개 / 신규행 297개
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\12.천안시\천안시.xlsx https://www.cheonan.go.kr cheonan.go.kr --write
|
||||||
|
$ utf8 _tab_phase234.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\12.천안시\천안시.xlsx _chungnam_phase234_all.py cheonan.go.kr #txt,#contents,main
|
||||||
|
대상 신규행: 283개
|
||||||
|
저장: D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\12.천안시\천안시.xlsx
|
||||||
|
L 분포: {'페이지': 225, '게시판': 58}
|
||||||
|
O 분포: {'미부착': 281, '1유형': 2}
|
||||||
|
접근실패: 0
|
||||||
|
|
||||||
|
━━━ 충청남도 청양군 (행~315, base=http://www.cheongyang.go.kr, domain=cheongyang.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\13.청양군\청양군.xlsx http://www.cheongyang.go.kr cheongyang.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
(신규행 없음 → Phase234 생략)
|
||||||
|
|
||||||
|
━━━ 충청남도 태안군 (행~315, base=https://www.taean.go.kr, domain=taean.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\14.태안군\태안군.xlsx https://www.taean.go.kr taean.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
(신규행 없음 → Phase234 생략)
|
||||||
|
|
||||||
|
━━━ 충청남도 홍성군 (행~315, base=https://www.hongseong.go.kr, domain=hongseong.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\15.홍성군\홍성군.xlsx https://www.hongseong.go.kr hongseong.go.kr
|
||||||
|
→ 확장그룹 67개 / 신규행 215개
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\15.홍성군\홍성군.xlsx https://www.hongseong.go.kr hongseong.go.kr --write
|
||||||
|
$ utf8 _tab_phase234.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\15.홍성군\홍성군.xlsx _chungnam_phase234_all.py hongseong.go.kr #txt,#contents,main
|
||||||
|
대상 신규행: 213개
|
||||||
|
저장: D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청남도\15.홍성군\홍성군.xlsx
|
||||||
|
L 분포: {'게시판': 41, '페이지': 172}
|
||||||
|
O 분포: {'미부착': 212, '4유형': 1}
|
||||||
|
접근실패: 0
|
||||||
|
|
||||||
|
━━━ 충청북도 괴산군 (행~315, base=https://www.goesan.go.kr, domain=goesan.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\1.괴산군\괴산군.xlsx https://www.goesan.go.kr goesan.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
(신규행 없음 → Phase234 생략)
|
||||||
|
|
||||||
|
━━━ 충청북도 단양군 (행~486, base=https://www.danyang.go.kr, domain=danyang.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\2.단양군\단양군.xlsx https://www.danyang.go.kr danyang.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
(신규행 없음 → Phase234 생략)
|
||||||
|
|
||||||
|
━━━ 충청북도 영동군 [weak-ssl] (행~674, base=https://www.yd21.go.kr, domain=yd21.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\4.영동군\영동군.xlsx https://www.yd21.go.kr yd21.go.kr --weak-ssl
|
||||||
|
→ 확장그룹 20개 / 신규행 143개
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\4.영동군\영동군.xlsx https://www.yd21.go.kr yd21.go.kr --weak-ssl --write
|
||||||
|
$ utf8 _tab_phase234.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\4.영동군\영동군.xlsx _chungbuk_phase234_all.py yd21.go.kr #txt,#contents,main --weak-ssl
|
||||||
|
대상 신규행: 146개
|
||||||
|
저장: D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\4.영동군\영동군.xlsx
|
||||||
|
L 분포: {None: 2, '페이지': 138, '게시판': 6}
|
||||||
|
O 분포: {None: 2, '미부착': 144}
|
||||||
|
접근실패: 2
|
||||||
|
|
||||||
|
━━━ 충청북도 옥천군 (행~564, base=https://www.oc.go.kr, domain=oc.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\5.옥천군\옥천군.xlsx https://www.oc.go.kr oc.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
(신규행 없음 → Phase234 생략)
|
||||||
|
|
||||||
|
━━━ 충청북도 음성군 (행~652, base=https://www.eumseong.go.kr, domain=eumseong.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\6.음성군\음성군.xlsx https://www.eumseong.go.kr eumseong.go.kr
|
||||||
|
→ 확장그룹 9개 / 신규행 18개
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\6.음성군\음성군.xlsx https://www.eumseong.go.kr eumseong.go.kr --write
|
||||||
|
$ utf8 _tab_phase234.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\6.음성군\음성군.xlsx _chungbuk_phase234_all.py eumseong.go.kr #contents,#txt,main
|
||||||
|
대상 신규행: 18개
|
||||||
|
저장: D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\6.음성군\음성군.xlsx
|
||||||
|
L 분포: {'페이지': 9, '게시판': 9}
|
||||||
|
O 분포: {'미부착': 18}
|
||||||
|
접근실패: 0
|
||||||
|
|
||||||
|
━━━ 충청북도 제천시 (행~644, base=https://www.jecheon.go.kr, domain=jecheon.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\7.제천시\제천시.xlsx https://www.jecheon.go.kr jecheon.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
(신규행 없음 → Phase234 생략)
|
||||||
|
|
||||||
|
━━━ 충청북도 증평군 (행~315, base=https://www.jp.go.kr, domain=jp.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\8.증평군\증평군.xlsx https://www.jp.go.kr jp.go.kr
|
||||||
|
→ 확장그룹 42개 / 신규행 175개
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\8.증평군\증평군.xlsx https://www.jp.go.kr jp.go.kr --write
|
||||||
|
$ utf8 _tab_phase234.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\8.증평군\증평군.xlsx _chungbuk_phase234_all.py jp.go.kr #txt,#contents,main
|
||||||
|
대상 신규행: 160개
|
||||||
|
저장: D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\8.증평군\증평군.xlsx
|
||||||
|
L 분포: {'게시판': 158, None: 1, '페이지': 1}
|
||||||
|
O 분포: {'미부착': 159, None: 1}
|
||||||
|
접근실패: 1
|
||||||
|
|
||||||
|
━━━ 충청북도 진천군 (행~334, base=https://www.jincheon.go.kr, domain=jincheon.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\9.진천군\진천군.xlsx https://www.jincheon.go.kr jincheon.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
(신규행 없음 → Phase234 생략)
|
||||||
|
|
||||||
|
━━━ 충청북도 청주시 (행~342, base=https://www.cheongju.go.kr, domain=cheongju.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\10.청주시\청주시.xlsx https://www.cheongju.go.kr cheongju.go.kr
|
||||||
|
→ 확장그룹 5개 / 신규행 97개
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\10.청주시\청주시.xlsx https://www.cheongju.go.kr cheongju.go.kr --write
|
||||||
|
$ utf8 _tab_phase234.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\10.청주시\청주시.xlsx _chungbuk_phase234_all.py cheongju.go.kr #contents,#txt,main
|
||||||
|
대상 신규행: 90개
|
||||||
|
저장: D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\10.청주시\청주시.xlsx
|
||||||
|
L 분포: {None: 1, '페이지': 76, '게시판': 13}
|
||||||
|
O 분포: {None: 1, '미부착': 89}
|
||||||
|
접근실패: 1
|
||||||
|
|
||||||
|
━━━ 충청북도 충주시 (행~403, base=https://www.chungju.go.kr, domain=chungju.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\11.충주시\충주시.xlsx https://www.chungju.go.kr chungju.go.kr
|
||||||
|
→ 확장그룹 3개 / 신규행 33개
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\11.충주시\충주시.xlsx https://www.chungju.go.kr chungju.go.kr --write
|
||||||
|
$ utf8 _tab_phase234.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\11.충주시\충주시.xlsx _chungbuk_phase234_all.py chungju.go.kr #contents,#txt,main
|
||||||
|
대상 신규행: 36개
|
||||||
|
저장: D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\충청북도\11.충주시\충주시.xlsx
|
||||||
|
L 분포: {None: 2, '게시판': 30, '페이지': 4}
|
||||||
|
O 분포: {None: 2, '미부착': 34}
|
||||||
|
접근실패: 2
|
||||||
|
|
||||||
|
━━━ 전북특별자치도 고창군 (행~381, base=https://www.gochang.go.kr, domain=gochang.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\1.고창군\고창군.xlsx https://www.gochang.go.kr gochang.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
(신규행 없음 → Phase234 생략)
|
||||||
|
|
||||||
|
━━━ 전북특별자치도 군산시 (행~639, base=https://www.gunsan.go.kr, domain=gunsan.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\2.군산시\군산시.xlsx https://www.gunsan.go.kr gunsan.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
(신규행 없음 → Phase234 생략)
|
||||||
|
|
||||||
|
━━━ 전북특별자치도 김제시 (행~454, base=https://www.gimje.go.kr, domain=gimje.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\3.김제시\김제시.xlsx https://www.gimje.go.kr gimje.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
(신규행 없음 → Phase234 생략)
|
||||||
|
|
||||||
|
━━━ 전북특별자치도 남원시 (행~333, base=https://www.namwon.go.kr, domain=namwon.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\4.남원시\남원시.xlsx https://www.namwon.go.kr namwon.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
(신규행 없음 → Phase234 생략)
|
||||||
|
|
||||||
|
━━━ 전북특별자치도 무주군 (행~315, base=https://www.muju.go.kr, domain=muju.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\5.무주군\무주군.xlsx https://www.muju.go.kr muju.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
(신규행 없음 → Phase234 생략)
|
||||||
|
|
||||||
|
━━━ 전북특별자치도 부안군 (행~315, base=https://www.buan.go.kr, domain=buan.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\6.부안군\부안군.xlsx https://www.buan.go.kr buan.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
(신규행 없음 → Phase234 생략)
|
||||||
|
|
||||||
|
━━━ 전북특별자치도 순창군 (행~454, base=https://www.sunchang.go.kr, domain=sunchang.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\7.순창군\순창군.xlsx https://www.sunchang.go.kr sunchang.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
(신규행 없음 → Phase234 생략)
|
||||||
|
|
||||||
|
━━━ 전북특별자치도 완주군 (행~331, base=https://www.wanju.go.kr, domain=wanju.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\8.완주군\완주군.xlsx https://www.wanju.go.kr wanju.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
(신규행 없음 → Phase234 생략)
|
||||||
|
|
||||||
|
━━━ 전북특별자치도 익산시 (행~537, base=https://www.iksan.go.kr, domain=iksan.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\9.익산시\익산시.xlsx https://www.iksan.go.kr iksan.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
(신규행 없음 → Phase234 생략)
|
||||||
|
|
||||||
|
━━━ 전북특별자치도 임실군 (행~315, base=https://www.imsil.go.kr, domain=imsil.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\10.임실군\임실군.xlsx https://www.imsil.go.kr imsil.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
(신규행 없음 → Phase234 생략)
|
||||||
|
|
||||||
|
━━━ 전북특별자치도 장수군 (행~315, base=https://www.jangsu.go.kr, domain=jangsu.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\11.장수군\장수군.xlsx https://www.jangsu.go.kr jangsu.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
(신규행 없음 → Phase234 생략)
|
||||||
|
|
||||||
|
━━━ 전북특별자치도 전주시 (행~315, base=https://www.jeonju.go.kr, domain=jeonju.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\12.전주시\전주시.xlsx https://www.jeonju.go.kr jeonju.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
(신규행 없음 → Phase234 생략)
|
||||||
|
|
||||||
|
━━━ 전북특별자치도 정읍시 (행~315, base=https://www.jeongeup.go.kr, domain=jeongeup.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\13.정읍시\정읍시.xlsx https://www.jeongeup.go.kr jeongeup.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
(신규행 없음 → Phase234 생략)
|
||||||
|
|
||||||
|
━━━ 전북특별자치도 진안군 (행~315, base=http://www.jinan.go.kr, domain=jinan.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\전북특별자치도\14.진안군\진안군.xlsx http://www.jinan.go.kr jinan.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
(신규행 없음 → Phase234 생략)
|
||||||
|
|
||||||
|
━━━ 제주특별자치도 서귀포시 (행~366, base=https://www.seogwipo.go.kr, domain=seogwipo.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\제주특별자치도\1.서귀포시\서귀포시.xlsx https://www.seogwipo.go.kr seogwipo.go.kr
|
||||||
|
→ 확장그룹 5개 / 신규행 24개
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\제주특별자치도\1.서귀포시\서귀포시.xlsx https://www.seogwipo.go.kr seogwipo.go.kr --write
|
||||||
|
$ utf8 _tab_phase234.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\제주특별자치도\1.서귀포시\서귀포시.xlsx _jeonbuk_phase234_all.py seogwipo.go.kr #main-contents,#content,#contents,.contents,#txt,main,#container,#sub
|
||||||
|
대상 신규행: 22개
|
||||||
|
저장: D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\제주특별자치도\1.서귀포시\서귀포시.xlsx
|
||||||
|
L 분포: {'페이지': 10, '게시판': 11, None: 1}
|
||||||
|
O 분포: {'미부착': 21, None: 1}
|
||||||
|
접근실패: 1
|
||||||
|
|
||||||
|
━━━ 제주특별자치도 제주시 (행~315, base=https://www.jejusi.go.kr, domain=jejusi.go.kr)
|
||||||
|
$ utf8 _tab_expand.py D:\01.프로젝트\DB수집\작업파일\광역_사이트맵\제주특별자치도\2.제주시\제주시.xlsx https://www.jejusi.go.kr jejusi.go.kr
|
||||||
|
→ 확장그룹 0개 / 신규행 0개
|
||||||
|
(신규행 없음 → Phase234 생략)
|
||||||
|
|
||||||
|
|
||||||
|
=== 배치 종료 ===
|
||||||
116
_스크립트/_tab_scan.py
Normal file
116
_스크립트/_tab_scan.py
Normal file
@ -0,0 +1,116 @@
|
|||||||
|
"""사이트 본문 탭(서브내비) 탐지 스캐너.
|
||||||
|
|
||||||
|
각 사이트 페이지를 받아 본문 탭 UL을 찾아, 탭이 2개 이상인 페이지를 보고한다.
|
||||||
|
공주시: ul.tab-ul. 다른 사이트도 흔한 탭 클래스 패턴을 함께 탐지.
|
||||||
|
|
||||||
|
사용법: python -X utf8 _tab_scan.py <엑셀경로> <도메인키워드>
|
||||||
|
예: python -X utf8 _tab_scan.py 충청남도/2.공주시/충청남도_공주시.xlsx gongju.go.kr
|
||||||
|
"""
|
||||||
|
import re
|
||||||
|
import sys
|
||||||
|
import warnings
|
||||||
|
from concurrent.futures import ThreadPoolExecutor, as_completed
|
||||||
|
from urllib.parse import urljoin, urlparse
|
||||||
|
|
||||||
|
import openpyxl
|
||||||
|
import requests
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
|
||||||
|
# 흔한 본문 탭(서브내비) 컨테이너 클래스 패턴 (소문자 비교)
|
||||||
|
TAB_CLASS_PATS = [
|
||||||
|
'tab-ul', # 공주시
|
||||||
|
'tab_ul',
|
||||||
|
'tabmenu', 'tab-menu', 'tab_menu',
|
||||||
|
'tablist', 'tab-list', 'tab_list',
|
||||||
|
'subtab', 'sub-tab', 'sub_tab',
|
||||||
|
'tab_wrap', 'tab-wrap',
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def find_tab_uls(soup):
|
||||||
|
"""본문 탭으로 보이는 ul들을 반환. (ul, [(text,href),...]) 리스트."""
|
||||||
|
results = []
|
||||||
|
seen = set()
|
||||||
|
for ul in soup.find_all('ul'):
|
||||||
|
cls = ' '.join(ul.get('class') or []).lower()
|
||||||
|
if not cls:
|
||||||
|
# 부모 div의 클래스도 확인 (ul 자체엔 클래스 없고 div.tab > ul 구조)
|
||||||
|
parent = ul.find_parent(['div'])
|
||||||
|
pcls = ' '.join(parent.get('class') or []).lower() if parent else ''
|
||||||
|
if not any(p in pcls for p in TAB_CLASS_PATS):
|
||||||
|
continue
|
||||||
|
elif not any(p in cls for p in TAB_CLASS_PATS):
|
||||||
|
continue
|
||||||
|
links = []
|
||||||
|
for a in ul.find_all('a'):
|
||||||
|
t = a.get_text(strip=True).replace('\xa0', '').strip()
|
||||||
|
h = (a.get('href') or '').strip()
|
||||||
|
if t:
|
||||||
|
links.append((t, h))
|
||||||
|
if len(links) >= 2:
|
||||||
|
key = tuple(t for t, _ in links)
|
||||||
|
if key not in seen:
|
||||||
|
seen.add(key)
|
||||||
|
results.append((cls, links))
|
||||||
|
return results
|
||||||
|
|
||||||
|
|
||||||
|
def scan_one(r, url):
|
||||||
|
try:
|
||||||
|
resp = requests.get(url, headers=H, timeout=15, verify=False)
|
||||||
|
resp.encoding = resp.apparent_encoding
|
||||||
|
soup = BeautifulSoup(resp.text, 'html.parser')
|
||||||
|
tabs = find_tab_uls(soup)
|
||||||
|
return r, url, tabs, None
|
||||||
|
except Exception as e:
|
||||||
|
return r, url, [], str(e)[:60]
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
xlsx = sys.argv[1]
|
||||||
|
domain = sys.argv[2]
|
||||||
|
wb = openpyxl.load_workbook(xlsx)
|
||||||
|
ws = wb.active
|
||||||
|
targets = []
|
||||||
|
for r in range(3, ws.max_row + 1):
|
||||||
|
u = ws.cell(r, 11).value
|
||||||
|
if isinstance(u, str) and domain in u:
|
||||||
|
targets.append((r, u))
|
||||||
|
print(f'스캔 대상: {len(targets)}개 ({domain})')
|
||||||
|
|
||||||
|
found = []
|
||||||
|
errs = []
|
||||||
|
with ThreadPoolExecutor(max_workers=8) as ex:
|
||||||
|
futs = [ex.submit(scan_one, r, u) for r, u in targets]
|
||||||
|
for fut in as_completed(futs):
|
||||||
|
r, url, tabs, err = fut.result()
|
||||||
|
if err:
|
||||||
|
errs.append((r, url, err))
|
||||||
|
elif tabs:
|
||||||
|
found.append((r, url, tabs))
|
||||||
|
|
||||||
|
found.sort()
|
||||||
|
print(f'\n=== 탭 발견: {len(found)}개 페이지 ===')
|
||||||
|
for r, url, tabs in found:
|
||||||
|
e = ws.cell(r, 5).value or ''
|
||||||
|
f = ws.cell(r, 6).value or ''
|
||||||
|
g = ws.cell(r, 7).value or ''
|
||||||
|
print(f'[행{r}] E={e} F={f} G={g}')
|
||||||
|
print(f' {url}')
|
||||||
|
for cls, links in tabs:
|
||||||
|
print(f' <ul class={cls}> 탭 {len(links)}개:')
|
||||||
|
for t, h in links:
|
||||||
|
print(f' - {t} -> {h}')
|
||||||
|
if errs:
|
||||||
|
print(f'\n=== 에러 {len(errs)}개 ===')
|
||||||
|
for r, url, err in errs[:20]:
|
||||||
|
print(f'[행{r}] {url} : {err}')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
1976
_스크립트/_wanju_menu.json
Normal file
1976
_스크립트/_wanju_menu.json
Normal file
File diff suppressed because it is too large
Load Diff
76
_스크립트/_zz_emptyboard.py
Normal file
76
_스크립트/_zz_emptyboard.py
Normal file
@ -0,0 +1,76 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""억제된 페이지→게시판 후보(아직 페이지인 행)를 빈게시판 신호로 재분석. 읽기전용."""
|
||||||
|
import sys, os, re, json, importlib.util, time
|
||||||
|
from concurrent.futures import ThreadPoolExecutor, as_completed
|
||||||
|
import openpyxl
|
||||||
|
|
||||||
|
HERE = os.path.dirname(os.path.abspath(__file__)); TEMP = os.path.join(HERE,'..','_temp')
|
||||||
|
MODULES = ['_chungnam_phase234_all.py','_chungbuk_phase234_all.py','_jeonbuk_phase234_all.py']
|
||||||
|
def load():
|
||||||
|
sites,mods={},{}
|
||||||
|
for f in MODULES:
|
||||||
|
sp=importlib.util.spec_from_file_location(f[:-3],os.path.join(HERE,f));m=importlib.util.module_from_spec(sp);sp.loader.exec_module(m)
|
||||||
|
for k,v in m.SITES.items(): sites[k]=v;mods[k]=m
|
||||||
|
return sites,mods
|
||||||
|
EMPTY=re.compile(r'게시물이?\s*없|등록된?\s*(?:게시물|자료|글|내용)\s*가?\s*없|자료가\s*없|검색된\s*(?:게시물|결과)\s*가?\s*없|데이터가\s*없')
|
||||||
|
BOARDDOM='.board_list,.bbs_list,table.board_list,.board_view,.board_wrap,.bbs,.tbl_list,.board,.list_wrap,.gallery_list,.photo_list,table.board'
|
||||||
|
WRITE=re.compile(r'글쓰기|글등록|등록하기|write|글작성')
|
||||||
|
def mkfetch(M,cfg):
|
||||||
|
if hasattr(M,'make_session'):
|
||||||
|
s=M.make_session(weak_ssl=cfg.get('weak_ssl',False)); return lambda u:M.fetch(s,u)
|
||||||
|
return lambda u:M.fetch(u)
|
||||||
|
def analyze(M,url,bsel,fetch):
|
||||||
|
soup=None
|
||||||
|
for k in range(3):
|
||||||
|
soup=fetch(url)
|
||||||
|
if soup is not None: break
|
||||||
|
time.sleep(0.4*(k+1))
|
||||||
|
if soup is None: return {'v':'fail'}
|
||||||
|
b=M.get_body(soup,bsel)
|
||||||
|
txt=b.get_text(' ',strip=True)
|
||||||
|
search=len([i for i in b.find_all('input') if (i.get('type') or 'text').lower() in ('text','search')])>0
|
||||||
|
rows=[el for el in b.select('table tbody tr, .board_list li, ul.bbs_list li, .bbs_list li') if el.find('a',href=True)]
|
||||||
|
empty=bool(EMPTY.search(txt))
|
||||||
|
boarddom=bool(b.select(BOARDDOM))
|
||||||
|
paging=bool(b.select('.pagination,.paging,nav.paging,.page_nav,.paginate'))
|
||||||
|
write=bool(WRITE.search(txt))
|
||||||
|
if empty or (boarddom and (paging or write)):
|
||||||
|
v='empty_board'
|
||||||
|
else:
|
||||||
|
v='page'
|
||||||
|
return {'v':v,'search':search,'rows':len(rows),'empty':empty,'boarddom':boarddom,'paging':paging,'write':write,
|
||||||
|
'snip':txt[:160]}
|
||||||
|
def main():
|
||||||
|
sites,mods=load(); names=[a for a in sys.argv[1:] if not a.startswith('--')]
|
||||||
|
for n in names:
|
||||||
|
cfg=sites[n];M=mods[n];fetch=mkfetch(M,cfg)
|
||||||
|
wb=openpyxl.load_workbook(cfg['xlsx'],read_only=True);ws=wb.active
|
||||||
|
d=json.load(open(os.path.join(TEMP,f'_audit_{n}.json'),encoding='utf-8'))
|
||||||
|
cands=[]
|
||||||
|
for x in d['L_diffs']:
|
||||||
|
if x['oldL']=='페이지' and x['newL']=='게시판':
|
||||||
|
if ws.cell(x['r'],12).value=='페이지': # 아직 페이지(억제됨)
|
||||||
|
cands.append(x)
|
||||||
|
wb.close()
|
||||||
|
res=[]
|
||||||
|
with ThreadPoolExecutor(max_workers=6) as ex:
|
||||||
|
futs={ex.submit(analyze,M,x['url'],cfg['body_sel'],fetch):x for x in cands}
|
||||||
|
for f in as_completed(futs):
|
||||||
|
x=futs[f]
|
||||||
|
try:r=f.result()
|
||||||
|
except:r={'v':'fail'}
|
||||||
|
res.append((x,r))
|
||||||
|
eb=[(x,r) for x,r in res if r['v']=='empty_board']
|
||||||
|
pg=[(x,r) for x,r in res if r['v']=='page']
|
||||||
|
nf=len([1 for _,r in res if r['v']=='fail'])
|
||||||
|
print('\n==== %s: 억제후보%d → 빈게시판의심%d 페이지%d 실패%d'%(n,len(cands),len(eb),len(pg),nf))
|
||||||
|
print('--- 빈게시판 의심(empty_board) 예시 ---')
|
||||||
|
for x,r in sorted(eb,key=lambda t:t[0]['r'])[:8]:
|
||||||
|
print(f" r{x['r']} {(x['label'] or '')[:20]} | empty={r['empty']} boarddom={r['boarddom']} paging={r['paging']} | {x['url'][:62]}")
|
||||||
|
print(f" · {r['snip'][:110]}")
|
||||||
|
print('--- 순수 페이지 예시 ---')
|
||||||
|
for x,r in sorted(pg,key=lambda t:t[0]['r'])[:5]:
|
||||||
|
print(f" r{x['r']} {(x['label'] or '')[:20]} | boarddom={r['boarddom']} | {x['url'][:62]}")
|
||||||
|
print(f" · {r['snip'][:110]}")
|
||||||
|
json.dump([{**x,**r} for x,r in res],open(os.path.join(TEMP,f'_emptyboard_{n}.json'),'w',encoding='utf-8'),ensure_ascii=False,indent=1)
|
||||||
|
if __name__=='__main__': main()
|
||||||
37
_스크립트/_공공기관2_napply.py
Normal file
37
_스크립트/_공공기관2_napply.py
Normal file
@ -0,0 +1,37 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""plan_{기관}.json 의 auto='어문'(콘텐츠이미지 없음) 행 → N에서 '이미지' 제거(강등).
|
||||||
|
candidate 행은 보존(사용자/추가 시각판정 영역). 하늘색 표시는 하지 않음(사용자 지시).
|
||||||
|
사용: python _공공기관_napply.py <기관명>
|
||||||
|
"""
|
||||||
|
import sys, os, json
|
||||||
|
import openpyxl
|
||||||
|
|
||||||
|
OUTDIR = r'D:\01.프로젝트\DB수집\작업파일\공공기관2'
|
||||||
|
TEMP = r'D:\01.프로젝트\DB수집\_temp'
|
||||||
|
|
||||||
|
|
||||||
|
def demote(N):
|
||||||
|
parts = [p for p in (N or '').split(',') if p and p != '이미지']
|
||||||
|
return ','.join(parts) if parts else '없음'
|
||||||
|
|
||||||
|
|
||||||
|
def run(name):
|
||||||
|
plan = json.load(open(os.path.join(TEMP, f'nshot_{name}', f'plan_{name}.json'), encoding='utf-8'))
|
||||||
|
xlsx = __import__('glob').glob(os.path.join(OUTDIR, f'*.{name}', f'{name}.xlsx'))[0]
|
||||||
|
wb = openpyxl.load_workbook(xlsx)
|
||||||
|
ws = wb.active
|
||||||
|
chg = 0
|
||||||
|
for it in plan:
|
||||||
|
if it['auto'] == '어문':
|
||||||
|
r = it['row']
|
||||||
|
cur = ws.cell(r, 14).value or ''
|
||||||
|
if '이미지' in cur:
|
||||||
|
ws.cell(r, 14).value = demote(cur)
|
||||||
|
chg += 1
|
||||||
|
wb.save(xlsx)
|
||||||
|
cand = sum(1 for it in plan if it['auto'] == 'candidate')
|
||||||
|
print(f'[{name}] 이미지강등 {chg}행 | 잔여이미지후보 {cand}행(시각판정 영역)')
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
run(sys.argv[1])
|
||||||
34
_스크립트/_공공기관2_nbatch.py
Normal file
34
_스크립트/_공공기관2_nbatch.py
Normal file
@ -0,0 +1,34 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""공공기관 N 이미지 몽타주 재판정 일괄: 각 기관 렌더+크기필터(nshot) → 이미지강등 적용(napply).
|
||||||
|
국토연구원은 이미 처리 → 스킵. 사용: python _공공기관_nbatch.py [기관명 ...]
|
||||||
|
"""
|
||||||
|
import sys, os, json, importlib.util
|
||||||
|
|
||||||
|
|
||||||
|
def load(mod):
|
||||||
|
spec = importlib.util.spec_from_file_location(mod, rf'D:\01.프로젝트\DB수집\_스크립트\{mod}.py')
|
||||||
|
m = importlib.util.module_from_spec(spec); spec.loader.exec_module(m)
|
||||||
|
return m
|
||||||
|
|
||||||
|
|
||||||
|
nshot = load('_공공기관2_nshot')
|
||||||
|
napply = load('_공공기관2_napply')
|
||||||
|
|
||||||
|
probe = {r['name']: r for r in json.load(open(r'D:\01.프로젝트\DB수집\_스크립트\_공공기관2_probe.json', encoding='utf-8'))}
|
||||||
|
order = sorted(probe.values(), key=lambda x: -int(x['num']))
|
||||||
|
only = sys.argv[1:]
|
||||||
|
if only:
|
||||||
|
order = [p for p in order if p['name'] in only or str(p['num']) in only]
|
||||||
|
|
||||||
|
DONE = {'국토연구원'}
|
||||||
|
for p in order:
|
||||||
|
name = p['name']
|
||||||
|
if name in DONE:
|
||||||
|
print(f'[{name}] 이미 처리 — 스킵'); continue
|
||||||
|
try:
|
||||||
|
nshot.run(name)
|
||||||
|
napply.run(name)
|
||||||
|
except Exception as e:
|
||||||
|
import traceback
|
||||||
|
print(f'[{name}] 실패: {e}'); traceback.print_exc()
|
||||||
|
print('\n=== N이미지 몽타주 재판정 일괄 완료 ===')
|
||||||
123
_스크립트/_공공기관2_nshot.py
Normal file
123
_스크립트/_공공기관2_nshot.py
Normal file
@ -0,0 +1,123 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""공공기관 N 이미지 몽타주 재판정 — 1단계: 렌더+콘텐츠이미지 크기필터 + 풀페이지 스샷 + 몽타주.
|
||||||
|
|
||||||
|
대상: N에 '이미지' 포함 & L!=사이트 & URL 있는 행.
|
||||||
|
출력: _temp\nshot_{기관}\ (스샷 png + montage_*.png) + plan_{기관}.json
|
||||||
|
plan: {row, url, L, N, n_imgs(필터통과), auto='어문'(이미지없음) | 'candidate'(몽타주판정)}
|
||||||
|
사용: python _공공기관_nshot.py <기관명>
|
||||||
|
"""
|
||||||
|
import sys, os, re, json, warnings
|
||||||
|
import openpyxl
|
||||||
|
from PIL import Image, ImageDraw, ImageFont
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
OUTDIR = r'D:\01.프로젝트\DB수집\작업파일\공공기관2'
|
||||||
|
TEMP = r'D:\01.프로젝트\DB수집\_temp'
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
NOISE = re.compile(r'(ico[_\-/]|/icon|logo|btn|bul[_\-]|bg[_\-]|banner|sns|blank|spacer|loading|arrow|/dot|line[_\.]|top_|foot|header|common|_icon|symbol|copyright|qr_|no_img|noimage|share|facebook|insta|twitter|naver|kakao|/skin/|/template/|/resources/|/images/common)', re.I)
|
||||||
|
|
||||||
|
JS_IMGS = """() => {
|
||||||
|
const out = [];
|
||||||
|
for (const i of document.querySelectorAll('img')) {
|
||||||
|
const r = i.getBoundingClientRect();
|
||||||
|
out.push({src: i.currentSrc||i.src||'', nw: i.naturalWidth, nh: i.naturalHeight, rw: Math.round(r.width), rh: Math.round(r.height)});
|
||||||
|
}
|
||||||
|
// 배경이미지도 일부
|
||||||
|
return out;
|
||||||
|
}"""
|
||||||
|
|
||||||
|
|
||||||
|
def content_imgs(items):
|
||||||
|
cnt = 0
|
||||||
|
for it in items:
|
||||||
|
src = it.get('src', '')
|
||||||
|
if not src or NOISE.search(src):
|
||||||
|
continue
|
||||||
|
nw, nh, rw, rh = it.get('nw', 0), it.get('nh', 0), it.get('rw', 0), it.get('rh', 0)
|
||||||
|
if nw >= 170 and nh >= 110 and rw >= 100 and rh >= 80:
|
||||||
|
cnt += 1
|
||||||
|
return cnt
|
||||||
|
|
||||||
|
|
||||||
|
def run(name):
|
||||||
|
from playwright.sync_api import sync_playwright
|
||||||
|
xlsx = __import__('glob').glob(os.path.join(OUTDIR, f'*.{name}', f'{name}.xlsx'))[0]
|
||||||
|
wb = openpyxl.load_workbook(xlsx)
|
||||||
|
ws = wb.active
|
||||||
|
targets = []
|
||||||
|
for r in range(3, ws.max_row + 1):
|
||||||
|
if ws.cell(r, 2).value is None:
|
||||||
|
break
|
||||||
|
L = ws.cell(r, 12).value
|
||||||
|
N = ws.cell(r, 14).value or ''
|
||||||
|
url = ws.cell(r, 11).value
|
||||||
|
if L != '사이트' and '이미지' in N and url and isinstance(url, str) and url.startswith('http'):
|
||||||
|
targets.append((r, url, L, N))
|
||||||
|
outdir = os.path.join(TEMP, f'nshot_{name}')
|
||||||
|
os.makedirs(outdir, exist_ok=True)
|
||||||
|
print(f'[{name}] 이미지후보 {len(targets)}행 렌더…')
|
||||||
|
plan = []
|
||||||
|
shots = [] # (row, N, path)
|
||||||
|
with sync_playwright() as p:
|
||||||
|
b = p.chromium.launch()
|
||||||
|
pg = b.new_page(user_agent=UA, viewport={'width': 1280, 'height': 1600})
|
||||||
|
for idx, (r, url, L, N) in enumerate(targets):
|
||||||
|
try:
|
||||||
|
pg.goto(url, timeout=25000, wait_until='domcontentloaded')
|
||||||
|
pg.wait_for_timeout(1800)
|
||||||
|
items = pg.evaluate(JS_IMGS)
|
||||||
|
nc = content_imgs(items)
|
||||||
|
if nc == 0:
|
||||||
|
plan.append({'row': r, 'url': url, 'L': L, 'N': N, 'n_imgs': 0, 'auto': '어문'})
|
||||||
|
else:
|
||||||
|
sp = os.path.join(outdir, f'r{r}.png')
|
||||||
|
try:
|
||||||
|
pg.screenshot(path=sp, full_page=False)
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
plan.append({'row': r, 'url': url, 'L': L, 'N': N, 'n_imgs': nc, 'auto': 'candidate'})
|
||||||
|
if os.path.exists(sp):
|
||||||
|
shots.append((r, N, sp))
|
||||||
|
except Exception as e:
|
||||||
|
plan.append({'row': r, 'url': url, 'L': L, 'N': N, 'n_imgs': -1, 'auto': '어문', 'err': str(e)[:40]})
|
||||||
|
if (idx + 1) % 20 == 0:
|
||||||
|
print(f' {idx+1}/{len(targets)}')
|
||||||
|
b.close()
|
||||||
|
# 몽타주 그리드 (후보만)
|
||||||
|
cols, cw, ch = 4, 300, 360
|
||||||
|
try:
|
||||||
|
font = ImageFont.truetype('malgun.ttf', 16)
|
||||||
|
except Exception:
|
||||||
|
font = ImageFont.load_default()
|
||||||
|
mont_paths = []
|
||||||
|
per = cols * 5 # 20개/장
|
||||||
|
for gi in range(0, len(shots), per):
|
||||||
|
chunk = shots[gi:gi + per]
|
||||||
|
rows_n = (len(chunk) + cols - 1) // cols
|
||||||
|
canvas = Image.new('RGB', (cols * cw, rows_n * ch), 'white')
|
||||||
|
d = ImageDraw.Draw(canvas)
|
||||||
|
for j, (r, N, sp) in enumerate(chunk):
|
||||||
|
try:
|
||||||
|
im = Image.open(sp).convert('RGB')
|
||||||
|
im.thumbnail((cw - 8, ch - 26))
|
||||||
|
except Exception:
|
||||||
|
continue
|
||||||
|
cx, cy = (j % cols) * cw, (j // cols) * ch
|
||||||
|
canvas.paste(im, (cx + 4, cy + 22))
|
||||||
|
d.rectangle([cx, cy, cx + cw - 1, cy + ch - 1], outline='gray')
|
||||||
|
d.text((cx + 4, cy + 3), f'r{r}', fill='red', font=font)
|
||||||
|
mp = os.path.join(outdir, f'montage_{name}_{gi//per+1}.png')
|
||||||
|
canvas.save(mp)
|
||||||
|
mont_paths.append(mp)
|
||||||
|
auto_amun = sum(1 for x in plan if x['auto'] == '어문')
|
||||||
|
cand = sum(1 for x in plan if x['auto'] == 'candidate')
|
||||||
|
with open(os.path.join(outdir, f'plan_{name}.json'), 'w', encoding='utf-8') as f:
|
||||||
|
json.dump(plan, f, ensure_ascii=False, indent=1)
|
||||||
|
print(f'[{name}] 자동강등(이미지無) {auto_amun} | 몽타주후보 {cand} | 몽타주 {len(mont_paths)}장')
|
||||||
|
for mp in mont_paths:
|
||||||
|
print(' MONT:', mp)
|
||||||
|
print(' PLAN:', os.path.join(outdir, f'plan_{name}.json'))
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
run(sys.argv[1])
|
||||||
393
_스크립트/_공공기관2_phase1.py
Normal file
393
_스크립트/_공공기관2_phase1.py
Normal file
@ -0,0 +1,393 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""공공기관 Phase 1: 사이트맵/메가메뉴 → 메뉴트리 D~K (Playwright 렌더 + 범용 중첩리스트 추출).
|
||||||
|
|
||||||
|
출력: D:\\01.프로젝트\\DB수집\\공공기관\\{기관}.xlsx (평면배치)
|
||||||
|
시트명: {번호}_{기관}
|
||||||
|
사용: python _공공기관_phase1.py [기관명 ...] (인자 없으면 OVERRIDES 정의된 전체)
|
||||||
|
"""
|
||||||
|
import sys, io, os, re, shutil, json, warnings
|
||||||
|
if __name__ == '__main__':
|
||||||
|
sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding='utf-8')
|
||||||
|
from copy import copy
|
||||||
|
from urllib.parse import urljoin, urlparse
|
||||||
|
import openpyxl
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
from openpyxl.styles import Alignment, Font
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
TEMPLATE = r'D:\01.프로젝트\DB수집\자료_취합_예시.xlsx'
|
||||||
|
OUTDIR = r'D:\01.프로젝트\DB수집\작업파일\공공기관2'
|
||||||
|
PROBE = r'D:\01.프로젝트\DB수집\_스크립트\_공공기관2_probe.json'
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
|
||||||
|
# 기관별 오버라이드: sitemap URL + 컨테이너 셀렉터(선택) + 렌더 대기(wait). probe json을 기본값으로.
|
||||||
|
OVERRIDES = {
|
||||||
|
'국립생태원': {'sitemap': 'https://www.nie.re.kr/nie/main/main.do', 'wait': 5000},
|
||||||
|
'국토안전관리원': {'use_home': True, 'sel': '.all-menu', 'wait': 4500},
|
||||||
|
'국립중앙의료원': {'use_home': True, 'sel': '#siteMap', 'wait': 4500},
|
||||||
|
'건설근로자공제회': {'use_home': True, 'wait': 4500},
|
||||||
|
'국민연금공단': {'sitemap': 'https://www.nps.or.kr/main.do', 'sel': '.allmenu-container', 'wait': 4000},
|
||||||
|
'국민건강보험공단': {'sitemap': 'https://www.nhis.or.kr/nhis/index.do', 'sel': 'nav.head-gnb', 'wait': 4000},
|
||||||
|
}
|
||||||
|
|
||||||
|
JS_PATS = [
|
||||||
|
# goMenuPage('MENUID','/realpath',...) 형: 2번째 인자가 실제 URL
|
||||||
|
re.compile(r"""go\w*(?:Menu|Page|Move)\w*\s*\(\s*['"][^'"]*['"]\s*,\s*['"]([^'"]+)['"]""", re.I),
|
||||||
|
re.compile(r"""go(?:Menu|SubMenu|Page|Link|Url|View|Move)\s*\(\s*['"]([^'"]+)['"]""", re.I),
|
||||||
|
re.compile(r"""(?:location\.href|window\.location(?:\.href)?)\s*=\s*(?:encodeURI\()?\s*['"]([^'"]+)['"]"""),
|
||||||
|
re.compile(r"""window\.open\s*\(\s*['"]([^'"]+)['"]"""),
|
||||||
|
re.compile(r"""(?:fn_?\w*|aLink|movePage|menuMove)\s*\(\s*['"]([^'"]+)['"]""", re.I),
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def js_href(a):
|
||||||
|
s = (a.get('href', '') or '') + ' ' + (a.get('onclick', '') or '')
|
||||||
|
for p in JS_PATS:
|
||||||
|
mm = p.search(s)
|
||||||
|
if mm:
|
||||||
|
u = mm.group(1).strip()
|
||||||
|
if u.startswith(('/', 'http', './', '?', '../')):
|
||||||
|
return u
|
||||||
|
return ''
|
||||||
|
|
||||||
|
|
||||||
|
def clean(s):
|
||||||
|
return re.sub(r'\s+', ' ', (s or '')).strip().replace('\xa0', '')
|
||||||
|
|
||||||
|
|
||||||
|
def render(url, wait=2500):
|
||||||
|
from playwright.sync_api import sync_playwright
|
||||||
|
with sync_playwright() as p:
|
||||||
|
b = p.chromium.launch()
|
||||||
|
pg = b.new_page(user_agent=UA, viewport={'width': 1440, 'height': 2400})
|
||||||
|
try:
|
||||||
|
pg.goto(url, timeout=30000, wait_until='domcontentloaded')
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
pg.wait_for_timeout(wait)
|
||||||
|
html = pg.content()
|
||||||
|
final = pg.url
|
||||||
|
b.close()
|
||||||
|
return final, html
|
||||||
|
|
||||||
|
|
||||||
|
SITEMAP_HINTS = ('sitemap', 'site_map', 'sitemapwrap', 'site-map', 'allmenu', 'all_menu', 'all-menu',
|
||||||
|
'menu_all', 'menuall', 'totalmenu', 'total_menu', 'totmenu', 'full_menu', 'fullmenu',
|
||||||
|
'gnb_all', 'gnball')
|
||||||
|
|
||||||
|
|
||||||
|
def pick_container(soup, prefer=None):
|
||||||
|
if prefer:
|
||||||
|
el = soup.select_one(prefer)
|
||||||
|
if el and len(el.find_all('a')) >= 8:
|
||||||
|
return el
|
||||||
|
cands = []
|
||||||
|
for d in soup.find_all(['div', 'section', 'main', 'nav', 'ul']):
|
||||||
|
cls = ' '.join(d.get('class', [])).lower() + ' ' + (d.get('id', '') or '').lower()
|
||||||
|
clsn = cls.replace(' ', '').replace('-', '').replace('_', '')
|
||||||
|
if any(x in cls for x in ['footer', 'aside']):
|
||||||
|
continue
|
||||||
|
ac = len(d.find_all('a'))
|
||||||
|
ul = len(d.find_all('ul'))
|
||||||
|
if ac < 12:
|
||||||
|
continue
|
||||||
|
score = ac + ul * 2
|
||||||
|
# 사이트맵/전체메뉴 컨테이너 강력 우대
|
||||||
|
if any(h in clsn for h in SITEMAP_HINTS):
|
||||||
|
score += 1000
|
||||||
|
# 전역 헤더/상단 gnb는 (사이트맵 페이지에선) 감점
|
||||||
|
if 'header' in cls or ('gnb' in clsn and not any(h in clsn for h in SITEMAP_HINTS)):
|
||||||
|
score -= 200
|
||||||
|
cands.append((score, ac, d))
|
||||||
|
if not cands:
|
||||||
|
return None
|
||||||
|
cands.sort(key=lambda x: x[0], reverse=True)
|
||||||
|
return cands[0][2]
|
||||||
|
|
||||||
|
|
||||||
|
NOISE_LABELS = {'메뉴없음', '메뉴 없음', '홈', 'home', 'home으로', '홈으로', 'eng', 'english', '로그인',
|
||||||
|
'login', '회원가입', '검색', 'search', '바로가기', '본문바로가기', '닫기', 'close', '전체메뉴',
|
||||||
|
'사이트맵', 'sitemap'}
|
||||||
|
|
||||||
|
|
||||||
|
def node_depth(el, container):
|
||||||
|
"""container 내부에서 el 위의 ul/ol/dl 조상 개수 = 중첩 깊이."""
|
||||||
|
d = 0
|
||||||
|
p = el.parent
|
||||||
|
while p is not None and p is not container:
|
||||||
|
if getattr(p, 'name', None) in ('ul', 'ol', 'dl'):
|
||||||
|
d += 1
|
||||||
|
p = p.parent
|
||||||
|
return d
|
||||||
|
|
||||||
|
|
||||||
|
def parse_table_sitemap(table):
|
||||||
|
"""테이블형 사이트맵: tr > th(대분류, 빈칸=이어받기) + th(중분류 a) + td(소분류 a들)."""
|
||||||
|
rows = []
|
||||||
|
curD = ''
|
||||||
|
for tr in table.find_all('tr'):
|
||||||
|
ths = tr.find_all('th', recursive=False)
|
||||||
|
if ths:
|
||||||
|
d_txt = clean(ths[0].get_text())
|
||||||
|
# 첫 th에 a가 없고 span/텍스트면 대분류 라벨(빈칸이면 이어받기)
|
||||||
|
if d_txt and not (len(ths) == 1 and ths[0].find('a')):
|
||||||
|
# 첫 th가 곧 중분류 a 단독인 경우는 제외 위해 a 유무 확인
|
||||||
|
if not ths[0].find('a') or len(ths) >= 2:
|
||||||
|
if not ths[0].find('a'):
|
||||||
|
curD = d_txt
|
||||||
|
E = ''; Eh = ''
|
||||||
|
if len(ths) >= 2:
|
||||||
|
ea = ths[1].find('a')
|
||||||
|
if ea:
|
||||||
|
E = clean(ea.get_text()); Eh = js_href(ea) or (ea.get('href') or '')
|
||||||
|
elif len(ths) == 1 and ths[0].find('a'):
|
||||||
|
ea = ths[0].find('a'); E = clean(ea.get_text()); Eh = js_href(ea) or (ea.get('href') or '')
|
||||||
|
td = tr.find('td')
|
||||||
|
leaves = td.find_all('a') if td else []
|
||||||
|
if leaves:
|
||||||
|
for la in leaves:
|
||||||
|
lab = clean(la.get_text())
|
||||||
|
if not lab:
|
||||||
|
continue
|
||||||
|
rows.append({'path': [curD, E, lab], 'href': js_href(la) or (la.get('href') or '')})
|
||||||
|
elif E:
|
||||||
|
rows.append({'path': [curD, E], 'href': Eh})
|
||||||
|
return [r for r in rows if r['path'] and any(r['path'])]
|
||||||
|
|
||||||
|
|
||||||
|
def extract_rows(container):
|
||||||
|
"""앵커 중심 깊이추출: 각 a/heading의 ul조상 수로 컬럼 결정(정규화). 범용. 테이블형 자동전환."""
|
||||||
|
tbls = container.find_all('table')
|
||||||
|
if tbls and sum(len(t.find_all('th')) for t in tbls) >= 4 and len(container.find_all('ul')) <= 1:
|
||||||
|
rows = []
|
||||||
|
for t in tbls:
|
||||||
|
rows += parse_table_sitemap(t)
|
||||||
|
if rows:
|
||||||
|
return rows
|
||||||
|
nodes = [] # [depth, label, href]
|
||||||
|
for el in container.find_all(['a', 'button', 'h2', 'h3', 'h4', 'h5', 'h6', 'dt']):
|
||||||
|
name = el.name
|
||||||
|
if name == 'a':
|
||||||
|
href = (el.get('href') or '').strip()
|
||||||
|
if href.startswith('#') or href.lower().startswith('javascript:') or not href:
|
||||||
|
href = js_href(el)
|
||||||
|
label = clean(el.get_text())
|
||||||
|
elif name == 'button':
|
||||||
|
cls = ' '.join(el.get('class', [])).lower()
|
||||||
|
if not (el.get('aria-expanded') is not None or el.get('aria-haspopup')
|
||||||
|
or any(x in cls for x in ('depth', 'trigger', 'gnb', '1d', '2d', '3d', 'menu'))):
|
||||||
|
continue
|
||||||
|
href = ''
|
||||||
|
label = clean(el.get('data-dir') or el.get('title') or el.get_text())
|
||||||
|
else:
|
||||||
|
if el.find('a'): # heading 안의 a는 따로 잡힘 → 중복방지
|
||||||
|
continue
|
||||||
|
href = ''
|
||||||
|
label = clean(el.get_text())
|
||||||
|
if not label or label.lower() in NOISE_LABELS or len(label) > 60:
|
||||||
|
continue
|
||||||
|
nodes.append([node_depth(el, container), label, href])
|
||||||
|
# 연속 중복 제거
|
||||||
|
dedup = []
|
||||||
|
for n in nodes:
|
||||||
|
if dedup and dedup[-1] == n:
|
||||||
|
continue
|
||||||
|
dedup.append(n)
|
||||||
|
nodes = dedup
|
||||||
|
if not nodes:
|
||||||
|
return []
|
||||||
|
mind = min(n[0] for n in nodes)
|
||||||
|
cols = [min(n[0] - mind, 6) for n in nodes]
|
||||||
|
out = []
|
||||||
|
path = [''] * 7
|
||||||
|
for i, (depth, label, href) in enumerate(nodes):
|
||||||
|
c = cols[i]
|
||||||
|
path[c] = label
|
||||||
|
for k in range(c + 1, 7):
|
||||||
|
path[k] = ''
|
||||||
|
is_leaf = (i == len(nodes) - 1) or (cols[i + 1] <= c)
|
||||||
|
if href or is_leaf:
|
||||||
|
out.append({'path': path[:c + 1], 'href': href})
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def rows_to_dicts(raw):
|
||||||
|
out = []
|
||||||
|
for it in raw:
|
||||||
|
p = it['path']
|
||||||
|
href = it.get('href', '')
|
||||||
|
if href.startswith('#') or href.lower().startswith('javascript:'):
|
||||||
|
href = ''
|
||||||
|
row = {'D': '', 'E': '', 'F': '', 'G': '', 'H': '', 'I': '', 'J': '', 'href': href}
|
||||||
|
for i, lab in enumerate(p[:7]):
|
||||||
|
row['DEFGHIJ'[i]] = lab
|
||||||
|
out.append(row)
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def write_excel(name, num, base, raw_rows):
|
||||||
|
domain = urlparse(base).netloc
|
||||||
|
output = os.path.join(OUTDIR, f'{num}.{name}', f'{name}.xlsx')
|
||||||
|
|
||||||
|
def abs_url(href):
|
||||||
|
if not href:
|
||||||
|
return ''
|
||||||
|
href = href.strip()
|
||||||
|
if href.startswith(('javascript:', '#')):
|
||||||
|
return ''
|
||||||
|
if href.startswith(('http://', 'https://')):
|
||||||
|
return href
|
||||||
|
return urljoin(base + '/', href)
|
||||||
|
|
||||||
|
def is_external(url):
|
||||||
|
return url.startswith(('http://', 'https://')) and domain not in url
|
||||||
|
|
||||||
|
# 부모-자식 URL 중복 제거 (전 깊이) — 부모행 D~lc 동일 + lc+1 채워짐 + K동일
|
||||||
|
def colval(r, c):
|
||||||
|
return r.get(c, '') or ''
|
||||||
|
final = []
|
||||||
|
i = 0
|
||||||
|
removed = 0
|
||||||
|
cols = 'DEFGHIJ'
|
||||||
|
while i < len(raw_rows):
|
||||||
|
row = raw_rows[i]
|
||||||
|
# 마지막 채워진 컬럼 찾기
|
||||||
|
lc = -1
|
||||||
|
for ci, c in enumerate(cols):
|
||||||
|
if colval(row, c) != '':
|
||||||
|
lc = ci
|
||||||
|
dup = False
|
||||||
|
if i + 1 < len(raw_rows) and lc >= 0 and lc < 6:
|
||||||
|
nxt = raw_rows[i + 1]
|
||||||
|
same = all(colval(row, cols[k]) == colval(nxt, cols[k]) for k in range(lc + 1))
|
||||||
|
if (same and colval(nxt, cols[lc + 1]) != '' and colval(row, cols[lc + 1]) == ''
|
||||||
|
and row.get('href', '') == nxt.get('href', '')):
|
||||||
|
dup = True
|
||||||
|
if dup:
|
||||||
|
removed += 1
|
||||||
|
i += 1
|
||||||
|
continue
|
||||||
|
final.append(row)
|
||||||
|
i += 1
|
||||||
|
|
||||||
|
if not final:
|
||||||
|
print(f' [{name}] 행 0개 — 스킵')
|
||||||
|
return 0
|
||||||
|
shutil.copy(TEMPLATE, output)
|
||||||
|
wb = openpyxl.load_workbook(output)
|
||||||
|
ws = wb.active
|
||||||
|
ws.title = f'{num:02d}_{name}'
|
||||||
|
HEADER = {'B1:R1', 'S1:W1', 'Y1:AA1'}
|
||||||
|
for rng in [str(m) for m in ws.merged_cells.ranges if str(m) not in HEADER]:
|
||||||
|
ws.unmerge_cells(rng)
|
||||||
|
for row in ws.iter_rows(min_row=3, max_row=ws.max_row, min_col=1, max_col=ws.max_column):
|
||||||
|
for cell in row:
|
||||||
|
cell.value = None
|
||||||
|
START = 3
|
||||||
|
template_r = 3
|
||||||
|
cur_max = ws.max_row
|
||||||
|
for idx, item in enumerate(final, start=START):
|
||||||
|
if idx > cur_max:
|
||||||
|
for c in range(1, ws.max_column + 1):
|
||||||
|
srcc = ws.cell(template_r, c)
|
||||||
|
tgt = ws.cell(idx, c)
|
||||||
|
if srcc.has_style:
|
||||||
|
tgt.font = copy(srcc.font); tgt.fill = copy(srcc.fill)
|
||||||
|
tgt.border = copy(srcc.border); tgt.alignment = copy(srcc.alignment)
|
||||||
|
tgt.number_format = srcc.number_format; tgt.protection = copy(srcc.protection)
|
||||||
|
url = abs_url(item.get('href', ''))
|
||||||
|
ws.cell(idx, 2).value = idx - 2
|
||||||
|
ws.cell(idx, 3).value = name
|
||||||
|
for ci, c in enumerate('DEFGHIJ'):
|
||||||
|
ws.cell(idx, 4 + ci).value = item.get(c, '')
|
||||||
|
ws.cell(idx, 11).value = url
|
||||||
|
if is_external(url):
|
||||||
|
ws.cell(idx, 19).value = '외부링크'
|
||||||
|
END = START + len(final) - 1
|
||||||
|
center = Alignment(horizontal='center', vertical='center', wrap_text=True)
|
||||||
|
left = Alignment(horizontal='left', vertical='center', wrap_text=False)
|
||||||
|
|
||||||
|
def merge_runs(col_letter, col_idx, group_cols=()):
|
||||||
|
runs = []
|
||||||
|
cur_val = ws.cell(START, col_idx).value
|
||||||
|
cur_grp = tuple(ws.cell(START, g).value for g in group_cols)
|
||||||
|
run_start = START
|
||||||
|
for r in range(START + 1, END + 1):
|
||||||
|
v = ws.cell(r, col_idx).value
|
||||||
|
g = tuple(ws.cell(r, gg).value for gg in group_cols)
|
||||||
|
if v == cur_val and g == cur_grp:
|
||||||
|
continue
|
||||||
|
if cur_val not in (None, '') and r - 1 > run_start:
|
||||||
|
runs.append((run_start, r - 1))
|
||||||
|
cur_val, cur_grp, run_start = v, g, r
|
||||||
|
if cur_val not in (None, '') and END > run_start:
|
||||||
|
runs.append((run_start, END))
|
||||||
|
for s, e in runs:
|
||||||
|
ws.merge_cells(f'{col_letter}{s}:{col_letter}{e}')
|
||||||
|
ws.cell(s, col_idx).alignment = center
|
||||||
|
return len(runs)
|
||||||
|
|
||||||
|
n_f = merge_runs('F', 6, group_cols=(4, 5))
|
||||||
|
n_e = merge_runs('E', 5, group_cols=(4,))
|
||||||
|
n_d = merge_runs('D', 4)
|
||||||
|
for r in range(START, END + 1):
|
||||||
|
for c in (4, 5, 6, 7):
|
||||||
|
if ws.cell(r, c).value is not None:
|
||||||
|
ws.cell(r, c).alignment = center
|
||||||
|
for r in range(1, END + 1):
|
||||||
|
ws.row_dimensions[r].height = 15
|
||||||
|
link_n = 0
|
||||||
|
for r in range(START, END + 1):
|
||||||
|
cell = ws.cell(r, 11)
|
||||||
|
u = cell.value
|
||||||
|
if u and isinstance(u, str) and u.startswith(('http://', 'https://')):
|
||||||
|
cell.hyperlink = u
|
||||||
|
old = cell.font
|
||||||
|
cell.font = Font(name=old.name or '맑은 고딕', size=old.size or 11,
|
||||||
|
bold=old.bold, italic=old.italic, color='0000FF', underline='single')
|
||||||
|
cell.alignment = left
|
||||||
|
link_n += 1
|
||||||
|
wb.save(output)
|
||||||
|
ext_n = sum(1 for r in range(START, END + 1) if ws.cell(r, 19).value == '외부링크')
|
||||||
|
print(f' [{name}] 원본{len(raw_rows)}→중복{removed}→{len(final)}행 | 병합D{n_d}E{n_e}F{n_f} | 외부{ext_n} K링크{link_n} → {output}')
|
||||||
|
return len(final)
|
||||||
|
|
||||||
|
|
||||||
|
def load_probe():
|
||||||
|
with open(PROBE, encoding='utf-8') as f:
|
||||||
|
return {r['name']: r for r in json.load(f)}
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
probe = load_probe()
|
||||||
|
targets = sys.argv[1:] if len(sys.argv) > 1 else list(probe.keys())
|
||||||
|
for name in targets:
|
||||||
|
p = probe.get(name)
|
||||||
|
if not p:
|
||||||
|
print(f'[{name}] probe 정보 없음 — 스킵'); continue
|
||||||
|
ov = OVERRIDES.get(name, {})
|
||||||
|
num = int(p['num'])
|
||||||
|
base = p['base']
|
||||||
|
if ov.get('use_home'):
|
||||||
|
sm = p['home']
|
||||||
|
else:
|
||||||
|
sm = ov.get('sitemap') or p.get('sitemap') or p['home']
|
||||||
|
wait = ov.get('wait', 3000)
|
||||||
|
print(f"\n=== {num}.{name} === render {sm}")
|
||||||
|
try:
|
||||||
|
final, html = render(sm, wait=wait)
|
||||||
|
soup = BeautifulSoup(html, 'html.parser')
|
||||||
|
cont = pick_container(soup, ov.get('sel'))
|
||||||
|
if not cont:
|
||||||
|
print(f' [{name}] 컨테이너 못찾음 (a없음)'); continue
|
||||||
|
raw = extract_rows(cont)
|
||||||
|
dicts = rows_to_dicts(raw)
|
||||||
|
write_excel(name, num, base, dicts)
|
||||||
|
except Exception as e:
|
||||||
|
import traceback
|
||||||
|
print(f' [{name}] 실패: {e}')
|
||||||
|
traceback.print_exc()
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
321
_스크립트/_공공기관2_phase234.py
Normal file
321
_스크립트/_공공기관2_phase234.py
Normal file
@ -0,0 +1,321 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""공공기관 Phase 2~4: L(게시판형태)·M(수량)·N(저작물유형)·O/P/Q(공공누리).
|
||||||
|
|
||||||
|
시·군용 _chungnam_phase234_all.py 로직 재사용 + 범용화(넓은 본문셀렉터·다양한 상세패턴·오디오·이미지노이즈필터).
|
||||||
|
출력: 공공기관\{기관}.xlsx 의 L~Q 채움 (역순 실행 권장).
|
||||||
|
사용: python _공공기관_phase234.py [기관명 ...] (없으면 번호 역순 전체)
|
||||||
|
"""
|
||||||
|
import sys, os, re, json, time, warnings
|
||||||
|
from urllib.parse import urljoin, urlparse
|
||||||
|
from concurrent.futures import ThreadPoolExecutor, as_completed
|
||||||
|
import requests
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
import openpyxl
|
||||||
|
|
||||||
|
warnings.filterwarnings('ignore')
|
||||||
|
UA = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
||||||
|
H = {'User-Agent': UA}
|
||||||
|
OUTDIR = r'D:\01.프로젝트\DB수집\작업파일\공공기관2'
|
||||||
|
PROBE = r'D:\01.프로젝트\DB수집\_스크립트\_공공기관2_probe.json'
|
||||||
|
|
||||||
|
TOTAL_PAT = re.compile(r'총\s*(?:게시물|건수)?\s*[:\-]?\s*(\d[\d,]*)\s*(?:건|개|page|페이지|item)', re.I)
|
||||||
|
TOTAL_PAT2 = re.compile(r'(?:전체|총|total)\s*[:\-]?\s*(\d[\d,]*)', re.I)
|
||||||
|
KOGL_IMG_PAT = re.compile(r'(?:new_)?img_open(?:type|code)(\d{1,2})\.(?:png|jpe?g|gif)', re.I)
|
||||||
|
KOGL_LINK_PAT = re.compile(r'kogl\.or\.kr/info/licenseType(\d)', re.I)
|
||||||
|
YOUTUBE_PAT = re.compile(r'(?:youtube\.com|youtu\.be|vimeo\.com)', re.I)
|
||||||
|
VIDEO_EXT = re.compile(r'\.(mp4|webm|mov|avi|m3u8)(?:\?|$)', re.I)
|
||||||
|
AUDIO_EXT = re.compile(r'\.(mp3|wav|m4a|ogg|flac)(?:\?|$|["\'&])', re.I)
|
||||||
|
IMG_NOISE = re.compile(r'(ico[_\-/]|/icon|logo|btn|bul[_\-]|bg[_\-]|banner|sns|blank|spacer|loading|arrow|/dot|line[_\.]|top_|foot|header|common|btn_|_icon|symbol|copyright|qr_|movie_ico|no_img|noimage|share|facebook|insta|youtube_ic|blog|twitter|naver|kakao)', re.I)
|
||||||
|
DETAIL_PAT = re.compile(r'(mode=V|view\.do|/view|read\.do|/read|detail\.do|/detail|bbsView|nttId=|articleNo=|boardSeq=|bbtSn=|idx=|seq=|bIdx=|board_no=|wr_id=|dataSid=|ntceSn=|brdId=|bcIdx=|page_idx=|=view)', re.I)
|
||||||
|
|
||||||
|
BODY_SEL = [
|
||||||
|
'#content', '#contents', '#contentsArea', '.contentsArea', '.content', '.contents',
|
||||||
|
'#sub_content', '.sub_content', '.sub_contents', '#subContent', '.subContent',
|
||||||
|
'#container .content', '.board_view', '.bbs_view', '.view_con', '.view_cont',
|
||||||
|
'.board', '.bbs', '#bbs', '.sub_cont', '#cont', '.cont_area', '#content_area',
|
||||||
|
'main', '#main', 'article', '.board_list', '.bbs_list',
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def make_session(weak=False):
|
||||||
|
s = requests.Session()
|
||||||
|
s.headers.update(H)
|
||||||
|
return s
|
||||||
|
|
||||||
|
|
||||||
|
def fetch(session, url, timeout=12):
|
||||||
|
try:
|
||||||
|
r = session.get(url, timeout=timeout, verify=False, allow_redirects=True)
|
||||||
|
meta = re.search(rb'charset=["\']?\s*([\w-]+)', r.content[:4096], re.I)
|
||||||
|
r.encoding = meta.group(1).decode(errors='ignore') if meta else r.apparent_encoding
|
||||||
|
if r.status_code == 200:
|
||||||
|
return BeautifulSoup(r.text, 'html.parser'), r.url
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
return None, None
|
||||||
|
|
||||||
|
|
||||||
|
def get_body(soup, selectors):
|
||||||
|
for sel in selectors:
|
||||||
|
try:
|
||||||
|
el = soup.select_one(sel)
|
||||||
|
if el and len(el.get_text(strip=True)) > 20:
|
||||||
|
return el
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
return soup
|
||||||
|
|
||||||
|
|
||||||
|
def detect_form(body):
|
||||||
|
has_paging = bool(body.select('.pagination, .paging, nav.paging, .page_nav, .paginate, .pgwrap, .board_paging, .num_wrap'))
|
||||||
|
text_inputs = [i for i in body.find_all('input')
|
||||||
|
if (i.get('type') or 'text').lower() in ('text', 'search')]
|
||||||
|
has_search = len(text_inputs) >= 1
|
||||||
|
has_listtable = bool(body.select('table.board_list, table.bbs_list, ul.board_list, .board_list tbody tr, .bbs_list li'))
|
||||||
|
txt = body.get_text(' ', strip=True)
|
||||||
|
m = TOTAL_PAT.search(txt) or TOTAL_PAT2.search(txt)
|
||||||
|
total = None
|
||||||
|
if m:
|
||||||
|
digits = m.group(1).replace(',', '')
|
||||||
|
if digits.isdigit():
|
||||||
|
total = int(digits)
|
||||||
|
is_board = has_paging or has_listtable or (total is not None and (has_search or has_paging or has_listtable))
|
||||||
|
if is_board:
|
||||||
|
return '게시판', total if total is not None else 0
|
||||||
|
if has_search and (has_paging or has_listtable):
|
||||||
|
return '게시판', total if total is not None else 0
|
||||||
|
return '페이지', 1
|
||||||
|
|
||||||
|
|
||||||
|
def extract_detail_urls(body, base_url, limit=6):
|
||||||
|
urls, seen = [], set()
|
||||||
|
for a in body.find_all('a', href=True):
|
||||||
|
h = a['href']
|
||||||
|
if not h or h.startswith('#') or h.lower().startswith('javascript:'):
|
||||||
|
continue
|
||||||
|
if DETAIL_PAT.search(h):
|
||||||
|
full = urljoin(base_url, h)
|
||||||
|
if full not in seen:
|
||||||
|
seen.add(full); urls.append(full)
|
||||||
|
if len(urls) >= limit:
|
||||||
|
break
|
||||||
|
return urls
|
||||||
|
|
||||||
|
|
||||||
|
def detect_media(body):
|
||||||
|
has_text = len(body.get_text(strip=True)) > 30
|
||||||
|
has_image = False
|
||||||
|
for img in body.find_all('img'):
|
||||||
|
src = img.get('src') or img.get('data-src') or ''
|
||||||
|
if not src or KOGL_IMG_PAT.search(src) or IMG_NOISE.search(src):
|
||||||
|
continue
|
||||||
|
w = img.get('width', '')
|
||||||
|
try:
|
||||||
|
if w and int(re.sub(r'\D', '', w) or 0) and int(re.sub(r'\D', '', w)) < 60:
|
||||||
|
continue
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
has_image = True
|
||||||
|
break
|
||||||
|
has_video = bool(body.find_all('video'))
|
||||||
|
if not has_video:
|
||||||
|
for ifr in body.find_all('iframe'):
|
||||||
|
if YOUTUBE_PAT.search(ifr.get('src', '')):
|
||||||
|
has_video = True; break
|
||||||
|
if not has_video:
|
||||||
|
for a in body.find_all('a', href=True):
|
||||||
|
if YOUTUBE_PAT.search(a['href']):
|
||||||
|
has_video = True; break
|
||||||
|
if not has_video and VIDEO_EXT.search(str(body)):
|
||||||
|
has_video = True
|
||||||
|
has_audio = bool(body.find_all('audio')) or bool(AUDIO_EXT.search(str(body)))
|
||||||
|
return has_image, has_video, has_audio, has_text
|
||||||
|
|
||||||
|
|
||||||
|
def n_string(has_text, has_image, has_video, has_audio):
|
||||||
|
parts = []
|
||||||
|
if has_text:
|
||||||
|
parts.append('어문')
|
||||||
|
if has_image:
|
||||||
|
parts.append('이미지')
|
||||||
|
if has_video:
|
||||||
|
parts.append('영상')
|
||||||
|
if has_audio:
|
||||||
|
parts.append('오디오')
|
||||||
|
return ','.join(parts) if parts else '없음'
|
||||||
|
|
||||||
|
|
||||||
|
def img_has_valid_anchor(img):
|
||||||
|
p = img.parent
|
||||||
|
while p is not None:
|
||||||
|
if p.name == 'a':
|
||||||
|
href = p.get('href', '')
|
||||||
|
return bool(href and not href.startswith('#') and not href.lower().startswith('javascript:'))
|
||||||
|
p = p.parent
|
||||||
|
return False
|
||||||
|
|
||||||
|
|
||||||
|
def detect_kogl(body):
|
||||||
|
types, q_y = set(), False
|
||||||
|
for a in body.find_all('a', href=True):
|
||||||
|
m = KOGL_LINK_PAT.search(a['href'])
|
||||||
|
if m:
|
||||||
|
types.add(int(m.group(1))); q_y = True
|
||||||
|
for img in body.find_all('img'):
|
||||||
|
m = KOGL_IMG_PAT.search(img.get('src', ''))
|
||||||
|
if m and 1 <= int(m.group(1)) <= 4:
|
||||||
|
types.add(int(m.group(1)))
|
||||||
|
if img_has_valid_anchor(img):
|
||||||
|
q_y = True
|
||||||
|
for el in body.find_all(style=True):
|
||||||
|
m = KOGL_IMG_PAT.search(el.get('style', ''))
|
||||||
|
if m and 1 <= int(m.group(1)) <= 4:
|
||||||
|
types.add(int(m.group(1)))
|
||||||
|
# 1·2·3·4 전부 = 범례페이지 = 미부착
|
||||||
|
if types == {1, 2, 3, 4}:
|
||||||
|
return set(), None
|
||||||
|
if not types:
|
||||||
|
return set(), None
|
||||||
|
return types, ('Y' if q_y else 'N')
|
||||||
|
|
||||||
|
|
||||||
|
def process_row(session, url, body_selectors):
|
||||||
|
out = {'L': '', 'M': '', 'N': '', 'O': '', 'P': '', 'Q': '', 'note': ''}
|
||||||
|
soup, final = fetch(session, url)
|
||||||
|
if soup is None:
|
||||||
|
out['note'] = '접근 실패'
|
||||||
|
return out
|
||||||
|
body = get_body(soup, body_selectors)
|
||||||
|
form, count = detect_form(body)
|
||||||
|
out['L'] = form
|
||||||
|
out['M'] = count if form == '게시판' else 1
|
||||||
|
has_img, has_vid, has_aud, has_txt = detect_media(body)
|
||||||
|
types_main, q_main = detect_kogl(body)
|
||||||
|
P = '게시판' if types_main else ''
|
||||||
|
types_all = set(types_main)
|
||||||
|
q_flags = [q_main] if q_main else []
|
||||||
|
if form == '게시판':
|
||||||
|
for du in extract_detail_urls(body, final or url, limit=6):
|
||||||
|
d_soup, _ = fetch(session, du, timeout=10)
|
||||||
|
if not d_soup:
|
||||||
|
continue
|
||||||
|
d_body = get_body(d_soup, body_selectors)
|
||||||
|
di, dv, da, dt = detect_media(d_body)
|
||||||
|
has_img |= di; has_vid |= dv; has_aud |= da; has_txt |= dt
|
||||||
|
dt_types, dt_q = detect_kogl(d_body)
|
||||||
|
if dt_types and not types_main and not P:
|
||||||
|
P = '게시물'
|
||||||
|
types_all |= dt_types
|
||||||
|
if dt_q:
|
||||||
|
q_flags.append(dt_q)
|
||||||
|
out['N'] = n_string(has_txt, has_img, has_vid, has_aud)
|
||||||
|
if not types_all:
|
||||||
|
out['O'] = '미부착'
|
||||||
|
else:
|
||||||
|
out['O'] = ','.join(f'{n}유형' for n in sorted(types_all))
|
||||||
|
out['P'] = P if P else '게시판'
|
||||||
|
out['Q'] = 'Y' if 'Y' in q_flags else 'N'
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def run_site(name, num, body_selectors, workers=8):
|
||||||
|
xlsx = os.path.join(OUTDIR, f'{num}.{name}', f'{name}.xlsx')
|
||||||
|
if not os.path.exists(xlsx):
|
||||||
|
print(f'[{name}] 파일 없음 — 스킵'); return None
|
||||||
|
wb = openpyxl.load_workbook(xlsx)
|
||||||
|
ws = wb.active
|
||||||
|
START = 3
|
||||||
|
END = START - 1
|
||||||
|
for r in range(START, ws.max_row + 1):
|
||||||
|
if ws.cell(r, 2).value is None:
|
||||||
|
break
|
||||||
|
END = r
|
||||||
|
if END < START:
|
||||||
|
print(f'[{name}] 데이터행 없음'); return None
|
||||||
|
# 이미 처리됨(이어하기): L열 채워진 비율 ≥90%면 스킵
|
||||||
|
filled = sum(1 for r in range(START, END + 1) if ws.cell(r, 12).value)
|
||||||
|
if '--force' not in sys.argv and filled >= (END - START + 1) * 0.9:
|
||||||
|
print(f'[{num}.{name}] 이미 처리됨({filled}/{END-START+1}) — 스킵')
|
||||||
|
return {'name': name, 'forms': {'(skip)': filled}, 'attach': 0, 'rows': END - START + 1}
|
||||||
|
tasks = []
|
||||||
|
for r in range(START, END + 1):
|
||||||
|
url = ws.cell(r, 11).value
|
||||||
|
is_ext = (ws.cell(r, 19).value == '외부링크')
|
||||||
|
tasks.append((r, url, is_ext))
|
||||||
|
n_ext = sum(1 for t in tasks if t[2])
|
||||||
|
print(f'\n[{num}.{name}] {len(tasks)}행 (외부 {n_ext}) 처리…')
|
||||||
|
t0 = time.time()
|
||||||
|
session = make_session()
|
||||||
|
results = {}
|
||||||
|
|
||||||
|
def worker(task):
|
||||||
|
row, url, is_ext = task
|
||||||
|
if is_ext:
|
||||||
|
return row, {'L': '사이트', 'M': '', 'N': '', 'O': '', 'P': '', 'Q': '', 'note': ''}
|
||||||
|
if not url or not isinstance(url, str) or not url.startswith('http'):
|
||||||
|
return row, {'L': '', 'M': '', 'N': '', 'O': '', 'P': '', 'Q': '', 'note': 'URL 없음'}
|
||||||
|
return row, process_row(session, url, body_selectors)
|
||||||
|
|
||||||
|
with ThreadPoolExecutor(max_workers=workers) as ex:
|
||||||
|
futs = [ex.submit(worker, t) for t in tasks]
|
||||||
|
done = 0
|
||||||
|
for fut in as_completed(futs):
|
||||||
|
row, res = fut.result()
|
||||||
|
results[row] = res
|
||||||
|
done += 1
|
||||||
|
if done % 50 == 0 or done == len(tasks):
|
||||||
|
print(f' {done}/{len(tasks)} ({time.time()-t0:.0f}s)')
|
||||||
|
|
||||||
|
for r in range(START, END + 1):
|
||||||
|
res = results.get(r)
|
||||||
|
if not res:
|
||||||
|
continue
|
||||||
|
if res.get('L'):
|
||||||
|
ws.cell(r, 12).value = res['L']
|
||||||
|
if res.get('M') != '':
|
||||||
|
ws.cell(r, 13).value = res['M']
|
||||||
|
if res.get('N'):
|
||||||
|
ws.cell(r, 14).value = res['N']
|
||||||
|
if res.get('O'):
|
||||||
|
ws.cell(r, 15).value = res['O']
|
||||||
|
if res.get('P'):
|
||||||
|
ws.cell(r, 16).value = res['P']
|
||||||
|
if res.get('Q'):
|
||||||
|
ws.cell(r, 17).value = res['Q']
|
||||||
|
if res.get('note') and not ws.cell(r, 19).value:
|
||||||
|
ws.cell(r, 19).value = res['note']
|
||||||
|
wb.save(xlsx)
|
||||||
|
forms, attach = {}, 0
|
||||||
|
for res in results.values():
|
||||||
|
forms[res.get('L', '')] = forms.get(res.get('L', ''), 0) + 1
|
||||||
|
if res.get('O') and res.get('O') != '미부착':
|
||||||
|
attach += 1
|
||||||
|
fs = ' '.join(f'{k}{v}' for k, v in forms.items() if k)
|
||||||
|
print(f' ✓ [{name}] {fs} | 공공누리부착 {attach} ({time.time()-t0:.0f}s)')
|
||||||
|
return {'name': name, 'forms': forms, 'attach': attach, 'rows': len(tasks)}
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
probe = {r['name']: r for r in json.load(open(PROBE, encoding='utf-8'))}
|
||||||
|
order = sorted(probe.values(), key=lambda x: -int(x['num'])) # 번호 역순
|
||||||
|
only = sys.argv[1:]
|
||||||
|
if only:
|
||||||
|
order = [p for p in order if p['name'] in only or str(p['num']) in only]
|
||||||
|
summ = []
|
||||||
|
for p in order:
|
||||||
|
try:
|
||||||
|
r = run_site(p['name'], int(p['num']), BODY_SEL)
|
||||||
|
if r:
|
||||||
|
summ.append(r)
|
||||||
|
except Exception as e:
|
||||||
|
import traceback
|
||||||
|
print(f"[{p['name']}] 실패: {e}")
|
||||||
|
traceback.print_exc()
|
||||||
|
print('\n=== 요약 ===')
|
||||||
|
for s in summ:
|
||||||
|
fs = ' '.join(f'{k}{v}' for k, v in s['forms'].items() if k)
|
||||||
|
print(f" {s['name']}: {fs} | 부착{s['attach']} /{s['rows']}행")
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
Some files were not shown because too many files have changed in this diff Show More
Loading…
Reference in New Issue
Block a user