|
![]() ![]()
|
There are a few different issues to consider when tuning the performance of Berkeley DB access method applications.
(액세스 메소드 성능과 관련된 튜닝 이슈들..)
(액세스 메소드는 성능에 중대한 영향을 미친다.고정길이의 레코드나 정수키를 사용하는 애플은 큐를 사용할때 성능이 높다.가변길이의 레코드일경우는 Btree메소드가 성능이 높고 대부분의 애플에서 해시나 Recno를 사용하는 경우보다 일반적으로 빠른 경향이 있다.액세스 메소드 API는 비슷하기때문에 엑세스메소드를 바꿔가면서 벤치마킹이 쉽다.)
(디폴트로 버클리 디비는 디비환경의 공유역역들을 메모리 백업 파일시스템에 생성한다.어떤 시스템에서는 dirty 페이지를 디스크로 플러싱할때 일반 파일시스템 페이지와 메모리 맵 페이지를 구별하지 않는다.이같은 이유로 버클리디비 캐시의 dirty 페이지는 파일시스템동작을 많이 발생시킬 수 있다.특히 쓰레드와 프로세스를 동기화하는 파일시스템일때는 더욱 그러하다.어떤경우에는 이것은 애플의 처리속도에 극적으로 많이 영향을 미칠수있다.이러한 문제를 회피하기 위해서는 시스템공유메모리(DB_SYSTEM_MEM )나 애플리케이션 프라이빗 메모리(DB_PRIVATE)에 공유영역을 생성하거나 이러한 문제에 대해 시스템적으로 설정가능하다면 메모리 맵페이지의 플러싱을 off시켜야 한다.)
(큰 키/데이타를 디비에 저장하는 것은 Btree,Hash,Recno디비의 성능특성을 변경시킨다.첫번째로 생각해볼 것은 디비 페이지 크기다.키/데이타 아이템이 너무크면 하나의 디비페이지에 위치될수 없다.이것은 일반적 디비 구조 밖의 오퍼블로우 페이지에 저장된다.(특히 페이지크기의 1/4보다 큰 아이템은 너무 크다고 생각해야 한다.).이 오버플로우 페이지를 접근하기 위해서는 일반적인 액세스에 비해 최소한 하나의 추가적 페이지 참조가 필요하다.그래서 많은 수의 오버플로우 페이지를 가진 디비를 생성하기보다는 페이지 크기를 늘리는것이 더 좋다.디비의 오버플러우 페이지수를 보기위해 db_stat utility (or the statistics returned by the DB->stat method)를 사용하라.)
The second issue is using large key/data items instead of duplicate data items. While this can offer performance gains to some applications (because it is possible to retrieve several data items in a single get call), once the key/data items are large enough to be pushed off-page, they will slow the application down. Using duplicate data items is usually the better choice in the long run.
(두번째 이슈는 복사된 데이타 아이템 대신 큰 키/데이타 아이템을 사용하는 문제다.이러한 사용이 단일 함수 호출로 여러개의 데이타 아이템을 가져올 가능성이 있기 때 어떤애플의 성능향상에 도움을 줄수도 있지만 키/데이타 아이템이 오프페이지에 푸시될만큼 크다면 결국 애플을 느리케한다.복사된 데이타 아이템은 오랜 수행을 할때는 더좋은 선택이다.)
A common question when tuning Berkeley DB applications is scalability. For example, people will ask why, when adding additional threads or processes to an application, the overall database throughput decreases, even when all of the operations are read-only queries.
(버클리 디비 애플 튜닝에 대한 흔한 질문중 하나는 확장성이다.예를들어 사람들은 왜 그리고 언제 쓰레드나 프로세스가 애플에 추가되어야 하는지 그리고 (심지어 모든 오퍼레이션이 읽기전용 쿼리일때도)전체디비의 처리량의 감소시키는지를 물어 볼것이다.)
First, while read-only operations are logically concurrent, they still have to acquire mutexes on internal Berkeley DB data structures. For example, when searching a linked list and looking for a database page, the linked list has to be locked against other threads of control attempting to add or remove pages from the linked list. The more threads of control you add, the more contention there will be for those shared data structure resources.
(첫째,읽기전용 오퍼레이션이 논리적인 동시성을 갖는다면 그것들은 버클리디비 내부적으로 뮤텍스를 얻는 동작을 하게 된다.예를 들어 링크드리스트를 검색하고 디비페이지를 찾을때 링크드리스는 링크드리스트에 페이지(역자주:링크드리스트의 아이템)의 추가삭제를 뮤텍스로 막하야 한다.더많은 쓰레드를 애플에 추가할수록 공유리소스를 쓰기위한 더많은 경쟁이 있게 된다.)
Second, once contention starts happening, applications will also start to see threads of control convoy behind locks (especially on architectures supporting only test-and-set spin mutexes, rather than blocking mutexes). On test-and-set architectures, threads of control waiting for locks must attempt to acquire the mutex, sleep, check the mutex again, and so on. Each failed check of the mutex and subsequent sleep wastes CPU and decreases the overall throughput of the system.
(두번째로 경쟁의 시작이 발생하게 되면 애플은 또한 락을 얻기위한 동작이 실행된다(뮤텍스를 블락킹하기보다 테스트-셋 스핀 뮤텍스만을 지원하는 아키텍서에서는 더더욱).테스트-셋 아키텍쳐에서는 락을 기다리는 쓰레드는 뮤텍스얻기시도,슬립,다시 뮤텍스얻기시도,슬립을 계속반복해야만 한다.뮤텍스얻기의 각각의 실패와 이어지는 슬립은 CPU를 소모시키고 전체적인 처리량을 감소시킨다.)
Third, every time a thread acquires a shared mutex, it has to shoot down other references to that memory in every other CPU on the system. Many modern snoopy cache architectures have slow shoot down characteristics.
(세번째..공유된 뮤텍스를 어던 쓰레드가 얻을때마다 이것(뮤텍스얻기)는 그 시스템의 다른 CPU에 대해 다른 메모리 참조를 shoot down시켜야 한다.현대의 많은 스누피캐시 아키텍쳐는 slow 슈다운 특성을 가지고 있다.)
Fourth, schedulers don't care what application-specific mutexes a thread of control might hold when de-schedule a thread. If a thread of control is descheduled while holding a shared data structure mutex, other threads of control will be blocked until the scheduler decides to run the blocking thread of control again. The more threads of control that are running, the smaller their quanta of CPU time, and the more likely they will be descheduled while holding a Berkeley DB mutex.
(네번째,스케쥴러는 쓰레드를 de-스케쥴링할때 어떤쓰레드가 가지고 있는 뮤텍스에 대해 고려하지 않는다.만약 뮤텍스를 가진 쓰레드가 de-스케쥴링되면 뮤텍스를 가진 쓰레드가 다시 스케쥴링될때까지 모든 다른 쓰레드는 블락킹되게 된다.더많은 쓰레드가 돌수록 CPU점유시간은 더 적고 뮤텍스를 가지고 있을동안 더 많이 de스케쥴링 될것이다.)
The results of adding new threads of control to an application, on the application's throughput, is application and hardware specific and almost entirely dependent on the application's data access pattern and hardware. In general, using operating systems that support blocking mutexes will often make a tremendous difference, and limiting threads of control to to some small multiple of the number of CPUs is usually the right choice to make.
(쓰레드를 애플에 추가시켰을때의 결과는 애플,하드웨어에 따라 다르고 거의 전체적으로 애플리케이션 액세스패턴과 하드웨어에 의존적이다.일반적으로 블락킹뮤텍스를 지원하는 운영체계를 사용하는 것은 대단한 차이점을 만든다(역자주:더좋다는 소리다.). 그리고 쓰레드수를 CPU수와 같을 정도로 적게 제한하는 것은 보통 좋은 선택이다. )
![]() ![]()
|
Copyright (c) 1996-2003 Sleepycat Software, Inc. - All rights reserved.